Resource scheduling method and device of energy system

By acquiring multi-source data of the energy system, combining multi-value stream reward function and deep Q network model to optimize resource scheduling strategies, the problem of low scheduling accuracy of traditional energy systems is solved, and higher scheduling accuracy and reliability are achieved.

CN120373813AInactive Publication Date: 2025-07-25GUANGZHOU RIMSEA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510864621.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional energy system resource scheduling solutions rely on single reference data and simple logic, resulting in low accuracy of scheduling strategies and difficult to meet the operation needs of complex factors.

Method used

By obtaining the operating data of the energy system, environmental data and power market fluctuation data, using the preset initial resource scheduling strategy as the basis of the current time step, combining the multi-value stream reward function, optimizing and adjusting the resource scheduling strategy to maximize the value of the multi-value stream reward function, and using technical means such as the deep Q network model and ant colony optimization algorithm to dynamically update the scheduling strategy.

Benefits of technology

It improves the accuracy and reliability of energy system resource scheduling, can better adapt to changes in system state, and achieve optimal balance in many aspects such as economic benefits and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373813A_ABST
    Figure CN120373813A_ABST
Patent Text Reader

Abstract

The invention relates to a resource scheduling method and device for an energy system. The method comprises the steps of obtaining operation data, environment data and power market fluctuation data of an energy system, taking a preset initial resource scheduling strategy as a resource scheduling strategy of a current time step, and according to the operation data, the environment data and the power market fluctuation data of the energy system, taking maximization of a preset multi-value flow reward function as a target; and updating the resource scheduling strategy of the current time step into a resource scheduling strategy which enables a preset multi-value flow reward function to obtain a maximum value, and scheduling resources of the energy system based on the resource scheduling strategy of the current time step. By adopting the method, the resource scheduling accuracy of the energy system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of resource scheduling, and in particular, to a resource scheduling method, device, computer device, storage medium, and computer program product for an energy system. Background Art

[0002] With the rapid development of renewable energy and distributed energy, the scale and quantity of renewable energy facilities such as solar power stations and wind power plants have been continuously increasing. Small biomass power generation devices and distributed energy storage devices are also increasingly connected to the energy system, resulting in a significant growth trend in the resource scheduling requirements of the energy system.

[0003] Traditional resource scheduling schemes for energy systems mainly rely on pre-set plans and empirical rules. For example, based on the historical load data, seasonal changes, and approximate electricity demand in the future of the energy system, the resource scheduling strategy of the energy system is formulated in advance to ensure the stable supply of electricity.

[0004] However, the reference data for the resource scheduling strategy in the traditional scheme is relatively single, and the logic for formulating the strategy is simple. The factors affecting the operation of the energy system are relatively complex, and the resource scheduling strategy obtained by the traditional scheme is difficult to meet the operation requirements of the energy system, that is, the accuracy of the traditional resource scheduling scheme is relatively low. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a resource scheduling method, device, computer device, computer-readable storage medium, and computer program product for an energy system that can improve the accuracy of resource scheduling for the energy system.

[0006] In a first aspect, the present application provides a resource scheduling method for an energy system. The method includes:

[0007] Obtain the operation data, environmental data, and power market fluctuation data of the energy system;

[0008] Use a preset initial resource scheduling strategy as the resource scheduling strategy for the current time step;

[0009] According to the operation data, environmental data, and power market fluctuation data of the energy system, with the goal of maximizing a preset multi-value flow reward function, update the resource scheduling strategy for the current time step to a resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value;

[0010] Based on the resource scheduling strategy for the current time step, schedule the resources of the energy system.

[0011] In a second aspect, the present application also provides a resource scheduling device for an energy system. The device includes:

[0012] A data acquisition module, configured to acquire the operation data, environmental data, and power market fluctuation data of the energy system;

[0013] A strategy formulation module, configured to use a preset initial resource scheduling strategy as the resource scheduling strategy for the current time step;

[0014] A strategy adjustment module, configured to update the resource scheduling strategy for the current time step to the resource scheduling strategy that maximizes the preset multi-value flow reward function according to the operation data, environmental data, and power market fluctuation data of the energy system, with the goal of maximizing the preset multi-value flow reward function;

[0015] A resource scheduling module, configured to schedule the resources of the energy system based on the resource scheduling strategy for the current time step.

[0016] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the embodiment of the resource scheduling method for the energy system described above are implemented.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the embodiment of the resource scheduling method for the energy system described above are implemented.

[0018] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the embodiment of the resource scheduling method for the energy system described above are implemented.

[0019] The above-mentioned resource scheduling method, device, computer device, storage medium, and computer program product for the energy system, different from the traditional solutions, by acquiring the operation data, environmental data, and power market fluctuation data of the energy system, these multi-source data cover various factors affecting the operation of the energy system, providing reliable data support for subsequent formulation of resource scheduling strategies. Further, according to the operation data, environmental data, and power market fluctuation data of the energy system, using a preset initial resource scheduling strategy as the resource scheduling strategy for the current time step, with the goal of maximizing the preset multi-value flow reward function, updating the resource scheduling strategy for the current time step to the resource scheduling strategy that maximizes the preset multi-value flow reward function. The updated resource scheduling strategy can better fit the state of the energy system at the current time step and can maximize the comprehensive benefits of the energy system as much as possible. Finally, executing the resource scheduling strategy for the current time step can improve the accuracy and reliability of the resource scheduling of the energy system. Description of the Drawings

[0020] Figure 1 It is an application environment diagram of the resource scheduling method for the energy system in an embodiment;

[0021] Figure 2 It is a schematic flowchart of the resource scheduling method for the energy system in an embodiment;

[0022] Figure 3 It is a schematic flowchart of the resource scheduling method for the energy system in another embodiment;

[0023] Figure 4 It is a schematic flowchart of the resource scheduling method for the energy system in yet another embodiment;

[0024] Figure 5 It is a schematic flowchart of the steps for updating the resource scheduling strategy at the current time step in an embodiment;

[0025] Figure 6 It is a schematic flowchart of the steps for formulating the initial resource scheduling strategy in an embodiment;

[0026] Figure 7 It is a schematic flowchart of the steps for adjusting the model parameters of the data prediction model in an embodiment;

[0027] Figure 8 It is a schematic flowchart of the resource scheduling method for the energy system in a detailed embodiment;

[0028] Figure 9 It is a structural block diagram of the deep Q - network model in an embodiment;

[0029] Figure 10 It is a structural block diagram of the resource scheduling device for the energy system in an embodiment;

[0030] Figure 11 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0031] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the content of this application will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0032] The resource scheduling method for the energy system provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers.

[0033] Specifically, it can be that the operator sends the operation data, environmental data, and power market fluctuation data of the energy system to the server 104 through the terminal 102. Then, the server 104 uses the preset initial resource scheduling strategy as the resource scheduling strategy for the current time step. Then, the server 104 updates the resource scheduling strategy for the current time step to the resource scheduling strategy that maximizes the preset multi-value flow reward function according to the operation data, environmental data, and power market fluctuation data of the energy system. Finally, the server 104 schedules the resources of the energy system based on the resource scheduling strategy for the current time step.

[0034] Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0035] In one embodiment, as Figure 2 shown, a resource scheduling method for an energy system is provided. Taking the method applied to the Figure 1 server 104 in it as an example for illustration, it includes the following steps:

[0036] S100, obtain the operation data, environmental data, and power market fluctuation data of the energy system.

[0037] Among them, the energy system includes but is not limited to renewable power generation energy devices, distributed energy generation devices, etc. The operation data of the energy system includes but is not limited to the power generation amount, power generation power, charge and discharge status, power quantity, and other operation parameters of various devices in the energy system (such as power generation devices, power transmission devices, energy storage devices, etc.), which reflect the operation status of the energy system. The environmental data of the energy system includes but is not limited to data such as light intensity, wind speed, temperature, and humidity. The renewable energy generation devices in the energy system are easily affected by environmental data. For example, light intensity affects the power generation power of solar power generation devices, and wind speed affects the power generation power of wind power generation devices, etc. The power market fluctuation data includes but is not limited to the electricity price fluctuation data, power supply and demand change data, etc. in the power market. The electricity price and power supply and demand will change continuously with factors such as the supply and demand balance and power generation cost in the power market, thus affecting the resource scheduling of the energy system.

[0038] It should be noted that the data obtained above can be the data obtained after preprocessing the original data. The corresponding original data includes, but is not limited to, original operation data, original environmental data, and original power market fluctuation data. These original data can be directly collected by sensors deployed in the energy system. Specifically, the preprocessing methods for the original data include, but are not limited to, feature engineering, outlier handling, data augmentation, etc.

[0039] S200, use the preset initial resource scheduling strategy as the resource scheduling strategy for the current time step.

[0040] Among them, the preset initial resource scheduling strategy refers to the preliminary energy resource allocation and scheduling plan formulated in advance, including, but not limited to, the power generation power allocation plan of power generation equipment in the energy system, the charge and discharge plan of energy storage equipment, the power transmission path planning, etc. The initial resource scheduling strategy is the basis for subsequent strategy optimization.

[0041] Specifically, it can be to analyze the initial operating state of the energy system, and then combine the basic constraints of the energy system (such as the power upper limit of power generation equipment, the capacity limit of energy storage equipment, etc.) to formulate the initial resource scheduling strategy in advance. For example, if it is predicted that the wind power generation will increase significantly and the electricity price is at a high level during a certain period, the preliminary resource scheduling strategy formulated can include increasing the proportion of wind power supply during this period; if it is predicted that the electricity price will be high during a future peak electricity consumption period, but the power supply of renewable energy is unstable during this period, the preliminary resource scheduling strategy formulated can include using more stable thermal power supply during this period; if it is predicted that the power generation of photovoltaic power generation equipment in the energy system far exceeds the local load during a certain period and the electricity price is low at this time, the preliminary resource scheduling strategy formulated can include storing the excess power generation of the photovoltaic power generation equipment into the energy storage equipment.

[0042] In the process of formulating the initial resource scheduling strategy, optimization algorithms such as genetic algorithm, particle swarm optimization algorithm, and ant colony optimization algorithm can be used for iterative calculation. In the iterative optimization process, the optimization algorithm will continuously search for optional resource scheduling strategies and evaluate the advantages and disadvantages of each resource scheduling strategy. After multiple optimization iterations, the optimization algorithm will gradually converge, and finally retain the optimal resource scheduling strategy as the initial resource scheduling strategy.

[0043] S300, according to the operation data, environmental data, and power market fluctuation data of the energy system, with the goal of maximizing the preset multi-value flow reward function, update the resource scheduling strategy of the current time step to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value.

[0044] Among them, the multi-value stream reward function is used to measure the performance of the adjustment actions taken during the process of adjusting the resource scheduling strategy on multiple value streams. For example, in the resource scheduling scenario of an energy system, the value streams may include, but are not limited to, market trading revenue, grid stability, load balancing degree, deviation from the target load curve, etc. The target load curve is a preset and ideal load change curve, which can be determined based on the historical operation data of the energy system, seasonal changes, electricity consumption habits, etc., and is used to characterize the desired load operation state of the energy system. The closer the actual load change curve is to the target load curve, the closer the current operation state of the energy system is considered to be to the target operation state. The updated resource scheduling strategy at the current time step is the optimized adjustment plan, and this resource scheduling strategy at the current time step can make the preset multi-value stream reward function reach the maximum value, achieving the comprehensive optimum of the energy system on multiple value streams at the current time step.

[0045] Specifically, according to the actual needs and goals of the energy system, the value streams affecting the performance of the energy system can be determined, and different weights can be assigned to different value streams to indicate different degrees of importance of different value streams. For example, if the current energy system pays more attention to the load balancing degree, a higher weight can be assigned to the load balancing degree in the multi-value stream reward function. Exemplarily, during the process of adjusting the resource scheduling strategy at the current time step, the server will generate candidate adjustment actions for the resource scheduling strategy at the current time step, calculate the multi-value stream reward function values corresponding to the candidate adjustment actions, and update the resource scheduling strategy at the current time step to the resource scheduling strategy that makes the preset multi-value stream reward function obtain the maximum value.

[0046] S400, based on the resource scheduling strategy at the current time step, schedule the resources of the energy system.

[0047] Following the above steps, the resource scheduling strategy at the current time step is the result obtained by comprehensively considering the real-time state of the energy system and optimizing with the goal of maximizing the multi-value stream reward function. By executing this resource scheduling strategy at the current time step, the optimal balance of the energy system in terms of economic benefits, stability, etc. can be achieved.

[0048] Specifically, the server can convert the target resource scheduling strategy into specific control instructions, and then convey these control instructions to each execution unit or each energy device in the energy system through the energy management system (EMS). The energy devices perform corresponding operation adjustments according to the received instructions. For example, the power generation device generates electricity according to the set power generation power and power generation time, and the energy storage device charges or discharges at the set time, etc.

[0049] The resource scheduling method of the above energy system, different from the traditional solution, obtains the operation data, environment data and power market fluctuation data of the energy system. These multi-source data cover various factors affecting the operation of the energy system, providing reliable data support for formulating subsequent resource scheduling strategies. Further, according to the operation data, environment data and power market fluctuation data of the energy system, with the preset initial resource scheduling strategy as the resource scheduling strategy for the current time step, aiming to maximize the preset multi-value flow reward function, the resource scheduling strategy for the current time step is updated to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value. The updated resource scheduling strategy can better fit the state of the energy system at the current time step and can maximize the comprehensive benefit of the energy system as much as possible. Finally, executing the resource scheduling strategy for the current time step can improve the accuracy and reliability of the resource scheduling of the energy system.

[0050] In one embodiment, after S400, the method further includes: updating the operation data, environment data and power market fluctuation data of the energy system, and returning to the step of S300 until a preset iteration end condition is reached.

[0051] Among them, theoretically speaking, as long as the energy system is in a working state, the above process of iteratively adjusting the resource scheduling strategy can be continuously looped. In fact, the preset iteration end condition can be that the energy system stops working and there is no longer a need for resource scheduling, or the operation data, environment data and power market fluctuation data of the energy system have reached a relatively stable state. In order to reduce the computing amount of the servers in the energy system, the resource scheduling strategy can no longer be adjusted so frequently.

[0052] It should be noted that after implementing the resource scheduling strategy for the current time step, the operation data of the energy system may change, and the environmental data of the energy system and the power market fluctuation data are constantly changing. Therefore, if the same resource scheduling strategy is always used, the resource scheduling strategy may become unsuitable for the current state of the energy system as the operation data, environmental data, and power market fluctuation data of the energy system change, and it may be difficult to meet the resource scheduling requirements of the energy system. Therefore, in this embodiment, with the time step as the minimum unit, for each time step, the operation data, environmental data, and power market fluctuation data of the energy system are updated, and within this time step, based on the updated operation data, environmental data, and power market fluctuation data, with the goal of maximizing the preset multi-value flow reward function, the resource scheduling strategy for the current time step is adjusted. In each adjustment, the server generates different candidate adjustment actions (such as changing the power generation power distribution of power generation equipment, adjusting the charge and discharge time of energy storage equipment, etc.), calculates the multi-value flow reward function values corresponding to each candidate adjustment action, and updates the resource scheduling strategy for the current time step to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value. Then, the data is updated again, and the above strategy optimization steps are repeated until the preset iteration end condition is met.

[0053] In this embodiment, by continuously optimizing the resource scheduling strategy according to the latest operation data, environmental data, and power market fluctuation data of the energy system, with the optimization goal of achieving the optimal balance of the energy system on the multi-value flow, the updated resource scheduling strategy can better meet the requirements of the energy system, improve the accuracy and reliability of the resource scheduling strategy, and thus improve the reliability and accuracy of resource scheduling.

[0054] In one embodiment, as Figure 3 shown, S300 includes:

[0055] S310, based on the operation data, environmental data, power market fluctuation data of the energy system, and the preset multi-value flow reward function, determine multiple candidate adjustment actions for the resource scheduling strategy for the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions.

[0056] S320, select the target candidate adjustment action when the multi-value reward function value reaches the maximum from the multiple candidate adjustment actions, and update the resource scheduling strategy for the current time step based on the target candidate adjustment action.

[0057] Among them, the operation data, environmental data, and power market fluctuation data of the energy system can jointly construct a high-dimensional space, also known as the state space, which is used to describe the comprehensive state of the energy system at a certain moment. Candidate adjustment actions refer to a series of adjustment measures that can be taken for the resource scheduling strategy at the current time step, such as adjusting the power generation of power generation equipment, changing the charge and discharge plan of energy storage equipment, etc. Candidate adjustment actions include but are not limited to adjusting the battery charge and discharge power, allocating photovoltaic power generation resources, buying / selling electricity, adjusting load distribution, etc. The target candidate adjustment action refers to the candidate adjustment action that maximizes the multi-value reward function value among multiple candidate adjustment actions, representing the most favorable adjustment direction for the resource scheduling strategy at the current time step in the current state space of the energy system.

[0058] Exemplarily, the multi-value flow optimization problem in this embodiment can be modeled as a Markov decision process in reinforcement learning. The operation data of the energy system (current load, equipment status such as the charging level of energy storage equipment, power generation, etc.), environmental data (weather in different time periods), and power market fluctuation data are constructed as the state space. It should be noted that the initial state space can also include other value flow parameters that affect the multi-value flow reward function value, such as the revenue of the wholesale power market, the load situation of transmission and distribution services, etc. It should be noted that the state of the energy system in this embodiment changes dynamically jointly determined by environmental data and the resource scheduling strategy. Due to the unknown dynamics, the transfer probability does not need to be explicitly modeled in the Markov decision modeling in this embodiment.

[0059] Furthermore, according to the initial state space, multiple candidate adjustment actions for the initial resource scheduling strategy can be determined. These candidate adjustment actions should be feasible in actual operation and can affect the state and performance of the energy system. For each candidate adjustment action, through predictive analysis, it is assumed that the resource scheduling strategy at the current time step is adjusted according to the candidate adjustment action, and the operating state of the energy system after applying the adjusted resource scheduling strategy at the current time step to the energy system is analyzed, and using the preset multi-value flow reward function, the multi-value reward function value of the energy system at this time is calculated, that is, the multi-value reward function value corresponding to the candidate adjustment action. Then, the target candidate adjustment action when the multi-value reward function value reaches the maximum is selected from multiple candidate adjustment actions, and the target candidate adjustment action is applied to the resource scheduling strategy at the current time step to update the resource scheduling strategy at the current time step.

[0060] In this embodiment, by comprehensively considering multiple important value dimensions of the energy system through the multi-value flow function, the energy system can achieve the optimal in multiple value dimensions, and the resource scheduling strategy of the energy system can quickly adapt to the changes in the operating state and environment of the energy system, improving the overall efficiency of the energy system.

[0061] In one embodiment, as Figure 4 shown, S310 includes:

[0062] S311, taking operation data, environmental data, and power market fluctuation data as inputs, calling the trained deep Q-network model, determining multiple candidate adjustment actions for the resource scheduling strategy at the current time step, and multiple multi-value reward function values corresponding to the multiple candidate adjustment actions.

[0063] Among them, the trained deep Q-network model includes a preset multi-value flow reward function, and the trained deep Q-network model is trained based on the historical operation data, historical environmental data, and historical power market fluctuation data of the energy system. The deep Q-network is a value-function-based reinforcement learning algorithm that can estimate the state s-action value a function (hereinafter referred to as the Q value), that is, the multi-value reward function value in this embodiment. Specifically, in the deep Q-network, represents the expected value of the long-term cumulative reward that can be obtained by taking action a (the candidate adjustment action in this embodiment) in a given state s (the operation data, environmental data, and power market fluctuation data of the energy system in this embodiment). In other words, measures the goodness or badness of executing a certain action in the current state of the energy system. During the policy adjustment process, the deep Q-network will tend to select larger candidate adjustment actions as the target adjustment actions.

[0064] It should be noted that the training process of the deep Q-network model can be as follows: First, construct an agent and an experience replay pool. The agent usually consists of an input layer, a hidden layer, and an output layer. Randomly initialize the model parameters in the agent. The experience replay pool is used to store training data. The training data can be represented by a quadruple , where is the state of the energy system at the current time step, is the adjustment action at the current time step, is the reward obtained after executing this adjustment action, is the state of the energy system at the next time step. Store the training data in the experience replay pool. During training, training data can be sampled from the experience replay pool for training the deep Q-network model. The deep Q-network model combines deep learning and Q learning. In the reinforcement learning framework, the agent interacts with the environment, changes the state of the energy system by executing actions, and obtains reward signals to learn the optimal behavior strategy. In the deep learning framework, the deep Q-network model can effectively process high-dimensional and complex state spaces by using the powerful function approximation ability of deep neural networks.

[0065] During the training process, the deep Q-network model can learn the rewards brought by executing candidate adjustment actions for the current resource scheduling strategy, that is, the multi-value reward function value in this embodiment. The multi-value reward function value can consider both immediate rewards and future rewards, and can also assign a discount factor to future rewards to measure the importance of immediate rewards and future rewards. By comparing the Q value calculated by the current deep Q-network model with the target Q value, where the target Q value is the learning objective of the deep Q-network model, the Q value output by the deep Q-network should be as close as possible to the target Q value. The target Q value is actually an approximation of the Bellman equation, which is an equation used in the field of reinforcement learning to describe the recursive relationship between state values and actions. Considering the complexity of the data, it is difficult to solve the Bellman equation precisely. Therefore, the target Q value can be used to approximate the Bellman equation. During the training process, by minimizing the error between the target Q value and the Q value calculated by the deep Q-network model, the gradient descent algorithm can be used to minimize this error, thereby continuously updating the model parameters of the deep Q-network model, enabling the deep Q-network model to better estimate the Q value.

[0066] During the training process, the deep Q-network model usually tends to select larger adjustment actions. However, in order to balance exploration and exploitation during the training process, a greedy strategy can be introduced into the deep Q-network model in this embodiment, that is, the server generates a random number. If the random number is less than the preset random number threshold, the deep Q-network model will randomly select an adjustment action, calculate and output the corresponding Q value for subsequent model parameter updates; if the random number is greater than or equal to the preset random number threshold (also known as the exploration rate, the larger the exploration rate, the more inclined to random exploration, and the smaller the exploration rate, the more inclined to utilize the knowledge already learned by the deep Q-network model), the deep Q-network model will output the Q value corresponding to the adjustment action with the largest Q value for subsequent model parameter updates. During training, the calculated by the deep Q-network model and the actual corresponding to this adjustment action can be used to calculate the mean square error, continuously optimizing the model parameters of the deep Q-network model until the preset accuracy requirement is met, obtaining the trained deep Q-network model. At this time, it can be understood that the preset multi-value flow reward function is "deployed" in the trained deep Q-network model, that is, the trained deep Q-network model can, based on the input data (operation data, environmental data, and electricity market fluctuation data), perform operations such as feature extraction on the above input data through a fully connected network in the hidden layer. When performing feature extraction, a convolutional neural network can also be combined to process multi-dimensional data, and finally output multiple candidate adjustment actions for the resource scheduling strategy at the current time step, as well as the multi-value reward function values corresponding to the multiple candidate adjustment actions.

[0067] It should be noted that the deep Q network model includes a main network and a target network. The model parameters of the main network and the target network are the same during initialization, and the model structures of the main network and the target network are also the same. During the training process, the model parameters of the main network are updated frequently, for example, the model parameters of the main network are updated after each round of training to continuously optimize the main network's estimate of the Q value, while the model parameters of the target network are updated less frequently, for example, the model parameters of the main network are copied to the target network only after a certain number of training rounds. This mechanism of regularly updating the model parameters of the target network can keep the target network relatively stable over a period of time, and is not easily affected by short-term drastic changes in the operating data of the energy system, environmental data, power market fluctuation data, etc., thereby providing a stable target for the training of the main network.

[0068] In this embodiment, the deep Q network model can effectively process high-dimensional and complex data such as operating data, environmental data and electricity market fluctuation data, and explore the potential relationships between the data, so as to determine reasonable candidate adjustment actions in a complex energy system environment, and comprehensively consider multiple important value dimensions of the energy system, so that the determined target adjustment action not only focuses on a single goal, but also seeks a balance between multiple goals, thereby maximizing the comprehensive benefits of the energy system and improving the reliability of resource scheduling in the energy system.

[0069] In one embodiment, the trained deep Q network model includes a first deep Q network model and a second deep Q network model, the first deep Q network model is used to determine multiple candidate adjustment actions for the resource scheduling strategy for the current time step, and to screen out a target candidate adjustment action when the multi-value reward function value reaches the maximum value from the multiple candidate adjustment actions, and the second deep Q network model is used to determine the multi-value reward function value corresponding to the target candidate adjustment action.

[0070] Specifically, the first deep Q network model focuses on determining multiple candidate adjustment actions for the resource scheduling strategy for the current time step, taking the operation data of the energy system, environmental data and electricity market fluctuation data as input. By analyzing and processing these complex data, the first deep Q network model uses its internal neural network structure and learned knowledge to output a series of candidate adjustment actions, as well as screen out the target candidate adjustment action when the multi-value reward function value reaches the maximum value from the multiple candidate adjustment actions.

[0071] The main task of the second deep Q-network model is to determine the multi-value reward function values corresponding to multiple candidate adjustment actions. After the first deep Q-network model gives the target candidate adjustment action, the second deep Q-network model will evaluate the target candidate adjustment action and calculate the multi-value reward function value corresponding to the target candidate adjustment action. Continuing from the above embodiment, the first deep Q-network model can be the main network, and the second deep Q-network model can be the target network. The main network is used to select the target adjustment action that can generate the maximum Q value, and the target network is used to calculate the Q value of the target adjustment action.

[0072] In this embodiment, by separating the selection of actions and the calculation of Q values, it is possible to avoid using the same network to perform these two tasks simultaneously. In this way, it can alleviate the problem that the Q value estimation is too high due to the error of the same Q network when selecting actions and calculating Q values, making the calculated Q value more accurate, improving the stability and convergence of the deep Q-network model, and also improving the accuracy of the resource scheduling strategy, thereby improving the reliability of the energy system resource scheduling.

[0073] In one embodiment, as Figure 5 shown, before S311, the method further includes:

[0074] S309, using historical operation data, historical environmental data, and historical power market fluctuation data as training data, performing priority sorting on the training data to obtain a priority sorting result.

[0075] S310, according to the priority sorting result, preferentially using the training data with a higher priority to train a preset initialized deep Q-network model to obtain a trained deep Q-network model.

[0076] Continuing from the above embodiment, the training data can be stored in an experience replay pool. Using historical operation data, historical environmental data, and historical power market fluctuation data as training data, in addition, the training data can also include the historical adjustment actions taken by the energy system at each time step in the past time period and the rewards brought by the historical adjustment actions (such as historical multi-value reward function values).

[0077] It should be noted that not all training data is equally important for model learning. Some training data is relatively important, such as the next state reached after taking different adjustment actions when the energy system is in different states, and the long-term rewards obtained. This type of training data can help the deep Q-network model learn the relationship between the energy system state, adjustment actions, and long-term rewards. Another example is the change situation of the historical environmental data and historical power market fluctuation data of the energy system, such as the change of load, market electricity price, equipment state, etc. This type of training data can help the deep Q-network model adapt to the dynamic changes of the environment and market and improve the accuracy of prediction. For example, the performance of the energy system when different resource scheduling strategies are adopted, such as the total revenue of the energy system, energy utilization efficiency, etc. These experiences can help the deep Q-network model learn to evaluate the advantages and disadvantages of different resource scheduling strategies.

[0078] By sorting the training data by priority and preferentially selecting the training data with high priority during the training process to train the preset initialized deep Q-network model, through continuous iterative training using the training data with high priority, the model parameters of the deep Q-network model are gradually optimized, and the estimation of the Q value by the deep Q-network model is also more and more accurate, obtaining the trained deep Q-network model.

[0079] In this embodiment, preferentially using the training data with a relatively high priority ranking to train the deep Q-network model can improve the efficiency and effect of model training, enabling the trained deep Q-network model to more accurately optimize the resource scheduling strategy, thereby improving the reliability of energy system resource scheduling.

[0080] In one embodiment, as Figure 6 shown, before S200, the method further includes:

[0081] S107: Taking the operation data, environmental data, and power market fluctuation data as inputs, calling the trained data prediction model to obtain the operation data prediction result and electricity price prediction result of the energy system.

[0082] S108: Determining the original resource scheduling strategy according to the operation data prediction result and the electricity price prediction result.

[0083] S109: Aiming at maximizing the preset performance index parameter, based on the preset ant colony optimization algorithm, iteratively optimizing the original resource scheduling strategy to obtain the initial resource scheduling strategy that makes the preset performance index parameter reach the maximum value.

[0084] Among them, the trained data prediction model is trained based on historical operation data, historical environmental data, and historical power market fluctuation data. It should be noted that historical operation data, historical environmental data, and historical power market fluctuation data are all preprocessed historical time series data. The preprocessing can include, but is not limited to, feature engineering, outlier handling, data augmentation, etc. Specifically, for the historical time series data originally collected by sensors, feature engineering processing is first performed to extract statistical features such as the mean, standard deviation, and change rate in the originally collected historical time series data, and external factors (such as calendar factors, including holiday flags, weekday flags, weekend flags, etc.) are introduced in combination with the actual business scenario. Then, outlier handling is carried out. For example, outliers are identified through the Z-score method and replaced by median interpolation. It is also possible to directly monitor and correct outliers through the Gaussian mixture model (GMM). Finally, data augmentation operations can be performed, including but not limited to time warping (randomly transforming historical time series data to simulate non-linear changes on the time axis), amplitude scaling (scaling the amplitude of the sequence within a reasonable range to enhance robustness), noise addition (introducing random noise to simulate measurement errors and improve the adaptability of the trained data prediction model to real data), etc. The operation data prediction result refers to the predicted value of the operation data of the energy system in a future period of time, such as the power generation and power consumption load at a certain future moment. The electricity price prediction result refers to the electricity price in the power market in a future period of time.

[0085] The ant colony optimization algorithm is an optimization algorithm that simulates the foraging behavior of ants. In this algorithm, each "ant" represents a possible resource scheduling scheme. The process of the "ant" searching for the optimal path in the search space is equivalent to searching for the optimal resource scheduling strategy among many possible resource scheduling strategies. The ant colony optimization algorithm guides the search direction of ants through the pheromone update mechanism. The path with a high pheromone concentration indicates that this path (i.e., the resource scheduling strategy) is better and will attract more ants to choose. Among them, the preset performance index parameters can include, but are not limited to, the economic benefit index of the energy system, the load stability index, the energy utilization rate index, etc.

[0086] Continuing from the above steps, it can be based on methods such as linear regression and decision trees in machine learning, or on recurrent neural networks and long short-term memory networks in deep learning to construct and train a data prediction model for predicting the operation data prediction result and the electricity price prediction result. It should be noted that this prediction process can be multi-time scale prediction. Both the operation data prediction result and the electricity price prediction result include short-term prediction results and medium- and long-term prediction results. The short-term prediction result can be a minute-level or hour-level prediction result for formulating real-time resource scheduling strategies, and the medium- and long-term prediction results can be daily-level or weekly-level predictions for formulating long-term resource scheduling strategies. Exemplarily, if it is predicted that the electricity price is high in a certain future period and the predicted power generation of a certain wind farm is also high, then the original resource scheduling strategy can be to arrange the wind farm to generate electricity at full load during this period and sell the excess electricity to the market; if it is predicted that the electricity price is low and the electricity load is small in a certain future period, then the original resource scheduling strategy can be to arrange the energy storage device to charge.

[0087] Furthermore, it can aim to maximize a preset performance metric parameter and iteratively optimize the original resource scheduling strategy through the ant colony optimization algorithm. In each iteration, the ant colony optimization algorithm generates a new set of candidate resource scheduling strategies, calculates the performance metric parameter values corresponding to each resource scheduling strategy, and then updates the pheromone based on the performance metric parameter values corresponding to each resource scheduling strategy to guide the search direction of the next iteration. After multiple iterations of optimization, the ant colony optimization algorithm gradually converges to the strategy that maximizes the preset performance metric parameter, and this strategy is the initial resource scheduling strategy.

[0088] In this embodiment, by combining the data prediction model, the ant colony optimization algorithm, etc., first determine the original resource scheduling strategy based on the operation data prediction result and the electricity price prediction result, and further continuously explore and adjust the original resource scheduling strategy through the ant colony optimization algorithm, which can make the initial resource scheduling strategy maximize the preset performance metric parameter as much as possible, improving the reliability and accuracy of resource scheduling.

[0089] In one embodiment, the trained data prediction model includes a long short-term memory network module, a Transformer module, and an output layer. S107 includes: using operation data, environmental data, and power market fluctuation data as input data, calling the long short-term memory network module to obtain the hidden state sequence of the input data, using the hidden state sequence as input, calling the Transformer module to obtain the output result of the Transformer module, concatenating the hidden state sequence and the output result of the Transformer module to obtain a concatenated vector, and using the concatenated vector as input, calling the output layer to obtain the operation data prediction result and the electricity price prediction result of the energy system.

[0090] Among them, the long short-term memory network module is responsible for extracting long-term and short-term dependencies from operation data, environmental data, and power market fluctuation data. The long short-term memory network is a special type of recurrent neural network that is good at processing sequential data and can capture long-term dependencies in the input data. In the scenario of energy system data prediction, operation data, environmental data, and power market fluctuation data all have time series characteristics. For example, the power generation power at different times, the environmental temperature at different periods, and the electricity price at different time points, etc. The long short-term memory network module can remember the information of these time series data within a relatively long time range, so as to make more accurate predictions of future data. Specifically, it controls the inflow, retention, and output of information through a gating mechanism (input gate, forget gate, and output gate), effectively solving the problem of gradient vanishing or gradient explosion in traditional recurrent neural networks. The Transformer model introduces a self-attention mechanism, which can process the input time series data in parallel and better capture the global dependencies of the time series data. When dealing with the complex data of the energy system, the Transformer module can pay attention to the mutual influence between different types of data. For example, how the wind speed and light intensity in environmental data affect the operation data of power generation equipment at the same time, and how these changes are related to the power market fluctuation data, etc. The Transformer model can establish direct connections between elements at different positions without being restricted by distance, so as to analyze the data more comprehensively. The output layer receives the results processed by the previous module and converts them into the final prediction results. In the prediction of energy system data, the output layer can be a fully connected layer, whose function is to output the prediction results of the operation data and electricity price prediction results of the energy system according to the concatenated vector.

[0091] Exemplarily, the long short-term memory network module in this embodiment includes a forgetting gate (for discarding unnecessary historical information), an input gate (for determining which new information to store), a memory cell update component (for combining the forgotten information and new information to generate a new memory cell), and an output gate (for generating the hidden state at the current time step). The operation data, environmental data, and power market fluctuation data are input into the long short-term memory network module, and the output data is a sequence of hidden states, which characterizes the dynamic characteristics in the input time series data, and this sequence of hidden states will be used as the input of the Transformer module. The Transformer module further processes the sequence of hidden states to capture global dependencies and complex patterns in the time series. Specifically, the Transformer module adjusts the dimension of the sequence of hidden states through a linear transformation and adds a positional encoding to it to identify the time step order. The Transformer module includes an encoder and a decoder. The encoder extracts the global features of the processed sequence of hidden states to capture the long-term dependencies between time series, and the decoder generates the predicted values for future time steps based on the long-term dependencies to obtain the output result of the Transformer module. Moreover, the Transformer module also deploys a multi-head attention mechanism to allow the Transformer module to capture global dependency information at different levels and from different perspectives, enabling it to better handle complex patterns in time series data. It should be noted that in addition to the sequence of hidden states, the input data of the Transformer module can also have additional inputs. One only needs to concatenate the sequence of hidden states and the additional inputs as the input data of the Transformer module. Further, by performing a weighted average or concatenation operation on the sequence of hidden states output by the long short-term memory network module and the output result of the Transformer module, a concatenated vector can be obtained, and then the predicted operation data result and electricity price prediction result of the energy system can be obtained through the output layer.

[0092] It should be noted that during the training process of the data prediction model, an attention mechanism can be introduced in the output of the long short-term memory network module and the Transformer module to dynamically assign weights to different features, so that the data prediction model can pay more attention to key time steps and key features. The mean squared error can be used as the main loss function during the training process, and at the same time, a regularization term (to prevent the model from overfitting) and context constraints (such as temperature, humidity, etc. will affect power consumption, holidays, weekdays, weekends will affect the power consumption pattern, electricity price fluctuations and market demand changes will directly affect power generation and consumption behaviors, and major events such as sports events and natural disasters will also have an impact on power consumption) are introduced.

[0093] In this embodiment, by combining the advantages of the long short-term memory network module and the Transformer module, it is possible to capture the long-term time dependence relationship of the data and mine the global dependence relationship of the data, so that the data prediction model can more comprehensively and accurately understand the complex characteristics of the energy system, thereby improving the prediction accuracy and further enhancing the reliability of the subsequent energy system resource scheduling.

[0094] In one embodiment, as Figure 7 shown, the method further includes:

[0095] S610, evaluate the prediction performance of the trained data prediction model according to the prediction results of the operation data of the energy system and the prediction results of the electricity price, and obtain an evaluation result.

[0096] S620, take the prediction results of the operation data of the energy system and the prediction results of the electricity price as inputs, call a preset graph drawing tool, and determine the operation data prediction curve and the electricity price prediction curve.

[0097] S630, visualize the evaluation result, the operation data prediction curve and the electricity price prediction curve.

[0098] S640, receive a model adjustment instruction sent by the user, the model adjustment instruction carries model adjustment parameters, and adjust the model parameters of the trained data prediction model according to the model adjustment parameters.

[0099] Among them, the preset graph drawing tool includes but is not limited to common visualization software, which can convert data into intuitive graphs. The operation data prediction curve can have time as the horizontal axis and operation data (such as power generation power, power consumption load, etc.) as the vertical axis, reflecting the change trend of the operation data over time. The electricity price prediction curve can have time as the horizontal axis and the electricity price as the vertical axis, reflecting the change trend of the electricity price over time.

[0100] Specifically, the prediction performance of the trained data prediction model can be evaluated by using the prediction results of the operation data of the energy system and the prediction results of the electricity price. For example, the prediction results of the operation data and the electricity price are respectively compared with the corresponding actual values, and the evaluation indicators can include but are not limited to mean square error, root mean square error, mean absolute error, etc., to obtain an evaluation result. The evaluation result can intuitively reflect the accuracy and reliability of the trained data prediction model in predicting the operation data and electricity price of the energy system. In addition, corresponding confidence intervals can be preset for the prediction results of the operation data and the electricity price respectively, and the uncertainty of the prediction results of the operation data and the electricity price can be evaluated through the confidence intervals.

[0101] Further, by using a preset graph drawing tool, generate an operating data prediction curve and an electricity price prediction curve, and visually display the evaluation results, the operating data prediction curve, and the electricity price prediction curve. For example, present the evaluation metrics in the form of numerical values or charts, and at the same time draw the operating data prediction curve and the electricity price prediction curve on the same interface, enabling users to comprehensively understand the prediction performance and prediction results of the data prediction model. In addition, based on the visually displayed results and combined with their own needs and experience, users can send model adjustment instructions to the server through an interactive dashboard. The model adjustment instructions carry model adjustment parameters, including but not limited to the learning rate, the number of iterations, modifications to the model network structure, etc. Then, the server can adjust the model parameters of the trained data prediction model according to the model adjustment parameters.

[0102] It should be noted that the above operations on the trained data prediction model can also be applied during the training process of the data prediction model. Users can adjust the model parameters or training parameters of the data prediction model according to the training results of the data prediction model, so that the data prediction model can better meet the application requirements.

[0103] In this embodiment, by evaluating the model prediction performance based on the operating data prediction results and the electricity price prediction results, it is possible to accurately analyze the deviations of the data prediction model in predicting the operation of the energy system and the electricity price, and allow users to continuously evaluate and adjust the model parameters, improving the prediction accuracy of the data prediction model, enabling it to better simulate and predict the operating data and electricity price of the energy system, and providing data support for formulating reliable resource scheduling strategies in the future.

[0104] To provide a clearer description of the resource scheduling method for the energy system provided in this application, the following is explained in combination with a detailed embodiment and an appendix Figure 8 The detailed embodiment includes the following steps:

[0105] S801, Obtain the operating data, environmental data, and power market fluctuation data of the energy system. Using the operating data, environmental data, and power market fluctuation data as input data, call the long short-term memory network module to obtain the hidden state sequence of the input data.

[0106] S802, Using the hidden state sequence as input, call the Transformer module to obtain the output result of the Transformer module. Concatenate the hidden state sequence and the output result of the Transformer module to obtain a concatenated vector. Using the concatenated vector as input, call the output layer to obtain the operating data prediction result and the electricity price prediction result of the energy system.

[0107] S803. Determine the original resource scheduling strategy based on the operation data prediction result and the electricity price prediction result. With the goal of maximizing the preset performance index parameters, based on the preset ant colony optimization algorithm, iteratively optimize the original resource scheduling strategy to obtain the initial resource scheduling strategy that maximizes the preset performance index parameters, and use the initial resource scheduling strategy as the resource scheduling strategy for the current time step.

[0108] S804. Using the operation data, environmental data, and power market fluctuation data as inputs, call the trained deep Q-network model to determine multiple candidate adjustment actions for the resource scheduling strategy at the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions.

[0109] S805. Screen out the target candidate adjustment action when the multi-value reward function value reaches the maximum from the multiple candidate adjustment actions, and update the resource scheduling strategy for the current time step based on the target candidate adjustment action.

[0110] S806. Based on the resource scheduling strategy for the current time step, schedule the resources of the energy system, and update the operation data, environmental data, and power market fluctuation data of the energy system.

[0111] S807. Return to step S804 until the preset iteration end condition is reached.

[0112] It should be noted that the schematic diagram of the structure of the deep Q-network model mentioned in this embodiment is as Figure 8 shown, Figure 8 in represents the environmental state in which the agent is located (which can be a numerical vector, an image, or other forms of data), represents the action output by the agent (i.e., the adjustment action), which can be determined by the deep neural network deployed in the agent according to the at the current time step. Figure 8 The environment in represents the external environment with which the agent interacts. In this embodiment, it can be an energy system. When the agent takes the action , the environment can generate a new state and feedback a reward

[0113] In addition, this embodiment can be applied to an energy management system. The system architecture of the energy management system includes a data acquisition layer, an intelligent computing layer, and an application layer. The data acquisition layer can collect the operation data of the energy system through IOT (Internet of Things devices) and edge computing nodes, and transmit the collected operation data to the server through the MQTT protocol, and clean and store it in real time. The IOT includes, but is not limited to, battery monitoring devices (real-time monitoring of battery status, such as voltage, current, temperature, charge and discharge status, etc.), photovoltaic power generation monitoring devices (which can be sensors to monitor the output power, light intensity, etc. of the photovoltaic system), wind power generation monitoring devices (monitoring wind speed, wind direction, and the status of wind turbines), and smart meters (used to collect data such as power consumption and load). The edge computing nodes include, but are not limited to, battery management system nodes (monitoring the health status and charge and discharge process of the battery, and can perform simple local calculations to optimize the battery usage efficiency), photovoltaic inverter nodes (real-time processing of the power data of the photovoltaic panel), and smart meter nodes (processing real-time power consumption data, performing preliminary data filtering and data compression and then uploading to the server). The edge computing nodes do not necessarily have to have a computing function, and can also only perform the function of obtaining data and uploading it to the server.

[0114] The intelligent computing layer includes an optimization engine, a prediction model, and a coordinated optimization module. The optimization engine is used to formulate an initial resource scheduling strategy according to the ant colony optimization algorithm. The prediction module is used to predict the operation data prediction result and electricity price prediction result based on the deep learning principle through the data prediction model. The coordinated optimization module is used to adopt a reinforcement learning algorithm (such as the deep Q network model) to dynamically optimize the resource scheduling strategy of the energy system under the superposition of multiple value streams (such as the wholesale market, transmission and distribution services, etc.). The application layer is used to provide a user-friendly interaction platform. Users can view the real-time status of the resources of the energy system, the resource scheduling strategy, the operation data prediction result, the electricity price prediction result, etc. through the interaction platform in real time. The interaction platform also supports visual energy asset portfolio management and analysis.

[0115] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0116] Based on the same inventive concept, an embodiment of the present application further provides a resource scheduling device for an energy system for implementing the resource scheduling method of the energy system involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the resource scheduling device for the energy system provided below can refer to the limitations on the resource scheduling method of the energy system in the above text, and will not be repeated here.

[0117] In one embodiment, as Figure 10 shown, a resource scheduling device 900 for an energy system is provided, including: a data acquisition module 910, a policy formulation module 920, a policy adjustment module 930, and a resource scheduling module 940, where:

[0118] The data acquisition module 910 is configured to acquire the operation data, environmental data, and power market fluctuation data of the energy system.

[0119] The policy formulation module 920 is configured to use a preset initial resource scheduling policy as the resource scheduling policy for the current time step.

[0120] The policy adjustment module 930 is configured to update the resource scheduling policy for the current time step to the resource scheduling policy that maximizes the preset multi-value flow reward function according to the operation data, environmental data, and power market fluctuation data of the energy system.

[0121] The resource scheduling module 940 is configured to schedule the resources of the energy system based on the resource scheduling policy for the current time step.

[0122] In one embodiment, the resource scheduling device 900 is further configured to update the operation data, environmental data, and power market fluctuation data of the energy system. The control strategy adjustment module 930 executes again the step of updating the resource scheduling strategy of the current time step to the resource scheduling strategy that maximizes the preset multi-value stream reward function according to the operation data, environmental data, and power market fluctuation data of the energy system until the preset iteration end condition is reached.

[0123] In one embodiment, the strategy adjustment module 930 is further configured to determine multiple candidate adjustment actions for the resource scheduling strategy of the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions according to the operation data, environmental data, power market fluctuation data, and the preset multi-value stream reward function of the energy system, and screen out the target candidate adjustment action when the multi-value reward function value reaches the maximum from the multiple candidate adjustment actions, and update the resource scheduling strategy of the current time step based on the target candidate adjustment action.

[0124] In one embodiment, the strategy adjustment module 930 is further configured to use the operation data, environmental data, and power market fluctuation data as inputs to call the trained deep Q-network model to determine multiple candidate adjustment actions for the resource scheduling strategy of the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions, wherein the trained deep Q-network model includes the preset multi-value stream reward function, and the trained deep Q-network model is trained based on the historical operation data, historical environmental data, and historical power market fluctuation data of the energy system.

[0125] In one embodiment, the trained deep Q-network model includes a first deep Q-network model and a second deep Q-network model. The first deep Q-network model is configured to determine multiple candidate adjustment actions for the resource scheduling strategy of the current time step, and screen out the target candidate adjustment action when the multi-value reward function value reaches the maximum from the multiple candidate adjustment actions. The second deep Q-network model is configured to determine the multi-value reward function value corresponding to the target candidate adjustment action.

[0126] In one embodiment, the resource scheduling device 900 is further configured to use the historical operation data, historical environmental data, and historical power market fluctuation data as training data, perform priority sorting on the training data to obtain a priority sorting result, and preferentially use the training data with a high priority sorting result to train the preset initial deep Q-network model to obtain the trained deep Q-network model.

[0127] In one embodiment, the resource scheduling device 900 is further configured to use the operation data, environment data, and power market fluctuation data as inputs, call the trained data prediction model, obtain the operation data prediction result and electricity price prediction result of the energy system, determine the original resource scheduling strategy according to the operation data prediction result and the electricity price prediction result, and take maximizing the preset performance index parameter as the goal. Based on the preset ant colony optimization algorithm, the original resource scheduling strategy is iteratively optimized to obtain the initial resource scheduling strategy that maximizes the preset performance index parameter. Among them, the trained data prediction model is trained based on historical operation data, historical environment data, and historical power market fluctuation data.

[0128] In one embodiment, the trained data prediction model includes a long short-term memory network module, a Transformer module, and an output layer. The resource scheduling device 900 is further configured to use the operation data, environment data, and power market fluctuation data as input data, call the long short-term memory network module to obtain the hidden state sequence of the input data, use the hidden state sequence as the input, call the Transformer module to obtain the output result of the Transformer module, splice the hidden state sequence and the output result of the Transformer module to obtain a spliced vector, and use the spliced vector as the input to call the output layer to obtain the operation data prediction result and electricity price prediction result of the energy system.

[0129] In one embodiment, the resource scheduling device 900 is further configured to evaluate the prediction performance of the trained data prediction model according to the operation data prediction result and the electricity price prediction result of the energy system to obtain an evaluation result. Using the operation data prediction result and the electricity price prediction result of the energy system as inputs, call the preset graph drawing tool to determine the operation data prediction curve and the electricity price prediction curve, visualize the evaluation result, the operation data prediction curve, and the electricity price prediction curve, receive the model adjustment instruction sent by the user, where the model adjustment instruction carries the model adjustment parameter, and adjust the model parameter of the trained data prediction model according to the model adjustment parameter.

[0130] Each module in the above resource scheduling device of the energy system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software to facilitate the processor to call and execute the operations corresponding to the above modules.

[0131] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as operation data of the energy system, environmental data, and power market fluctuation data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a resource scheduling method for an energy system.

[0132] Those skilled in the art can understand that Figure 11 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0133] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above-mentioned embodiment of the resource scheduling method for the energy system.

[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above-mentioned embodiment of the resource scheduling method for the energy system.

[0135] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned embodiment of the resource scheduling method for the energy system.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards in the relevant regions.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0139] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A resource scheduling method for an energy system, characterized in that, The method includes: Obtaining the operation data, environmental data, and power market fluctuation data of the energy system; Using the preset initial resource scheduling strategy as the resource scheduling strategy for the current time step; According to the operation data, environmental data, and power market fluctuation data of the energy system, with the goal of maximizing the preset multi-value flow reward function, updating the resource scheduling strategy for the current time step to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value; Based on the resource scheduling strategy for the current time step, scheduling the resources of the energy system.

2. The method according to claim 1, wherein After scheduling the resources of the energy system based on the resource scheduling strategy for the current time step, the method further includes: Updating the operation data, environmental data, and power market fluctuation data of the energy system; Returning to the step of updating the resource scheduling strategy for the current time step to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value according to the operation data, environmental data, and power market fluctuation data of the energy system, with the goal of maximizing the preset multi-value flow reward function, until the preset iteration end condition is reached.

3. The method according to claim 1, wherein The step of updating the resource scheduling strategy for the current time step to the resource scheduling strategy that makes the preset multi-value flow reward function reach the maximum value according to the operation data, environmental data, and power market fluctuation data of the energy system includes: Determining multiple candidate adjustment actions for the resource scheduling strategy for the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions, according to the operation data, environmental data, power market fluctuation data, and preset multi-value flow reward function of the energy system; Selecting the target candidate adjustment action when the multi-value reward function value reaches the maximum value from the multiple candidate adjustment actions, and updating the resource scheduling strategy for the current time step based on the target candidate adjustment action.

4. The method according to claim 3, wherein The step of determining multiple candidate adjustment actions for the resource scheduling strategy for the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions, according to the operation data, environmental data, power market fluctuation data, and preset multi-value flow reward function of the energy system includes: Using the operation data, environmental data, and power market fluctuation data as inputs, calling the trained deep Q-network model to determine multiple candidate adjustment actions for the resource scheduling strategy for the current time step, and the multi-value reward function values corresponding to the multiple candidate adjustment actions; Wherein, the trained deep Q-network model includes a preset multi-value flow reward function, and the trained deep Q-network model is trained based on the historical operation data, historical environmental data, and historical power market fluctuation data of the energy system.

5. The method according to claim 4, wherein The trained deep Q-network model includes a first deep Q-network model and a second deep Q-network model. The first deep Q-network model is used to determine multiple candidate adjustment actions for the resource scheduling strategy at the current time step, and screen out the target candidate adjustment action when the multi-value reward function value reaches the maximum from the multiple candidate adjustment actions. The second deep Q-network model is used to determine the multi-value reward function value corresponding to the target candidate adjustment action.

6. The method according to claim 4, characterized in that, Before calling the trained deep Q-network model with the operation data, the environment data, and the power market fluctuation data as inputs, it further includes: Using the historical operation data, historical environment data, and historical power market fluctuation data as training data, performing priority sorting on the training data to obtain a priority sorting result; According to the priority sorting result, preferentially using the training data with a higher priority sorting to train a preset initial deep Q-network model to obtain a trained deep Q-network model.

7. The method according to any one of claims 1 to 6, characterized in that, Before using the preset initial resource scheduling strategy as the resource scheduling strategy at the current time step, the method further includes: Using the operation data, the environment data, and the power market fluctuation data as inputs, calling a trained data prediction model to obtain the operation data prediction result and the electricity price prediction result of the energy system; Determining an original resource scheduling strategy according to the operation data prediction result and the electricity price prediction result; Taking maximizing a preset performance index parameter as the goal, based on a preset ant colony optimization algorithm, iteratively optimizing the original resource scheduling strategy to obtain an initial resource scheduling strategy that makes the preset performance index parameter reach the maximum; Among them, the trained data prediction model is trained based on historical operation data, historical environment data, and historical power market fluctuation data.

8. The method according to claim 7, wherein The trained data prediction model includes a long short-term memory network module, a Transformer module, and an output layer; Using the operation data, the environment data, and the power market fluctuation data as inputs, calling a trained data prediction model to obtain the operation data prediction result and the electricity price prediction result of the energy system, including: Using the operation data, the environment data, and the power market fluctuation data as input data, calling the long short-term memory network module to obtain a hidden state sequence of the input data; Using the hidden state sequence as an input, calling the Transformer module to obtain the output result of the Transformer module; Concatenating the hidden state sequence and the output result of the Transformer module to obtain a concatenated vector; Using the concatenated vector as an input, calling the output layer to obtain the operation data prediction result and the electricity price prediction result of the energy system.

9. The method according to claim 7, characterized in that The method further includes: Evaluating the prediction performance of the trained data prediction model according to the operation data prediction result and the electricity price prediction result of the energy system to obtain an evaluation result; Taking the operation data prediction result and electricity price prediction result of the energy system as inputs, call a preset graph drawing tool to determine the operation data prediction curve and the electricity price prediction curve; Visualize the evaluation result, the operation data prediction curve, and the electricity price prediction curve; Receive a model adjustment instruction sent by a user, where the model adjustment instruction carries model adjustment parameters; According to the model adjustment parameters, adjust the model parameters of the trained data prediction model.

10. A resource scheduling device for an energy system, characterized in that, The device includes: A data acquisition module for acquiring the operation data, environmental data, and power market fluctuation data of the energy system; A strategy formulation module for using a preset initial resource scheduling strategy as the resource scheduling strategy for the current time step; A strategy adjustment module for, based on the operation data, environmental data, and power market fluctuation data of the energy system, aiming to maximize a preset multi-value flow reward function, update the resource scheduling strategy for the current time step to a resource scheduling strategy that maximizes the preset multi-value flow reward function; A resource scheduling module for scheduling the resources of the energy system based on the resource scheduling strategy for the current time step.

Citation Information

Patent Citations

  • Resource scheduling method and device of virtual power plant, computer equipment and storage medium

    CN117933667A

  • Park-level adjustable load-oriented resource regulation and control method, device and equipment and medium

    CN119539424A

  • Power dispatching optimization method and system based on reinforcement learning

    CN119647898A

  • Virtual power plant peak regulation optimization scheduling method and system, electronic equipment and medium

    CN119783997A

  • Dynamic spectrum sharing based on machine learning

    US20230217264A1