Scheduling strategy generation method and device, equipment, storage medium and product

By employing a pre-defined two-layer game model and particle swarm optimization technology in virtual power plant scheduling, the optimal scheduling strategy is generated, solving the problem of low scheduling efficiency caused by a single entity or single objective, and achieving a balance between the revenue of virtual power plant operators and virtual power plants, as well as improving scheduling accuracy.

CN121787770APending Publication Date: 2026-04-03CHINA RESOURCES NEW ENERGY INVESTMENT CO LTD SHANXI BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing research on virtual power plant dispatching focuses on a single entity or a single objective, resulting in low overall dispatching efficiency.

Method used

An initial scheduling strategy is generated by adopting a pre-defined two-layer game model, combined with particle swarm optimization and non-uniform mutation operator mechanism. The optimal scheduling strategy is obtained through iterative optimization, taking into account the benefits of both the virtual power plant operator and the virtual power plant.

Benefits of technology

It improves overall scheduling efficiency. By fully considering the benefits of both parties through a two-level game model, it achieves a balance of interests between virtual power plant operators and virtual power plants, thereby improving the accuracy and convergence speed of the scheduling scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787770A_ABST
    Figure CN121787770A_ABST
Patent Text Reader

Abstract

The invention discloses a scheduling strategy generation method and device, equipment, a storage medium and a product, and relates to the technical field of virtual power plants, and the scheduling strategy generation method comprises the steps: randomly generating an initial scheduling strategy of a preset scheduling time period, inputting the initial scheduling strategy to a preset double-layer game model, and obtaining a fitness score, the upper layer of the preset double-layer game model is a virtual power plant operator income optimization model, and the lower layer of the preset double-layer game model is a virtual power plant income optimization model; and performing iterative optimization on the initial scheduling strategy based on the fitness score to obtain an optimal scheduling strategy meeting an iteration termination condition. The upper layer of the preset double-layer game model is the virtual power plant operator income optimization model, and the lower layer of the preset double-layer game model is the virtual power plant income optimization model, so that the comfort score output by the model fully considers the income of both the virtual power plant operator and the virtual power plant, and the overall scheduling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual power plant technology, and in particular to a scheduling strategy generation method, apparatus, equipment, storage medium and product. Background Technology

[0002] Virtual power plants, as the core carriers that aggregate distributed energy, energy storage, and controllable loads, can improve the flexibility of the power system through coordinated dispatch, and have become one of the key technologies for the construction of new power systems.

[0003] Current research on the economic dispatch of virtual power plants often focuses on a single entity (such as optimizing only the operating costs of virtual power plants) or a single objective (such as considering only economic efficiency), resulting in low overall dispatch efficiency. Summary of the Invention

[0004] The main purpose of this application is to provide a scheduling strategy generation method, apparatus, device, storage medium and product, which aims to solve the technical problem of low overall scheduling efficiency caused by a single subject or single target.

[0005] To achieve the above objectives, this application proposes a scheduling policy generation method, which includes: An initial scheduling strategy for a preset scheduling period is randomly generated. The initial scheduling strategy is then input into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. Based on the fitness score, the initial scheduling strategy is iteratively optimized to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0006] In one embodiment, the step of randomly generating the initial scheduling strategy for a preset scheduling period includes: A particle swarm is randomly generated, and a scheduling decision vector is set for each particle in the particle swarm. The scheduling decision vector includes the particle position and update speed. The particle position includes the purchase and sale price of electricity by the virtual power plant operator and the purchase and sale power of the virtual power plant. The scheduling parameters at the particle positions are used as the initial scheduling strategy.

[0007] In one embodiment, the step of iteratively optimizing the initial scheduling strategy based on the fitness score to obtain an optimal scheduling strategy that satisfies the iteration termination condition includes: Based on the fitness score, determine the current optimal scheduling strategy; The adjustment factor for each particle is calculated, and the particle update mechanism is determined based on the adjustment factor. The particle update mechanism includes a particle swarm optimization mechanism and a non-uniform mutation operator mechanism. Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is iteratively updated to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0008] In one embodiment, the step of iteratively updating the scheduling decision vector of each particle based on the particle update mechanism and the current globally optimal scheduling strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition includes: Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is updated to obtain the updated scheduling decision vector. Determine whether the updated scheduling decision vector satisfies the iteration termination condition; If not satisfied, determine the fitness score of the updated scheduling decision vector and return the step of determining the current optimal scheduling strategy based on the fitness score, until the change of the updated target scheduling decision vector from the scheduling decision vector obtained after the previous iteration is less than the threshold. The current optimal scheduling strategy is taken as the optimal scheduling strategy that satisfies the iteration termination condition.

[0009] In one embodiment, the step of determining the current optimal scheduling strategy based on the fitness score includes: Based on the fitness score, the current individual optimal strategy for each particle is determined; Determine the optimal strategy for the target individual with the highest fitness score among the individual optimal strategies; Determine whether the fitness score of the optimal strategy for the target individual is greater than the fitness score of the optimal scheduling strategy; If the value is greater than the target value, the optimal scheduling policy will be updated to the target optimal policy.

[0010] In one embodiment, the step of updating the scheduling decision vector of each particle based on the particle update mechanism and the optimal scheduling strategy to obtain the updated scheduling decision vector includes: When the particle update mechanism is a non-uniform mutation operator mechanism, perturbation random numbers are randomly generated. Based on the perturbation random number, the migration direction and migration offset are determined; The scheduling decision vector is shifted along the migration direction by the migration offset to obtain the updated scheduling decision vector.

[0011] Furthermore, to achieve the above objectives, this application also proposes a scheduling strategy generation apparatus, which includes: The generation module is used to randomly generate an initial scheduling strategy for a preset scheduling period, and input the initial scheduling strategy into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. An optimization module is used to iteratively optimize the initial scheduling strategy based on the fitness score to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0012] In addition, to achieve the above objectives, this application also proposes a scheduling policy generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the scheduling policy generation method as described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the scheduling policy generation method described above.

[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the scheduling policy generation method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: Compared to related technologies that often focus on a single entity (e.g., optimizing only the operating costs of virtual power plants) or a single objective (e.g., considering only economics), leading to low overall scheduling efficiency, this application randomly generates an initial scheduling strategy for a preset scheduling period. This initial scheduling strategy is then input into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. Based on the fitness score, the initial scheduling strategy is iteratively optimized to obtain an optimal scheduling strategy that meets the iteration termination condition. This application uses a preset two-layer game model to determine the comfort score of the initial scheduling strategy for a preset scheduling period. Based on the comfort score output by the preset two-layer game model, the initial scheduling strategy is iteratively optimized to obtain an optimal scheduling strategy that meets the iteration termination condition. Since the upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model, the comfort score output by this model fully considers the revenue of both the virtual power plant operator and the virtual power plant, thus improving overall scheduling efficiency. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the scheduling strategy generation method of this application in Embodiment 1. Figure 2 This is a diagram illustrating the energy-sharing system architecture of the virtual power plant and the virtual power plant operator in the scheduling strategy generation method of this application. Figure 3 This is a flowchart illustrating Embodiment 2 of the scheduling strategy generation method of this application; Figure 4 This is a schematic diagram of the module structure of the scheduling strategy generation device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the scheduling strategy generation method in this application embodiment.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solution of this application embodiment is: randomly generating an initial scheduling strategy for a preset scheduling period, inputting the initial scheduling strategy into a preset two-layer game model to obtain a fitness score, wherein the upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model; based on the fitness score, iteratively optimizing the initial scheduling strategy to obtain an optimal scheduling strategy that satisfies the iteration termination condition.

[0023] In contrast to related technologies that often focus on a single entity (e.g., optimizing only the operating costs of virtual power plants) or a single objective (e.g., considering only economic efficiency), resulting in low overall scheduling efficiency, this application randomly generates an initial scheduling strategy for a preset scheduling period. This initial scheduling strategy is then input into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. Based on the fitness score, the initial scheduling strategy is iteratively optimized to obtain an optimal scheduling strategy that meets the iteration termination condition.

[0024] This application uses a pre-defined two-layer game model to determine the comfort score of the initial scheduling strategy for a pre-defined scheduling period. Based on the comfort score output by the pre-defined two-layer game model, the initial scheduling strategy is iteratively optimized to obtain the optimal scheduling strategy that meets the iteration termination condition. Since the upper layer of the pre-defined two-layer game model is a virtual power plant operator revenue optimization model and the lower layer is a virtual power plant revenue optimization model, the comfort score output by the model fully considers the revenue of both the virtual power plant operator and the virtual power plant, thus improving the overall scheduling efficiency.

[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or scheduling strategy generation device capable of performing the above functions. The following description uses a scheduling strategy generation device as an example to illustrate this embodiment and the subsequent embodiments.

[0026] Based on this, embodiments of this application provide a scheduling policy generation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the scheduling strategy generation method of this application.

[0027] In this embodiment, the scheduling policy generation method includes steps S10 to S20: Step S10: Randomly generate an initial scheduling strategy for a preset scheduling period, input the initial scheduling strategy into a preset two-layer game model, and obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. It should be noted that the execution entity in this embodiment is the scheduling strategy generation device. The initial scheduling strategy includes an initial combination of power purchase and sale and electricity price of virtual power plants and virtual power plant operators set before the algorithm iteration begins, serving as the starting point for optimization. The preset scheduling period refers to a predefined time period in economic scheduling, such as 24 hours, corresponding to the total number of scheduling periods NT, used to divide the optimization interval. The fitness score is an evaluation index obtained by calculating the objective function value of the two-level game model, used to measure the quality of the scheduling strategy. Within the preset scheduling period (e.g., 24 hours), the scheduling strategy generation device randomly generates a set of initial scheduling strategies, inputs the initial scheduling strategies into the preset two-level game model, and finally outputs the fitness score.

[0028] Specifically, the process of constructing the pre-defined two-layer game model is as follows: 1. Model Body and Boundary Definition (1) Participating entities: The lower layer consists of multiple virtual power plants, each containing wind power, photovoltaic, and energy storage batteries; the upper layer consists of system operators, who are responsible for formulating the transaction price between virtual power plants and the power purchase and sale strategy of the distribution network.

[0029] (2) Transaction Boundaries: Virtual power plants can trade electricity through two channels—horizontal transactions with other virtual power plants and vertical transactions with virtual power plant operators; virtual power plant operators can purchase electricity from the distribution network (to supplement the power supply gap of virtual power plants) or sell electricity (surplus electricity of virtual power plants), referring to Figure 2 , Figure 2 It provides an energy-sharing system architecture diagram for virtual power plants and virtual power plant operators.

[0030] 2. Lower-level virtual power plant revenue optimization model (1) Objective function: Minimize the total cost of the virtual power plant (operating cost + environmental pollution cost) - Revenue from electricity purchase and sale), the formula is as follows:

[0031] in, The total cost of the virtual power plant. The operating cost of a virtual power plant, including wind power and solar power maintenance costs, curtailment costs, and battery operating costs, is calculated using the following formula:

[0032] in, , This is the operating cost coefficient. To contribute to actual wind power, To contribute to the actual development of photovoltaics The charging and discharging power of the energy storage battery. For wind curtailment power, For abandoned light power, , This is the penalty cost coefficient.

[0033] The environmental pollution cost of a virtual power plant, including electromagnetic pollution from transmission lines and chemical leakage costs from batteries, is calculated using the following formula:

[0034] in, This refers to the electromagnetic pollution cost coefficient per unit transmission power of a power transmission line. The chemical pollution cost coefficient per unit charge / discharge power of energy storage batteries. The charging and discharging power of the energy storage battery.

[0035] , The power purchased and sold by a virtual power plant to a virtual power plant operator; , The purchase and sale prices of electricity set by virtual power plant operators; NT: Total number of scheduling periods (e.g., 24 hours).

[0036] (2) Constraints Load constraint: The total energy supply of the virtual power plant must meet the load demand, as shown in the formula:

[0037] Energy storage constraints: The battery's state of charge (SOC) must be within a safe range, as shown in the formula: Power exchange constraints: The power purchased and sold between virtual power plants and virtual power plant operators must be within limits, as shown in the formula:

[0038] 3. Revenue Optimization Model for Upper-Level Virtual Power Plant Operators (1) Objective function: Maximize the total revenue of the virtual power plant operator (including the difference between the purchase and sale price of virtual power plants and the revenue from transactions with the distribution network), as shown in the following formula:

[0039] Total electricity sold and purchased by all virtual power plants to virtual power plant operators; , Electricity purchase and sale prices in the distribution network; Power balance between virtual power plant operators and distribution networks ( (Energy storage charging and discharging power for virtual power plant operators).

[0040] (2) Constraints: Virtual power plant operators must simultaneously have both power purchase and power sale virtual power plants to ensure transaction feasibility. The formula is as follows: ( The number of virtual power plants for electricity purchase. (Number of virtual power plants for electricity sales).

[0041] 4. Definition of Nash Equilibrium A two-level game reaches Nash equilibrium when the following conditions are met: At the virtual power plant level: If any virtual power plant unilaterally changes its power purchase and sale strategy, its total cost will not decrease, as shown in the formula:

[0042] At the virtual power plant operator level: If a virtual power plant operator unilaterally changes its electricity price or power balancing strategy, its revenue will not increase, as shown in the formula:

[0043] in, To balance electricity prices, For virtual power plant power purchase and sale balance strategies, To balance power.

[0044] Furthermore, the above parameters must also satisfy system constraints: 1. Distributed energy output constraints: Actual wind / solar power output must fluctuate within ±30% of the predicted value, as shown in the formula:

[0045] 2. Energy storage charging and discharging constraints: The energy storage charging and discharging power must be within the rated range, as shown in the formula:

[0046] Step S20: Based on the fitness score, iteratively optimize the initial scheduling strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0047] Understandably, the iteration termination condition refers to the criteria for stopping the algorithm's iteration, including reaching the maximum number of iterations T, or the difference between the optimal fitness scores of two consecutive iterations being less than or equal to a preset accuracy threshold ξ. The scheduling strategy generation device iteratively optimizes the scheduling strategy based on its current fitness score, dynamically adjusting the strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition. This effectively avoids the problem of traditional optimization algorithms easily getting trapped in local optima, significantly improving the convergence speed and the accuracy of the scheduling scheme, and ultimately obtaining the optimal scheduling strategy that brings the interests of both the virtual power plant and the operator close to a Nash equilibrium.

[0048] For example, the scheduling strategy generation device starts with the particle position and velocity corresponding to the initial scheduling strategy and enters the iterative calculation stage of the IRLA algorithm: the device first calculates the adjustment factor β for each particle. When β ≥ 0.5, the particle swarm optimization framework is used to update the particle's velocity and position; when β < 0.5, a non-uniform mutation operator is used to migrate the particles to expand the search range; subsequently, the device combines a population feedback mechanism to further optimize the velocity update using historical particle information; in each iteration, the device recalculates the fitness score (i.e., the objective function value of the two-layer game model) corresponding to each particle and updates the individual optimal and global optimal solutions; as shown Figure 1 In the system structure shown, the iterative process continues until the number of iterations reaches the preset T=500 or the change in the optimal fitness score is less than ξ=0.001. At this point, the device outputs the global optimal solution as the final optimal scheduling strategy to complete the economic scheduling.

[0049] In this embodiment, a two-layer game model is constructed. The lower layer aims to minimize the operating costs (including maintenance and curtailment costs) and environmental pollution costs (electromagnetic pollution and battery leakage costs) of the virtual power plant, while the upper layer aims to maximize the revenue of the virtual power plant operator (VPO) from purchasing and selling electricity, thereby achieving a balance of interests among multiple stakeholders.

[0050] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Before step S10, the scheduling strategy generation method further includes steps S01~S02: Step S01: Randomly generate a particle swarm and set a scheduling decision vector for each particle in the particle swarm. The scheduling decision vector includes the particle position and update speed. The particle position includes the purchase and sale price of electricity by the virtual power plant operator and the purchase and sale power of the virtual power plant. It should be noted that the particle swarm refers to a set of candidate solutions used in parallel search of the solution space in improved reinforcement learning algorithms, and its population size is pre-set according to the complexity of the scheduling problem. The scheduling decision vector is the mathematical expression representing a complete scheduling scheme, containing all decision variables that need to be optimized during algorithm iteration. Particle position is a core component of the scheduling decision vector, representing a specific scheduling scheme, including the virtual power plant operator's purchase price Cb(t) and sales price Cs(t), and the purchase power Pb^i(t) and sales power Ps^i(t) of each virtual power plant i. The update rate is a core parameter in improved reinforcement learning algorithms, representing the direction and step size of particle position changes during iteration, used to guide the swarm towards the optimal solution region.

[0051] Understandably, the scheduling strategy generation device maps complex scheduling decisions to position vectors in the particle swarm optimization algorithm, incorporating key decision variables such as electricity price and power into a unified optimization framework. This lays the search foundation for subsequent iterative optimization based on improved reinforcement learning algorithms, enabling the algorithm to systematically explore a high-dimensional solution space that includes economic transactions and physical constraints.

[0052] Step S02: Use the scheduling parameters in the particle position as the initial scheduling strategy.

[0053] It should be noted that the scheduling policy generation device directly maps particle positions to initial scheduling policies, establishing a bridge between the improved reinforcement learning algorithm and the two-layer game model. This enables the candidate solutions generated by the algorithm to be accurately evaluated by the objective function, providing a feasible starting point for subsequent iterative optimization.

[0054] For example, when the scheduling policy generation device executes the improved reinforcement learning algorithm, it first initializes a particle swarm containing K particles: the device randomly generates a scheduling decision vector for each particle k, where the particle position xk contains the sequence of the purchase price Cs(t) and sales price Cb(t) of the virtual power plant operator to the virtual power plant in the next 24 hours (NT=96 time periods), as well as the sequence of the purchase power Pb^i(t) and sales power Ps^i(t) of virtual power plants A, B, C, etc., respectively with the operator; at the same time, it randomly initializes an update rate vk for each decision variable dimension.

[0055] In one feasible implementation, step S10 includes: Based on the fitness score, determine the current optimal scheduling strategy; Understandably, the scheduling strategy generation device obtains the global historical best scheduling strategy (gbest) by comparing the fitness scores of all particles. That is, the strategy with the highest score that the entire particle swarm has found in all iterations is taken as the global historical best scheduling strategy (gbest).

[0056] For example, after completing one iteration of calculation, the scheduling policy generation device evaluates the fitness scores of all K particles at their new positions: the device first compares the new score of each particle k with its own recorded individual historical best scheduling policy pbest score. If the new score is better (lower cost for the lower-level model and higher benefit for the upper-level model), the device updates the particle's pbest with the new particle position (i.e., the new combination of electricity price and power). Subsequently, the device compares the pbest of all particles with the currently recorded global historical best scheduling policy gbest, and selects the policy with the best score as the new gbest.

[0057] The adjustment factor for each particle is calculated, and the particle update mechanism is determined based on the adjustment factor. The particle update mechanism includes a particle swarm optimization mechanism and a non-uniform mutation operator mechanism. It should be noted that the adjustment factor refers to the key parameter β used in the improved reinforcement learning algorithm for dynamically selecting the particle update strategy. The particle swarm optimization mechanism refers to the swarm intelligence-based search method used when the adjustment factor β ≥ 0.5, updating particle velocity and position through the individual historical best (pbest) and the global historical best (gbest). The non-uniform mutation operator mechanism refers to the mutation operation used when the adjustment factor β < 0.5, applying non-linear perturbation to the particle position. The scheduling strategy generation device calculates the adjustment factor for the current iteration based on the initially set adjustment factor or the adjustment factor calculated in the previous iteration.

[0058] Specifically, the adjustment factor maintains particle diversity and avoids local optima; the formula is as follows:

[0059] in, ( (These are random numbers from a standard normal distribution). is the adjustment factor for particle k in the d-th iteration; This is the penalty coefficient.

[0060] Furthermore, when At that time, the PSO framework is used to update the particle position and velocity:

[0061] when At this time, a non-uniform mutation operator is used to migrate particles, thereby expanding the search range:

[0062] Population feedback: Optimize speed updates by incorporating historical particle information; the formula is:

[0063] S(·) is the tanh activation function; , These are random numbers distributed according to a standard normal distribution. This represents the position of historical particles.

[0064] in, (T is the maximum number of iterations, and α is the uniformity parameter of variation).

[0065] Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is iteratively updated to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0066] It should be noted that the scheduling strategy generation device iteratively updates by combining the dynamic selection update mechanism with the historical best strategy, enabling the particle swarm to continuously explore a better solution space based on the discovery of high-quality solutions. This ensures the convergence of the algorithm and effectively prevents premature convergence, thereby systematically approximating the globally optimal scheduling scheme that balances the interests of both the virtual power plant and the operator.

[0067] In one feasible implementation, the step of iteratively updating the scheduling decision vector of each particle based on the particle update mechanism and the current globally optimal scheduling strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition includes: Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is updated to obtain the updated scheduling decision vector. Understandably, the scheduling strategy generation device recalculates and replaces the position and velocity values ​​in the particle scheduling decision vector based on the update mechanism determined in the current iteration.

[0068] Determine whether the updated scheduling decision vector satisfies the iteration termination condition; It should be noted that the scheduling strategy generation device determines whether the algorithm state has reached the preset convergence criterion after each iteration, providing a clear stopping criterion for the optimization process. This ensures that the algorithm can terminate in time when a satisfactory solution is found to save computing resources, and also prevents the situation where the algorithm fails to find the truly optimal solution due to premature termination, thus guaranteeing the quality of the scheduling scheme and the efficiency of the algorithm.

[0069] If not satisfied, determine the fitness score of the updated scheduling decision vector and return the step of determining the current optimal scheduling strategy based on the fitness score, until the change of the updated target scheduling decision vector from the scheduling decision vector obtained after the previous iteration is less than the threshold. Understandably, the scheduling strategy generation device re-inputs the newly generated scheduling decision vector into the preset two-layer game model, calculates the new fitness score, and returns the new fitness score to the step of determining the current optimal scheduling strategy based on the fitness score. This process continues until the absolute value of the difference between the global historical optimal fitness scores obtained from two consecutive iterations, i.e., |gbest_d - gbest_{d-1}|, is less than the preset accuracy threshold ξ. By establishing a "calculation-evaluation-feedback" cyclic optimization mechanism, the algorithm can automatically continue the search process when the convergence criterion is not reached. It continuously improves the scheduling strategy using the new information obtained in each iteration, ensuring that the final solution can closely approximate the theoretical optimal value, thereby effectively improving the accuracy and reliability of the economic scheduling of the virtual power plant.

[0070] The current optimal scheduling strategy is taken as the optimal scheduling strategy that satisfies the iteration termination condition.

[0071] It should be noted that the current optimal scheduling strategy refers to the scheduling scheme with the best fitness score found by the entire particle swarm so far during the execution of the improved reinforcement learning algorithm, i.e., the global historical optimal solution gbest. This solution includes the optimal purchase and sale price sequence for virtual power plant operators and the optimal purchase and sale power sequence for each virtual power plant. The scheduling strategy generation device ensures that the solution simultaneously considers the operational economics of virtual power plants and the maximization of revenue for virtual power plant operators by determining the global optimal solution at the final convergence point as the optimal scheduling strategy for the system. This balances the interests of both parties near the Nash equilibrium point, providing a scientific and effective decision-making basis for virtual power plants to participate in the electricity market.

[0072] In one feasible implementation, the step of determining the current optimal scheduling strategy based on the fitness score includes: Based on the fitness score, the current individual optimal strategy for each particle is determined; Understandably, the individual optimal strategy refers to the scheduling decision vector with the best fitness score found by each particle in its own search history up to the current iteration, i.e., the individual historical optimal solution (pbest). After the scheduling strategy generation device calculates the new fitness scores of all particles in the d-th iteration, it performs the following operation on particle k: compares the score of the particle's new position with its recorded pbest score. If the new score is better (lower total cost for the lower-level virtual power plant model, and higher total revenue for the upper-level operator model), then the particle's pbest is updated using the new particle position (including the new electricity price and power combination); otherwise, the original pbest is retained.

[0073] Determine the optimal strategy for the target individual with the highest fitness score among the individual optimal strategies; It should be noted that the target individual optimal strategy refers to the single strategy with the best fitness score selected from the historical best strategies (pbest) of all particles. In the d-th iteration, after updating the pbest of all particles, the scheduling strategy generation device iterates through and compares the pbest fitness scores of particles 1 to K: it finds that the pbest of particle 15 (corresponding to a specific combination of electricity price and power) has the highest score among all current individual optimal strategies (lowest cost for the lower-level model and highest benefit for the upper-level model), and thus determines this strategy as the target individual optimal strategy (i.e., the new gbest) for this iteration.

[0074] Determine whether the fitness score of the optimal strategy for the target individual is greater than the fitness score of the optimal scheduling strategy; It is understandable that the optimal scheduling policy refers to the globally best policy (gbest) recorded by the algorithm before the current iteration. After the scheduling policy generation device determines the optimal policy of the target individual (i.e., the new candidate gbest) in the d-th iteration, it compares its fitness score with the currently recorded score of the globally best policy.

[0075] If the value is greater than the target value, the optimal scheduling policy will be updated to the target optimal policy.

[0076] It should be noted that the scheduling strategy generation device completely replaces the currently stored global historical best strategy (original gbest) with the newly determined target individual optimal strategy (candidate gbest). In the d-th iteration, after the device determines that the fitness score (total cost of the lower-level model 10500) of the target individual optimal strategy (such as pbest of particle 15) is greater than (i.e. better than) the score (cost 10600) of the current optimal scheduling strategy, it immediately performs an update operation: replacing the content of the global historical best gbest with all the scheduling parameters (including its electricity price sequence and power sequence) contained in pbest of particle 15.

[0077] In one feasible implementation, the step of updating the scheduling decision vector of each particle based on the particle update mechanism and the optimal scheduling strategy to obtain the updated scheduling decision vector includes: When the particle update mechanism is a non-uniform mutation operator mechanism, perturbation random numbers are randomly generated. Understandably, the perturbation random number refers to the random variable r used to determine the direction and magnitude of particle mutation during the non-uniform mutation process. This random number follows a uniform distribution in the interval [0,1] and is a key parameter controlling the mutation operation. After determining that particle k adopts the non-uniform mutation operator mechanism, the scheduling strategy generation device first generates a perturbation random number that follows a uniform distribution in [0,1]. If r=0.3, since r<0.5, the device performs positive mutation on the particle position, where the Δ function ensures that the mutation magnitude gradually decreases as the number of iterations increases.

[0078] Based on the perturbation random number, the migration direction and migration offset are determined; It should be noted that the migration direction refers to the direction in which the particle position mutates in the solution space, determined by the comparison between the perturbation random number r and the threshold 0.5. When r < 0.5, the particle migrates towards the upper bound of the solution space; when r ≥ 0.5, it migrates towards the lower bound of the solution space. The migration offset refers to the specific adjustment magnitude of the particle position during the mutation operation. The scheduling strategy generation device achieves adaptive control of particle position mutation by combining random numbers with a nonlinear mutation function. This ensures that the algorithm has strong global exploration capabilities in the early stages of the search to escape local optima, and also has refined local exploitation capabilities in the later stages of the search to improve convergence accuracy.

[0079] The scheduling decision vector is shifted along the migration direction by the migration offset to obtain the updated scheduling decision vector.

[0080] Understandably, the scheduling strategy generation device applies the calculated theoretical offset to the particle position, completing the core operation of the non-uniform mutation operator mechanism, with the specific direction determined by the migration direction.

[0081] In this embodiment, by integrating particle swarm optimization (PSO) and reinforcement learning techniques, and introducing adjustment factors, non-uniform mutation operators, and population feedback mechanisms, the algorithm is prevented from getting trapped in local optima, thereby improving convergence speed and scheduling accuracy.

[0082] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the scheduling strategy generation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0083] This application also provides a scheduling strategy generation device, please refer to... Figure 4 The scheduling strategy generation device includes: The generation module 10 is used to randomly generate an initial scheduling strategy for a preset scheduling period, input the initial scheduling strategy into a preset two-layer game model, and obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. The optimization module 20 is used to iteratively optimize the initial scheduling strategy based on the fitness score to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0084] Optionally, the generation module includes: A submodule is configured to randomly generate a particle swarm and set a scheduling decision vector for each particle in the swarm. The scheduling decision vector includes the particle position and update velocity. The particle position includes the purchase and sale price of electricity by the virtual power plant operator and the purchase and sale power of the virtual power plant. The scheduling parameters in the particle position are used as the initial scheduling strategy.

[0085] Optionally, the setting submodule includes: The computational unit is used to determine the current optimal scheduling strategy based on the fitness score; calculate the adjustment factor for each particle; determine the particle update mechanism based on the adjustment factor, wherein the particle update mechanism includes a particle swarm optimization mechanism and a non-uniform mutation operator mechanism; and iteratively update the scheduling decision vector of each particle based on the particle update mechanism and the optimal scheduling strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

[0086] Optionally, the computing unit includes: The first judgment subunit is used to update the scheduling decision vector of each particle based on the particle update mechanism and the optimal scheduling strategy to obtain the updated scheduling decision vector; determine whether the updated scheduling decision vector satisfies the iteration termination condition; if not, determine the fitness score of the updated scheduling decision vector, and return to the step of determining the current optimal scheduling strategy based on the fitness score, until the change in the updated target scheduling decision vector compared with the scheduling decision vector obtained after the previous iteration is less than a threshold; and take the current optimal scheduling strategy as the optimal scheduling strategy that satisfies the iteration termination condition.

[0087] The second judgment subunit is used to determine the current individual optimal strategy of each particle based on the fitness score; determine the target individual optimal strategy with the highest fitness score among the individual optimal strategies; determine whether the fitness score of the target individual optimal strategy is greater than the fitness score of the optimal scheduling strategy; if it is greater, then update the optimal scheduling strategy to the target optimal strategy.

[0088] Optionally, the first determination subunit includes: An update component is used to randomly generate perturbation random numbers when the particle update mechanism is a non-uniform mutation operator mechanism; determine the migration direction and migration offset based on the perturbation random numbers; and move the scheduling decision vector along the migration direction by the migration offset to obtain the updated scheduling decision vector.

[0089] The scheduling strategy generation apparatus provided in this application, employing the scheduling strategy generation method in the above embodiments, can solve the technical problem of scheduling strategy generation. Compared with the prior art, the beneficial effects of the scheduling strategy generation apparatus provided in this application are the same as those of the scheduling strategy generation method provided in the above embodiments, and other technical features in the scheduling strategy generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0090] This application provides a scheduling policy generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the scheduling policy generation method in Embodiment 1 above.

[0091] The following is for reference. Figure 5The diagram illustrates a structural schematic of a scheduling policy generation device suitable for implementing embodiments of this application. The scheduling policy generation device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The scheduling policy generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0092] like Figure 5 As shown, the scheduling policy generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the scheduling policy generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the scheduling policy generation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows scheduling policy generation devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0093] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0094] The scheduling strategy generation device provided in this application, employing the scheduling strategy generation method in the above embodiments, can solve the technical problem of scheduling strategy generation. Compared with the prior art, the beneficial effects of the scheduling strategy generation device provided in this application are the same as those of the scheduling strategy generation method provided in the above embodiments, and other technical features in this scheduling strategy generation device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0095] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0097] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the scheduling policy generation method in the above embodiments.

[0098] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0099] The aforementioned computer-readable storage medium may be included in the scheduling policy generation device; or it may exist independently and not be assembled into the scheduling policy generation device.

[0100] The aforementioned computer-readable storage medium carries one or more programs. When the one or more programs are executed by the scheduling strategy generation device, the scheduling strategy generation device: randomly generates an initial scheduling strategy for a preset scheduling period; inputs the initial scheduling strategy into a preset two-layer game model to obtain a fitness score, wherein the upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model; and iteratively optimizes the initial scheduling strategy based on the fitness score to obtain an optimal scheduling strategy that satisfies the iteration termination condition.

[0101] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0104] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described scheduling policy generation method, thereby solving the technical problem of scheduling policy generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the scheduling policy generation method provided in the above embodiments, and will not be repeated here.

[0105] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the scheduling policy generation method described above.

[0106] The computer program product provided in this application can solve the technical problem of scheduling policy generation. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the scheduling policy generation method provided in the above embodiments, and will not be repeated here.

[0107] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.

Claims

1. A method for generating scheduling strategies, characterized in that, The scheduling policy generation method includes: An initial scheduling strategy for a preset scheduling period is randomly generated. The initial scheduling strategy is then input into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. Based on the fitness score, the initial scheduling strategy is iteratively optimized to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

2. The scheduling strategy generation method as described in claim 1, characterized in that, The step of randomly generating the initial scheduling strategy for a preset scheduling period includes the following: A particle swarm is randomly generated, and a scheduling decision vector is set for each particle in the particle swarm. The scheduling decision vector includes the particle position and update speed. The particle position includes the purchase and sale price of electricity by the virtual power plant operator and the purchase and sale power of the virtual power plant. The scheduling parameters at the particle positions are used as the initial scheduling strategy.

3. The scheduling strategy generation method as described in claim 2, characterized in that, The step of iteratively optimizing the initial scheduling strategy based on the fitness score to obtain the optimal scheduling strategy that satisfies the iteration termination condition includes: Based on the fitness score, determine the current optimal scheduling strategy; The adjustment factor for each particle is calculated, and the particle update mechanism is determined based on the adjustment factor. The particle update mechanism includes a particle swarm optimization mechanism and a non-uniform mutation operator mechanism. Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is iteratively updated to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

4. The scheduling strategy generation method as described in claim 3, characterized in that, The step of iteratively updating the scheduling decision vector of each particle based on the particle update mechanism and the current globally optimal scheduling strategy to obtain the optimal scheduling strategy that satisfies the iteration termination condition includes: Based on the particle update mechanism and the optimal scheduling strategy, the scheduling decision vector of each particle is updated to obtain the updated scheduling decision vector. Determine whether the updated scheduling decision vector satisfies the iteration termination condition; If not satisfied, determine the fitness score of the updated scheduling decision vector and return the step of determining the current optimal scheduling strategy based on the fitness score, until the change of the updated target scheduling decision vector from the scheduling decision vector obtained after the previous iteration is less than the threshold. The current optimal scheduling strategy is taken as the optimal scheduling strategy that satisfies the iteration termination condition.

5. The scheduling strategy generation method as described in claim 3, characterized in that, The step of determining the current optimal scheduling strategy based on the fitness score includes: Based on the fitness score, the current individual optimal strategy for each particle is determined; Determine the optimal strategy for the target individual with the highest fitness score among the individual optimal strategies; Determine whether the fitness score of the optimal strategy for the target individual is greater than the fitness score of the optimal scheduling strategy; If the value is greater than the target value, the optimal scheduling policy will be updated to the target optimal policy.

6. The scheduling strategy generation method as described in claim 4, characterized in that, The step of updating the scheduling decision vector of each particle based on the particle update mechanism and the optimal scheduling strategy to obtain the updated scheduling decision vector includes: When the particle update mechanism is a non-uniform mutation operator mechanism, perturbation random numbers are randomly generated. Based on the perturbation random number, the migration direction and migration offset are determined; The scheduling decision vector is shifted along the migration direction by the migration offset to obtain the updated scheduling decision vector.

7. A scheduling strategy generation device, characterized in that, The device includes: The generation module is used to randomly generate an initial scheduling strategy for a preset scheduling period, and input the initial scheduling strategy into a preset two-layer game model to obtain a fitness score. The upper layer of the preset two-layer game model is a virtual power plant operator revenue optimization model, and the lower layer is a virtual power plant revenue optimization model. An optimization module is used to iteratively optimize the initial scheduling strategy based on the fitness score to obtain the optimal scheduling strategy that satisfies the iteration termination condition.

8. A scheduling strategy generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the scheduling policy generation method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the scheduling policy generation method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the scheduling policy generation method as described in any one of claims 1 to 6.