A device activation and task offloading collaborative optimization method, device and medium for frequency regulation of a virtual power plant

CN122801296APending Publication Date: 2026-09-22SHANDONG JIANZHU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610807488.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

随着参与设备数量增加,系统需要处理的状态数据量和调频相关任务量也会同步增加,从而加重通信传输和协同计算负担;与此同时,通信条件较差或处理速度较慢的个别设备还可能成为整体调频过程中的瓶颈,导致系统综合响应时延增大,影响控制指令统一下发的及时性

Benefits of technology

[0016]本发明的优点在于:本发明将设备激活决策和任务卸载决策进行联合优化,避免了现有方法中先固定设备集合、再单独优化任务卸载方式所导致的决策割裂问题。通过在每个调频控制周期内动态选择参与调频的设备集合,系统能够在满足目标调节功率需求的基础上,减少不必要的设备参与数量,降低状态上传规模和协同计算负担。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801296A_ABST
    Figure CN122801296A_ABST
Patent Text Reader

Abstract

The application provides a device activation and task unloading cooperative optimization method and device for frequency regulation of a virtual power plant and a medium, and belongs to the technical field of frequency regulation of a virtual power plant. The method constructs a state vector by collecting system operation information, inputs a trained diffusion strategy network to generate a joint decision action containing device activation and task unloading. Accordingly, the device uploads information, and the frequency regulation task is allocated to edge and cloud server cooperative processing. Based on the processing result, the time delay and cost are calculated to construct a reward function, the network is updated in parameters, and the cycle iteration is performed until convergence to obtain an optimal strategy. The application can effectively improve the real-time performance, stability and economy of the frequency regulation service of the virtual power plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, apparatus, and medium for collaborative optimization of equipment activation and task offloading for frequency regulation in virtual power plants, belonging to the field of virtual power plant frequency regulation technology. Background Technology

[0002] Virtual power plants participate in grid frequency regulation services by aggregating distributed resources such as energy storage devices and adjustable loads. During frequency regulation control, operations such as equipment status acquisition, data communication transmission, frequency regulation-related calculations, and control command issuance typically need to be completed within a short control cycle, thus placing high demands on the system's real-time performance and coordination. Especially in a cloud-edge-device collaborative architecture, the effectiveness of frequency regulation control depends not only on the adjustability of the equipment itself but also on the combined influence of communication link status, dynamic changes in edge-side computing resources, and cloud-side computing resources.

[0003] In actual operation, the wireless communication link between terminal devices and edge nodes is easily affected by factors such as device location, obstruction, interference, and channel fading, resulting in random fluctuations in data transmission rate. Simultaneously, the available computing power on both the cloud and edge sides dynamically changes with the workload. Because frequency regulation tasks are characterized by short control cycles and high response time requirements, significant delays in status information uploading, task processing, or control command generation can easily lead to lag in frequency regulation command issuance. This, in turn, affects the timely output of the actual regulated power, reduces the frequency regulation performance of the virtual power plant, and increases overall operating costs.

[0004] In existing technologies, most optimization methods for frequency regulation in virtual power plants focus on a single aspect. One type of method typically predetermines the set of devices participating in frequency regulation, optimizing only the offloading method of computing tasks between the cloud and the edge. Another type of method uses static rules, empirical thresholds, or fixed strategies to select devices participating in frequency regulation, making it difficult to dynamically adjust the set of devices and task processing methods based on real-time communication conditions, computing power status, and target regulation power requirements. These methods generally fail to provide unified modeling and collaborative optimization for device activation decisions and computation offloading decisions, thus making it difficult to simultaneously address multiple requirements such as target regulation power attainability, communication and computation latency, and cloud-edge resource consumption.

[0005] Furthermore, in virtual power plant frequency regulation scenarios, optimal control results cannot be achieved simply by having all adjustable devices participate in frequency regulation within each control cycle. As the number of participating devices increases, the amount of status data and frequency regulation-related tasks that the system needs to process also increases, thus increasing the burden on communication transmission and collaborative computing. At the same time, individual devices with poor communication conditions or slow processing speeds may become bottlenecks in the overall frequency regulation process, leading to increased overall system response latency and affecting the timeliness of unified control command issuance. Therefore, existing "full activation" or coarse-grained device selection methods are insufficient to meet the real-time frequency regulation requirements under short control cycle conditions.

[0006] Therefore, there is an urgent need to provide a cloud-edge-device collaborative optimization technology solution for virtual power plant frequency regulation services. This solution can jointly optimize the activation decisions of equipment involved in frequency regulation and the offloading decisions of computing tasks in dynamic communication and computing environments. While meeting the target regulation power requirements, it can reduce the overall latency and operating costs of communication and computing, thereby improving the real-time performance, stability and economy of virtual power plant frequency regulation services. Summary of the Invention

[0007] The purpose of this invention is to provide a method, apparatus, and medium for collaborative optimization of equipment activation and task offloading in virtual power plant frequency regulation. This invention addresses the problems in existing technologies, such as the separation of equipment activation decisions and computational task offloading decisions, difficulty in adapting to dynamic communication and computing environments, and difficulty in simultaneously considering the achievability of target regulation power, response latency, and operating costs. The invention aims to improve the real-time performance, stability, and economy of virtual power plant frequency regulation services.

[0008] To achieve the above objectives, the present invention employs the following technical solution: A collaborative optimization method for equipment activation and task offloading in virtual power plant frequency regulation includes the following steps: S1: Collect system operation information within the current frequency modulation control cycle and construct the system state vector; S2: Input the system state vector into the trained diffusion strategy network, and the diffusion strategy network generates a joint decision action for the current control cycle. The joint decision action includes a device activation decision and a task offloading decision. The device activation decision is used to determine the set of target devices participating in frequency modulation, and the task offloading decision is used to determine the allocation ratio of frequency modulation-related tasks generated by the activated devices between the edge server and the cloud server. S3: Based on the device activation decision, control the activated device to upload status information and frequency modulation tasks, and based on the task unloading decision, allocate the effective frequency modulation task volume to the edge server and cloud server for collaborative processing; S4: Calculate the comprehensive latency and collaborative operation cost based on the effective task volume of edge servers and cloud servers, establish an instant reward function, and update the parameters of the diffusion strategy network based on the instant reward. S5: Execute S1-S4 repeatedly to continuously optimize equipment activation decisions and task unloading decisions until the instantaneous reward function converges, obtain the optimal diffusion strategy network, collect the operation information of the virtual power plant dispatch system to be optimized, and obtain the optimization strategy through the optimal diffusion strategy network.

[0009] Preferably, the system operation information includes the target adjustment power, the maximum available adjustment capability of each frequency modulation device, the wireless channel status, the device workload, the device ramp-up time, the available computing power of the edge server, and the available computing power of the cloud server.

[0010] Preferably, the joint decision action for the current control cycle generated by the diffusion strategy network specifically includes: Using a diffusion policy network with the system state vector as a condition, a joint action latent representation is generated through a reverse denoising process; The first N-dimensional variables of the potential representation of the joint action are transformed by the Sigmoid function and thresholded to obtain discrete device activation results; The last N-dimensional variables of the potential representation of the joint action are mapped to the [0,1] interval using the Sigmoid function to obtain the continuous task unloading ratio; Where N is the total number of adjustable devices in the virtual power plant.

[0011] Preferably, the method for obtaining the effective frequency modulation task quantity specifically includes: For activated devices, multiply their original task volume by the device activation decision to obtain the effective frequency modulation task volume; The effective frequency modulation task volume is obtained by calculating the effective frequency modulation task volume processed by the edge server based on the task unloading ratio in the task unloading decision. The effective frequency modulation task volume processed by the cloud server is obtained by subtracting the effective frequency modulation task volume processed by the edge server from the effective task volume.

[0012] Preferably, the instant reward function is as follows: , in, For instant reward function, For frequency modulation performance coefficient, Indicates the first Frequency modulation performance coefficient within each control cycle The weighting coefficients corresponding to the cloud-edge computing cost. Indicates the first The cost of coordinated operation within a control cycle; The long-term optimization objective is: , in, For strategy The corresponding long-term expected cumulative discount reward function, Indicates the strategy The generated state-action trajectory is expected. As a discount factor, These are strategy parameters.

[0013] Preferably, the frequency modulation performance coefficient is calculated as follows: , in, For equipment In the The target regulation power undertaken within each control cycle This indicates the target regulating power for this control cycle. Indicates the control period. Indicates the frequency modulation settlement cycle. For the first The overall delay within each control cycle For equipment The time spent climbing the hill, For the first A set of activated devices for each control cycle. Indicates the number of control cycles contained within a single settlement cycle; The method for calculating the collaborative operation cost is as follows: , in, This represents the cost of edge server computation per unit of time. Represents edge server In the The computational delay within each control cycle This represents the cost of cloud server computing per unit time. Indicates the cloud server is in the The computational delay within each control cycle This indicates the number of edge servers.

[0014] Preferably, the comprehensive delay calculation method is as follows: , , , , , , , , in, For edge servers In the Task completion delay within each control cycle For cloud servers in the first Task completion delay within each control cycle Calculate the latency based on the total number of tasks received by the cloud server. For cloud server task arrival latency, For cloud server task upload latency, For equipment In the The amount of tasks allocated to the cloud server within a control cycle. For equipment The wireless transmission rate to its access edge server, For edge servers The return speed between the cloud server and the server For edge server task arrival latency, For edge servers The calculation delay for the total amount of received tasks. For the number of edge servers, For edge server task upload latency, For equipment In the The amount of tasks allocated to the edge server within a control cycle The number of CPU cycles required per unit bit of task data. For edge servers Total number of tasks received For edge servers In the Computing capacity within a control cycle The number of devices can be adjusted. Calculate the latency based on the total number of tasks received by the cloud server. This represents the total number of tasks received by the cloud server. This represents the computing power of the cloud server during the kth control cycle.

[0015] A device for collaborative optimization of equipment activation and task offloading for frequency regulation in virtual power plants includes: The state awareness and modeling module is used to collect system operation information within the current frequency regulation control cycle and construct the system state vector. The joint decision generation module is used to input the system state vector into the diffusion policy network and generate joint decision actions that include device activation decisions and task unloading decisions; The activation and unloading execution module is used to determine the set of devices to be activated based on the device activation decision, and to allocate valid tasks to the edge side and cloud side based on the task unloading decision; The cloud-edge collaborative processing module is used to control the edge server and cloud server to perform task processing and determine the overall latency and operating cost; The frequency modulation control and performance evaluation module is used to generate and issue frequency modulation control commands, and at the same time evaluate the aggregated frequency modulation response results. The reward calculation and strategy update module is used to calculate real-time rewards based on latency, power satisfaction, response results, and cost, and update the diffusion strategy network parameters.

[0016] The advantages of this invention are as follows: This invention jointly optimizes device activation decisions and task offloading decisions, avoiding the decision-making fragmentation problem caused by fixing the device set first and then optimizing the task offloading method separately in existing methods. By dynamically selecting the set of devices participating in frequency regulation within each frequency regulation control cycle, the system can reduce the number of unnecessary devices involved, reduce the scale of state uploads, and decrease the burden of collaborative computing while meeting the target regulation power requirements.

[0017] This invention comprehensively considers the wireless channel status, the available computing power of the edge server, and the available computing power of the cloud server to dynamically determine the allocation ratio of frequency modulation-related tasks between the edge and cloud sides. Compared with fixed offloading or single-side processing methods, this invention can adjust the task processing position according to the real-time communication environment and computing resource status, thereby reducing the overall communication-computing latency.

[0018] This invention combines communication-computing integrated latency with the dynamic response process of the equipment, and further considers the impact of the equipment ramp-up time on the frequency modulation response effect. This makes the frequency modulation performance evaluation not only reflect whether the target regulation power is met, but also reflect the impact of the control command activation lag and the equipment dynamic response lag on the aggregated frequency modulation power.

[0019] This invention employs a diffusion strategy network to generate joint decision-making actions, enabling unified modeling of a hybrid action space that includes discrete device activation variables and continuous task unloading ratios, thereby improving the expressive and optimization capabilities of joint decision-making in complex dynamic environments. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0021] Figure 1This is a schematic diagram of the device activation and task unloading collaborative optimization method provided in an embodiment of the present invention.

[0022] Figure 2 This is a functional structure block diagram of the device activation and task unloading collaborative optimization device provided in an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1 like Figure 1 As shown, to address the problems in existing technologies such as the separation of equipment activation decisions and computational task offloading decisions, difficulty in adapting to dynamic communication and computing environments, and difficulty in simultaneously considering the achievability of target regulation power, response latency, and operating costs, this invention provides a collaborative optimization method for equipment activation and task offloading in virtual power plant frequency regulation. In one embodiment, the method generates equipment activation and task offloading decisions based on a diffusion strategy network. This method and apparatus, designed for virtual power plant frequency regulation service scenarios, jointly optimize the set of equipment participating in frequency regulation and the processing location of frequency regulation-related tasks within a cloud-edge-device collaborative architecture, thereby improving the real-time performance, stability, and economy of virtual power plant frequency regulation services.

[0025] Specifically, the method is executed within a cloud-edge-device collaborative architecture comprised of a cloud server, an edge server, and a terminal frequency modulation device; this invention operates on two time scales: the frequency modulation settlement cycle and the frequency modulation control cycle. The frequency modulation settlement cycle is... The control cycle is A single settlement cycle contains One control cycle, satisfying Within each control cycle, the system sequentially performs status acquisition, joint decision generation, task uploading, cloud-edge collaborative processing, control command generation, device response execution, and reward feedback update.

[0026] Diffusion policy networks are essentially policy networks in reinforcement learning, but they incorporate the action generation mechanism of a diffusion model. In other words, they belong to a type of diffusion-based reinforcement learning policy network used to generate decisions on device participation and task offloading based on the current system state.

[0027] Includes the following steps: S1: Collect system operation information within the current frequency modulation control cycle and construct the system state vector; S2: Input the system state vector into the trained diffusion strategy network, and the diffusion strategy network generates a joint decision action for the current control cycle. The joint decision action includes a device activation decision and a task offloading decision. The device activation decision is used to determine the set of target devices participating in frequency modulation, and the task offloading decision is used to determine the allocation ratio of frequency modulation-related tasks generated by the activated devices between the edge server and the cloud server. S3: Based on the device activation decision, control the activated device to upload status information and frequency modulation tasks, and based on the task unloading decision, allocate the effective frequency modulation task volume to the edge server and cloud server for collaborative processing; S4: Calculate the comprehensive latency and collaborative operation cost based on the effective task volume of edge servers and cloud servers, establish an instant reward function, and update the parameters of the diffusion strategy network based on the instant reward. S5: Execute S1-S4 repeatedly to continuously optimize equipment activation decisions and task unloading decisions until the instantaneous reward function converges, obtain the optimal diffusion strategy network, collect the operation information of the virtual power plant dispatch system to be optimized, and obtain the optimization strategy through the optimal diffusion strategy network.

[0028] As a refinement of the above embodiments, the system operation information in step S1 includes the target regulation power, the maximum available regulation capability of each frequency modulation device, the wireless channel status, the device workload, the device ramp-up time, the available computing power of the edge server, and the available computing power of the cloud server. Based on the above information, a system state vector for the current control cycle is constructed: , Wherein, the system state vector It is used to characterize the frequency modulation demand, equipment capabilities, communication environment, and cloud-edge computing resource status within the current frequency modulation control cycle, and serves as input to the diffusion strategy network. Adjust power to the target. For equipment Maximum available adjustment capacity For equipment The wireless channel status, For equipment The time spent climbing the hill, For equipment The amount of frequency modulation tasks. Edge server Available computing power Available computing power of cloud servers.

[0029] Through this step, the system can sense the target adjustment power demand, terminal device operating status, wireless link quality, and cloud-edge computing power changes in each control cycle, providing real-time status information for subsequent joint decision-making.

[0030] As a refinement of the above embodiment, in step S2, the system state vector of the current control cycle is obtained. Then, it is input into the diffusion strategy network, which generates the joint decision action for the current control cycle. The joint decision action includes device activation decision and task unloading decision.

[0031] Equipment activation decisions are used to determine the set of target equipment to participate in the frequency regulation service of the current control cycle. For equipment... Define the activation variables for device activation decisions: , when When it is 1, it indicates that the device In the It is activated and participates in frequency modulation within each control cycle; when When it is 0, it indicates that the device The device does not participate in frequency modulation during the current control cycle. Therefore, the set of activated devices for the current control cycle is: , Task offloading decisions are used to determine the allocation ratio of frequency modulation tasks generated by activated devices between the edge and cloud sides. For devices... Define the task unloading ratio: , in, Indicates equipment The proportion of effective frequency modulation tasks that are offloaded to the edge side (edge ​​server) for processing. Indicates equipment The proportion of effective frequency modulation tasks that are offloaded to the cloud side (cloud server) for processing.

[0032] In this invention, the device activation variable It belongs to discrete decision variables, while the task unloading ratio These are continuous decision variables. Together, they constitute a hybrid action space. If a standard single-action output structure is used, problems arise such as difficulty in uniformly modeling discrete and continuous actions, high dimensionality of the action space, complex combinations of candidate actions, and unstable policy search. Therefore, this invention employs a diffusion policy network to generate a joint action latent representation and simultaneously obtains device activation results and task offloading ratios through action mapping.

[0033] Specifically, the joint action generated by the diffusion strategy network can be potentially represented as: , Among them, the former The dimension is used to indicate the activation tendency of each device, after which... The dimension is used to represent the task offloading tendency of each device. The diffusion policy network uses the system state vector. As conditional information, the potential representation of the joint action is generated.

[0034] During training, the diffusion policy network progressively adds noise to the latent representation of joint actions through a forward diffusion process, gradually transforming the original joint action representation into a noisy representation that follows a Gaussian distribution. The forward diffusion process can be represented as follows: , in, Indicates the first The potential representation of the action in each diffusion step, Indicates the first The noise figure for each diffusion step. Represents the identity matrix. Indicates a Gaussian distribution. This represents the conditional transition distribution of the forward diffusion process.

[0035] During the action generation process, the diffusion policy network diffuses random noise. Starting with the system state vector Given the condition, the joint action latent representation is gradually recovered through a multi-step inverse denoising process. The reverse denoising process can be represented as: , in, Indicates the network parameters of the diffusion strategy. and Let represent the mean and variance of the output from the diffusion strategy network, respectively. This represents the inverse denoising conditional probability distribution parameterized by the diffusion strategy network. This is achieved by introducing the system state. As a condition, the diffusion strategy network can adjust power, device capabilities, wireless channel status, and cloud-edge computing power status according to the current target to generate joint actions that match the current control cycle.

[0036] In obtaining the potential representation of joint actions Then, it is mapped to executable device activation and task unloading decisions. For the device activation part, the first N dimensions of latent variables are taken and converted into activation probabilities using the Sigmoid function. The activation result is mapped to the activation probability and the preset threshold. When the activation probability is greater than or equal to the preset threshold, the variable is activated. The value is 1, otherwise it is 0.

[0037] For the task unloading portion, the remaining N-dimensional latent variables are taken and mapped to the [0,1] interval using the Sigmoid function to obtain the task unloading ratio. .

[0038] Therefore, the diffusion strategy network can transform the generation process of high-dimensional joint actions into a conditional denoising process, and obtain discrete device activation actions and continuous task unloading actions simultaneously through action mapping.

[0039] Through the above approach, the diffusion strategy network in this invention can handle two coupled decision problems—device activation and task offloading—within the same decision framework. The device activation result affects the set of devices participating in frequency modulation, the degree to which the target regulation power is met, and the scale of state uploads; the task offloading ratio affects the allocation of tasks between the edge and cloud sides, the combined communication-computing latency, and the cost of cloud-edge collaborative operation. Therefore, this joint decision generation process provides a basis for subsequent latency assessment, cost assessment, frequency modulation performance evaluation, and reward feedback updates.

[0040] As a refinement of the above embodiment, in step S3, based on the device activation decision output by the diffusion strategy network, the system determines the set of activated devices for the current control cycle. Only activated devices upload status information and generate valid frequency modulation tasks within the current control cycle.

[0041] equipment In the The original frequency modulation task within each control cycle is Then its effective frequency modulation workload is: , in, For equipment In the Effective frequency modulation task volume within a control cycle.

[0042] Therefore, when the equipment When not activated, =0, its effective frequency modulation task is zero; when the device When activated, =1, its effective frequency modulation task quantity is equal to the original frequency modulation task quantity.

[0043] Based on task uninstallation ratio ,equipment Frequency modulation workload allocated to edge processing for: , The amount of frequency modulation tasks allocated to cloud-side processing is : , equipment The accessed edge server is Then the edge server In the The total frequency modulation task received within each control cycle is: , Cloud servers in The total frequency modulation task received within each control cycle is: , Device activation control is used to avoid all devices participating in frequency modulation indiscriminately while meeting the target regulation power requirements, thereby reducing unnecessary status uploads, wireless communication, and collaborative computing overhead. Task offloading execution is used to dynamically allocate frequency modulation-related tasks between the edge side and the cloud side based on the wireless channel status, the available computing power of the edge server, and the available computing power of the cloud server, thereby reducing the overall communication-computing latency within the current control cycle.

[0044] In one implementation, to ensure that the target adjustment power is achievable within the current control cycle, the device activation decision satisfies the following constraints: , in, Indicates equipment In the Available regulation capacity within each control cycle and meets the requirements This constraint is used to ensure that the set of activated devices has the capability to meet the target regulation power requirements.

[0045] As a refinement of the above embodiments, step S4 includes: After completing the device activation and task offloading decisions, the activated device uploads its status information and frequency modulation task to the edge server it is connected to. For tasks offloaded to the edge, the edge server processes them; for tasks offloaded to the cloud, they are first uploaded via the wireless link from the terminal device to the edge server, and then transmitted to the cloud for processing via the backhaul link from the edge server to the cloud server.

[0046] During cloud-edge-device collaborative frequency modulation, participating devices need to upload their local status information and frequency modulation tasks to the edge server they are connected to. Tasks can be processed locally at the edge or forwarded to the cloud via the edge server. Therefore, the overall system response latency is primarily determined by the wireless transmission latency from the terminal to the edge, the backhaul latency from the edge to the cloud, and the cloud-edge computing latency.

[0047] Record equipment To its access edge server The wireless transmission rate is According to Shannon's formula as follows: , in, For channel bandwidth, For equipment Upload transmission power, For noise power, For equipment The equivalent channel gain between it and the edge server it connects to.

[0048] Considering the combined effects of path loss, shadowing fading, and small-scale fading, the channel gain is expressed as: , The large-scale fading term is as follows: , The small-scale fading term is: , In the formula, The channel gain at the reference distance. For equipment Instead of connecting to the edge server The distance between them This is the path loss index. For shadow fading factor, This represents the Rayleigh fading coefficient at a small scale. It is a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1.

[0049] For edge-side processing of frequency modulation tasks, the device only needs to upload the frequency modulation task to its access edge server via a wireless link. Therefore, the upload latency for edge-side frequency modulation tasks is: , For cloud-based FM tasks, the FM task must first be uploaded to the edge server via a wireless link, and then forwarded to the cloud by the edge server via a backhaul link. (Note: The "edge server" is a separate, unrelated term.) The return rate between the cloud server and the cloud server is The upload latency for the cloud frequency modulation task is: , Considering that the data volume of frequency modulation control commands is usually smaller than that of device status information and frequency modulation task calculation, and that downlink control commands can be quickly issued via broadcast or short message, this paper mainly considers the communication-computing integrated latency generated by uplink frequency modulation task transmission and cloud-edge computing processes.

[0050] The number of CPU cycles required for a unit bit frequency modulation task data is Edge server In the The computing power within each control cycle is The cloud server's computing power is Edge server The total frequency modulation task received is Then its calculation delay is: , The total number of frequency modulation tasks received by the cloud server is Its calculation delay is: , This method does not further subdivide the CPU resources within the same edge server at the task level. Instead, it treats tasks received by the same edge server as aggregated tasks for processing. This assumption can reflect the impact of the total task load of edge nodes on computation latency while maintaining the simplicity of the model.

[0051] Edge servers need to complete local computation after receiving the frequency modulation task they are responsible for processing. For edge servers... The arrival delay of its edge-side frequency modulation task can be expressed as: , Therefore, edge servers The completion delay of the frequency modulation task is: , Similarly, the cloud needs to wait for the frequency modulation task offloaded to the cloud to arrive before it can perform calculations. The arrival delay of the frequency modulation task is: The latency for completing the cloud-based frequency modulation task is: , Since the generation of frequency modulation control commands depends on the collaborative processing results of related frequency modulation tasks on the edge side and in the cloud, the system in the first... The communication-computation integrated delay (device dynamic response model) within a control cycle is defined as follows: , Equipment involvement in decision-making changes the effective frequency modulation task scale. and Task unloading decisions are made by changing and The load distribution of frequency modulation tasks between the edge and the cloud affects the overall communication-computing latency.

[0052] In cloud-edge-device collaborative computing, effective frequency modulation tasks generated by participating frequency modulation devices can be processed on edge servers or cloud servers. Due to differences in deployment costs, service capabilities, and billing methods between edge and cloud computing resources, different offloading decisions not only affect communication-computation latency but also further impact system computing costs. To characterize the economic costs of using cloud-edge computing resources, this paper establishes a computing cost model based on the computing processing time of edge and cloud servers: , in, and These represent the edge computing cost and cloud computing cost per unit time, respectively. Represents edge server In the The computational delay within each control cycle Indicates the cloud server is in the The computational delay within each control cycle.

[0053] After cloud-edge collaborative processing is completed, the system generates a frequency modulation control command based on the target adjustment power, device activation result, task processing result, and communication-computation integrated latency, and sends it to the activated device. Upon receiving the frequency modulation control command, the device does not instantly reach the target adjustment power, but rather performs a dynamic power response under the combined effects of communication-computation integrated latency and device ramp-up time.

[0054] equipment In the The target regulation power undertaken within each control cycle is ,equipment The climbing time is ,equipment In the Time within each control cycle The actual output regulation power is The virtual power plant is in the first... The aggregated frequency modulation power (aggregated frequency modulation power model) within each control cycle is: , To evaluate the overall tracking capability of the virtual power plant to the target regulating power during the settlement period, a frequency regulation performance coefficient is introduced. : , in, Indicates the frequency modulation settlement cycle. Indicates the control period. This indicates the number of control cycles contained within a single settlement cycle. This indicates the time of the virtual power plant during the k-th control cycle. Aggregated frequency modulation power, This indicates the target regulating power for this control cycle.

[0055] Substituting the equipment dynamic response model and the aggregated frequency modulation power model into the above equation, we can obtain the expanded expression for the frequency modulation performance coefficient as follows: , in, For equipment In the The target regulation power undertaken within each control cycle This indicates the target regulating power for this control cycle. Indicates the control period. Indicates the frequency modulation settlement cycle. For the first The overall delay within each control cycle For equipment The time spent climbing the hill, This indicates the number of control cycles contained within a single settlement cycle.

[0056] As shown in the above formula, the frequency modulation performance coefficient is jointly determined by the power satisfaction ratio term and the effective response time term. The power satisfaction ratio term characterizes the degree to which the target regulation power undertaken by the activated equipment in the current control cycle meets the target regulation power of that cycle. It complements the normalized power deficit term; that is, the lower the power satisfaction ratio, the larger the corresponding power deficit. The effective response time term characterizes the impact of communication-computation latency and equipment dynamic ramp-up time on the frequency modulation response time. When the power satisfaction ratio decreases, i.e., the power deficit increases, the frequency modulation performance coefficient decreases; similarly, when the communication-computation latency increases or the equipment dynamic ramp-up time lengthens, the frequency modulation performance coefficient also decreases. Therefore, this invention can uniformly map the power deficit impact, latency impact, and dynamic response impact at the control cycle level to the frequency modulation performance index at the settlement cycle level.

[0057] The objectives to be optimized in practice are as follows: In terms of economics, the revenue loss caused by the decline in frequency modulation performance during the settlement period. Recorded as: , in, This represents the frequency regulation revenue that a virtual power plant can obtain within the settlement period under ideal, zero-latency conditions. Simultaneously, it represents the cumulative operating costs generated by cloud-edge collaborative processing within the settlement period. Recorded as: , Therefore, the overall optimization objective in practice of this invention is: , That is, while ensuring the achievability of the target adjustment power and the frequency modulation performance, the time delay impact, power allocation impact and cloud-edge computing cost within the control cycle are uniformly mapped to the overall loss at the settlement cycle level, thereby achieving joint optimization of device activation decision and task offloading decision.

[0058] The aforementioned overall loss is used to characterize the frequency modulation performance loss and cloud-edge collaborative operation cost at the settlement cycle level. During reinforcement learning training, it can be decomposed into communication-computation integrated latency, target regulation power deficit, and cloud-edge operation cost at the control cycle level, and used as a component of the immediate reward.

[0059] The instant reward function is as follows: After the current frequency modulation control cycle ends, the system calculates an immediate reward based on the execution results of the joint decision-making actions, and updates the parameters of the diffusion strategy network based on the immediate reward. The reward function is used to feed back the frequency modulation effect and cloud-edge collaborative operation overhead within the current control cycle to the strategy network, enabling the diffusion strategy network to learn better device participation and task offloading decisions during continuous frequency modulation control.

[0060] The optimization objective of this invention is to reduce the cost of cloud-edge collaborative operation while ensuring the frequency regulation effect of the virtual power plant. Therefore, during the reinforcement learning training process, the frequency regulation performance coefficient and the cloud-edge computing cost are both included as components of the reward function. Specifically, the first... The instant reward for each control cycle is defined as follows: , in, For instant reward function, For frequency modulation performance coefficient, Indicates the first The frequency regulation performance coefficient within each control cycle is used to characterize the degree to which the virtual power plant meets the target regulation power and the dynamic response effect of the equipment. The weighting coefficients corresponding to the cloud-edge computing cost. Indicates the first The collaborative operation cost within each control cycle. Through the above reward design, the diffusion policy network can simultaneously improve frequency modulation performance and reduce cloud-edge collaborative computing costs while maximizing immediate rewards.

[0061] The optimization objective of the diffusion strategy network is to maximize the long-term cumulative discount reward. Let the policy parameters be... The discount factor is The long-term optimization objective is: , The system observes the status during each control cycle. Joint actions are generated by the diffusion strategy network. The environment performs the joint action and returns to the next state. and instant rewards Subsequently, the network parameters of the diffusion strategy are updated based on reward feedback, thereby continuously optimizing device activation decisions and task unloading decisions in subsequent control cycles.

[0062] Through the above steps, the present invention achieves closed-loop optimization of device activation, task unloading, cloud-edge collaborative processing, frequency modulation dynamic response, and reward feedback update.

[0063] Example 2 In one implementation, such as Figure 2 As shown, a collaborative optimization device for equipment activation and task offloading for frequency regulation in virtual power plants can be deployed as a control terminal, edge control node, frequency regulation decision server, or hardware and software entity with equivalent computing, communication, and control functions on the frequency regulation control side of a virtual power plant. The device includes a processor, a memory, a communication interface, and a power management unit. The memory stores program instructions that can be executed by the processor. When the processor executes the program instructions, it performs functions such as system state vector construction, diffusion strategy network inference, device activation decision generation, task offloading ratio calculation, communication-computation integrated latency assessment, frequency modulation control instruction generation, reward calculation, and diffusion strategy network parameter update.

[0064] The communication interface is used to establish data connections with the terminal FM device, the edge server, and the cloud server, respectively. Through this communication interface, the device can receive status information uploaded by the terminal FM device, receive task processing results and resource status information fed back by the edge server and the cloud server, issue FM control commands to the activated device, and send task unloading and collaborative processing control information to the edge server and the cloud server.

[0065] In terms of functional structure, the device includes a state awareness and modeling module, a joint decision generation module, an activation and unloading execution module, a cloud-edge collaborative processing module, a frequency regulation control and performance evaluation module, a reward calculation and strategy update module, and a communication interaction module. Each module works collaboratively according to the frequency regulation control cycle to complete the frequency regulation optimization control of the virtual power plant under the cloud-edge-device collaborative architecture. This includes: (1) State awareness and construction module This module is used to collect system operation information within the current frequency modulation control cycle and construct a system state vector based on the collected results. The system operation information includes at least the target regulation power, available regulation capability of the equipment, equipment workload, wireless channel status, equipment ramp-up time, available computing power of the edge server, and available computing power of the cloud server.

[0066] The system state vector output by this module is used to characterize the frequency regulation requirements, equipment capabilities, communication status, and cloud-edge resource status within the current control cycle, and serves as the input to the joint decision generation module.

[0067] (2) Joint Decision Generation Module This module is used to generate joint decision actions for the current control cycle based on the diffusion policy network. The joint decision actions include device activation decisions and task unloading decisions.

[0068] Among them, the device activation decision is used to determine the set of target devices participating in the frequency modulation service of the current control cycle; the task offloading decision is used to determine the allocation ratio of frequency modulation-related tasks generated by the activated devices between the edge side and the cloud side.

[0069] This module generates and maps the potential representation of joint actions through a diffusion policy network, thereby achieving unified modeling of discrete device activation actions and continuous task unloading actions.

[0070] (3) Activating and Unloading the Execution Module This module is used to determine the set of activated devices in the current control cycle based on the device activation decision output by the joint decision generation module, and to control the activated devices to upload status information and frequency modulation related tasks.

[0071] Simultaneously, based on the task offloading ratio output by the joint decision generation module, this module allocates effective tasks generated by activated devices to the edge and cloud sides. For inactive devices, this module does not assign frequency modulation-related tasks, thereby reducing unnecessary status uploads, communication transmissions, and computational processing overhead.

[0072] This module enables the device to control the number of devices participating in frequency modulation and the task allocation structure while meeting the target regulation power requirements, thereby improving the real-time performance and resource utilization efficiency of the cloud-edge-device collaborative frequency modulation process.

[0073] (4) Cloud-edge collaborative processing module This module controls the edge server and cloud server to perform collaborative task processing. Specifically, this module is used to hand over tasks offloaded to the edge side to the edge server for processing, to send tasks offloaded to the cloud side back to the cloud server for processing via the edge server, and to receive the task processing results returned by the edge server and the cloud server.

[0074] This module is also used to determine the edge-side completion latency, cloud-side completion latency, and combined communication-computing latency within the current control period based on the task upload process, edge-side processing process, and cloud-side processing process. Furthermore, this module can calculate the cloud-edge collaborative operation cost within the current control period based on the task processing volume, computing power consumption, energy consumption, or resource usage price of the edge server and cloud server.

[0075] (5) Frequency modulation control and performance evaluation module This module is used to generate frequency modulation control commands based on the target adjustment power, device activation results, cloud-edge collaborative processing results, and communication-computing integrated latency, and then send the control commands to the activated device through the communication interaction module.

[0076] This module is also used to calculate the aggregated frequency regulation power of the virtual power plant in the current control cycle based on the dynamic power response process of the activated equipment, and to further evaluate the frequency regulation performance. The frequency regulation performance evaluation considers at least the impact of target regulation power satisfaction, communication-computation integrated latency, and equipment ramp-up time on the actual frequency regulation response.

[0077] Through this module, the device can transform equipment activation results, task unloading results, and cloud-edge processing results into actual frequency regulation control behavior, and evaluate the impact of this control behavior on the frequency regulation performance of the virtual power plant.

[0078] (6) Reward Calculation and Strategy Update Module This module is used to calculate real-time rewards based on the communication-computation integrated latency, target adjustment power satisfaction, aggregated frequency modulation response results, and cloud-edge collaborative operation costs within the current control cycle.

[0079] This module is also used to update the diffusion strategy network parameters based on immediate rewards and state transition results, enabling the joint decision generation module to output better device activation and task unloading decisions in subsequent control cycles.

[0080] Through this module, the device can form a closed-loop learning mechanism based on execution results, enabling the device activation strategy and task offloading strategy to be continuously optimized as the virtual power plant's operating status, communication environment, and cloud-edge resource status change.

[0081] As a refinement of the above embodiments, a communication interaction module is also included. The communication interaction module runs through all the above method steps and is used to realize bidirectional data communication between the device and the terminal frequency modulation device, the edge server, and the cloud server.

[0082] Specifically, the communication interaction module is used to receive device status information, available adjustment capabilities, task load and channel-related information uploaded by the terminal frequency modulation device; to receive available computing power, task processing results and running status information returned by the edge server and cloud server; to send task unloading control information to the edge server and cloud server; and to issue frequency modulation control commands to the activated device.

[0083] Through the communication interaction module, the device can form a complete data interaction link with nodes at the cloud, edge, and terminal layers, thereby supporting the closed-loop operation of system state perception, joint decision generation, cloud-edge collaborative processing, frequency modulation control execution, and reward feedback update.

[0084] The device operates according to the following process in each frequency modulation control cycle.

[0085] First, the state perception and modeling module collects the target regulation power, available regulation capability of the equipment, equipment workload, wireless channel status, equipment ramp-up time, available computing power of the edge server and available computing power of the cloud server within the current control cycle through the communication interaction module, and constructs the system state vector based on the collection results.

[0086] Secondly, the joint decision generation module inputs the system state vector into the diffusion strategy network, which generates joint decision actions for the current control cycle. These joint decision actions include device activation decisions and task unloading decisions.

[0087] Then, the activation and unloading execution module determines the set of activated devices for the current control cycle based on the device activation decision, and determines the allocation ratio of the effective tasks of the activated devices between the edge side and the cloud side based on the task unloading decision.

[0088] Next, the cloud-edge collaborative processing module controls the activated device to upload status information and frequency modulation-related tasks through the communication interaction module, and controls the edge server and cloud server to perform collaborative processing according to the task unloading results. Based on the task upload and cloud-edge computing processing results, this module determines the combined communication-computing latency and cloud-edge collaborative operation cost.

[0089] Subsequently, the frequency regulation control and performance evaluation module generates frequency regulation control commands based on the target regulation power, equipment activation results, task processing results, and communication-computation integrated time delay. It then sends control commands to the activated equipment through the communication interaction module, enabling the activated equipment to perform dynamic power response, forming virtual power plant aggregated frequency regulation power, and further evaluating the frequency regulation performance within the current control cycle and settlement cycle.

[0090] Finally, the reward calculation and strategy update module calculates real-time rewards based on the communication-computation integrated latency, target adjustment power satisfaction, aggregated frequency modulation response results, and cloud-edge collaborative operation costs. Based on the real-time rewards, the module updates the diffusion strategy network parameters, enabling the device to continuously optimize device activation decisions and task offloading decisions in subsequent control cycles.

[0091] Through the above-described working process, the device of the present invention can achieve integrated closed-loop optimization of equipment activation, cloud-edge task unloading, frequency regulation control execution, and strategy feedback update during the frequency regulation process of a virtual power plant in dynamic communication and computing environments.

[0092] This disclosure also provides a device for coordinated optimization of equipment activation and task offloading for frequency regulation in virtual power plants, including a processor and a memory. Optionally, the device may further include a communication interface and a bus. The processor, communication interface, and memory can communicate with each other via the bus. The communication interface can be used for information transmission. The processor can call logical instructions in the memory to execute the device activation and task offloading coordinated optimization method for frequency regulation in virtual power plants described in the above embodiments.

[0093] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0094] Memory, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor executes functional applications and data processing by running the program instructions / modules stored in the memory, thereby realizing the equipment activation and task offloading collaborative optimization method for virtual power plant frequency regulation described in the above embodiments.

[0095] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory may include high-speed random access memory and may also include non-volatile memory.

[0096] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described collaborative optimization method for device activation and task offloading in virtual power plant frequency regulation.

[0097] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0098] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code. It can also be a transient storage medium.

[0099] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A collaborative optimization method for equipment activation and task offloading in virtual power plant frequency regulation, characterized in that, The method is executed based on a cloud-edge-device collaborative architecture and includes the following steps: S1: Collect system operation information within the current frequency modulation control cycle and construct the system state vector; S2: Input the system state vector into the trained diffusion strategy network, and the diffusion strategy network generates a joint decision action for the current control cycle. The joint decision action includes a device activation decision and a task offloading decision. The device activation decision is used to determine the set of target devices participating in frequency modulation, and the task offloading decision is used to determine the allocation ratio of frequency modulation-related tasks generated by the activated devices between the edge server and the cloud server. S3: Based on the device activation decision, control the activated device to upload status information and frequency modulation tasks, and based on the task unloading decision, allocate the effective frequency modulation task volume to the edge server and cloud server for collaborative processing; S4: Calculate the comprehensive latency and collaborative operation cost based on the effective task volume of edge servers and cloud servers, establish an instant reward function, and update the parameters of the diffusion strategy network based on the instant reward. S5: Execute S1-S4 repeatedly to continuously optimize equipment activation decisions and task unloading decisions until the instantaneous reward function converges, obtain the optimal diffusion strategy network, collect the operation information of the virtual power plant dispatch system to be optimized, and obtain the optimization strategy through the optimal diffusion strategy network.

2. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 1, characterized in that, The system operation information includes target adjustment power, maximum available adjustment capability of each frequency modulation device, wireless channel status, device workload, device ramp-up time, available computing power of edge servers, and available computing power of cloud servers.

3. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 1, characterized in that, The joint decision action for the current control cycle is generated by the diffusion strategy network, specifically including: Using a diffusion policy network with the system state vector as a condition, a joint action latent representation is generated through a reverse denoising process; The first N-dimensional variables of the potential representation of the joint action are transformed by the Sigmoid function and thresholded to obtain discrete device activation results; The last N-dimensional variables of the potential representation of the joint action are mapped to the [0,1] interval using the Sigmoid function to obtain the continuous task unloading ratio; Where N is the total number of adjustable devices in the virtual power plant.

4. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 1, characterized in that, The specific methods for obtaining the effective frequency modulation task quantity include: For activated devices, multiply their original task volume by the device activation decision to obtain the effective frequency modulation task volume; The effective frequency modulation task volume is obtained by calculating the effective frequency modulation task volume processed by the edge server based on the task unloading ratio in the task unloading decision. The effective frequency modulation task volume processed by the cloud server is obtained by subtracting the effective frequency modulation task volume processed by the edge server from the effective task volume.

5. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 1, characterized in that, The instant reward function is as follows: , in, For instant reward function, For frequency modulation performance coefficient, Indicates the first Frequency modulation performance coefficient within each control cycle The weighting coefficients corresponding to the cloud-edge computing cost. Indicates the first The cost of coordinated operation within a control cycle; The long-term optimization objective is: , in, For strategy The corresponding long-term expected cumulative discount reward function, Indicates the strategy The generated state-action trajectory is expected. As a discount factor, These are strategy parameters.

6. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 5, characterized in that, The frequency modulation performance coefficient is calculated as follows: , in, For equipment In the The target regulation power undertaken within each control cycle This indicates the target regulating power for this control cycle. Indicates the control period. Indicates the frequency modulation settlement cycle. For the first The overall delay within each control cycle For equipment The time spent climbing the hill, For the first A set of activated devices for each control cycle. Indicates the number of control cycles contained within a single settlement cycle; The method for calculating the collaborative operation cost is as follows: , in, This represents the cost of edge server computation per unit of time. Represents edge server In the The computational delay within each control cycle This represents the cost of cloud server computing per unit time. Indicates the cloud server is in the The computational delay within each control cycle This indicates the number of edge servers.

7. The collaborative optimization method for equipment activation and task offloading for frequency regulation in virtual power plants according to claim 6, characterized in that, The comprehensive delay calculation method is as follows: , , , , , , , , in, For edge servers In the Task completion delay within each control cycle For cloud servers in the first Task completion delay within each control cycle Calculate the latency based on the total number of tasks received by the cloud server. For cloud server task arrival latency, For cloud server task upload latency, For equipment In the The amount of tasks allocated to the cloud server within a control cycle. For equipment The wireless transmission rate to its access edge server, For edge servers The return speed between the cloud server and the server For edge server task arrival latency, For edge servers The calculation delay for the total amount of received tasks. For the number of edge servers, For edge server task upload latency, For equipment In the The amount of tasks allocated to the edge server within a control cycle The number of CPU cycles required per unit bit of task data. For edge servers Total number of tasks received For edge servers In the Computing capacity within a control cycle The number of devices can be adjusted. Calculate the latency based on the total number of tasks received by the cloud server. This represents the total number of tasks received by the cloud server. For cloud servers in the first Computing capacity within a control cycle.

8. A collaborative optimization device for equipment activation and task offloading in virtual power plant frequency regulation, characterized in that, include: The state awareness and modeling module is used to collect system operation information within the current frequency regulation control cycle and construct the system state vector. The joint decision generation module is used to input the system state vector into the diffusion policy network and generate joint decision actions that include device activation decisions and task unloading decisions; The activation and unloading execution module is used to determine the set of devices to be activated based on the device activation decision, and to allocate valid tasks to the edge side and cloud side based on the task unloading decision; The cloud-edge collaborative processing module is used to control the edge server and cloud server to perform task processing and determine the overall latency and operating cost; The frequency modulation control and performance evaluation module is used to generate and issue frequency modulation control commands, and at the same time evaluate the aggregated frequency modulation response results. The reward calculation and strategy update module is used to calculate real-time rewards based on latency, power satisfaction, response results, and cost, and update the diffusion strategy network parameters.

9. A device for collaborative optimization of equipment activation and task offloading for frequency regulation in virtual power plants, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute, when running the program instructions, the collaborative optimization method for device activation and task offloading for virtual power plant frequency regulation as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the collaborative optimization method for equipment activation and task offloading for frequency regulation of virtual power plants as described in any one of claims 1-7.