A system and method for trusted collaborative operation optimization of a virtual power plant
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]为了克服上述虚拟电厂调用能力不足、运营效率低的问题,本发明提供一种虚拟电厂的可信协同运行优化系统方法
本发明提供一种虚拟电厂的可信协同运行优化系统和方法,通过联邦学习子系统基于本地状态数据以隐私保护方式进行协同模型训练,使得各分布式节点无需上传原始数据即可参与模型共建。从根本上消除了数据集中带来的隐私泄露风险,破解了用户参与的核心顾虑,解决用户隐私泄露的问题。通过区块链存证子系统对模型训练贡献信息、调度指令集及执行结果进行链上存证并基于存证信息自动执行激励结算,构建了一个不可篡改、可追溯的客观记录与自动执行体系。该特征将原本模糊的贡献与行为转化为链上可验证、可编程的数据,使得激励结算过程透明、公平、自动,从而有效激励用户贡献数据与调节能力。
Smart Images

Figure CN122534084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual power plant technology, and more specifically to a trusted collaborative operation optimization system and method for virtual power plants. Background Technology
[0002] Virtual power plants are a key technology for aggregating massive, heterogeneous, and dispersed distributed energy resources. By aggregating various distributed resources such as distributed power sources, adjustable loads, and energy storage, they serve as a new type of operating entity to collaboratively participate in power system optimization and power market transactions, forming a power operation organization model.
[0003] In actual operation, virtual power plants currently require owners of distributed resources to upload detailed operational data, raising serious concerns about privacy leaks and causing many potential users to refuse to participate. Furthermore, the inability to quantify users' regulatory behavior and data contributions leads to disputes over incentive settlement, further dampening the enthusiasm of high-value users. This results in insufficient virtual power plant call capacity and low operational efficiency. Summary of the Invention
[0004] To overcome the problems of insufficient virtual power plant call capacity and low operational efficiency, this invention provides a reliable collaborative operation optimization system method for virtual power plants.
[0005] On one hand, the present invention provides a trusted collaborative operation optimization system for a virtual power plant, comprising: The federated learning subsystem is used to train a collaborative model in a privacy-preserving manner based on the local state data of each distributed node in the virtual power plant, and obtain a global prediction model. A multi-agent cooperative decision-making subsystem, which is communicatively connected to the federated learning subsystem, is used to generate a cooperative scheduling instruction set based on the power prediction data obtained by the global prediction model, so that each distributed node executes the corresponding scheduling instruction; The blockchain evidence storage subsystem is communicatively connected to both the federated learning subsystem and the multi-agent collaborative decision-making subsystem. It receives contribution information from the federated learning subsystem during the collaborative model training process, the collaborative scheduling instruction set from the multi-agent collaborative decision-making subsystem, and the instruction execution results from each distributed node, and stores these on-chain. Based on the stored evidence information, it automatically executes incentive settlement and verifies the execution deviations of the scheduling instructions. In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.
[0006] Optionally, the federated learning subsystem includes: Multiple edge clients, each deployed on a distributed energy resource side, are used to perform privacy processing on local state data locally and train local prediction models based on the privacy-processed data; The cloud server is used to receive model parameter update information uploaded by each edge client, and to update the global prediction model by fusing the update information of each model parameter. The updated global prediction model is then distributed to each edge client. The privacy processing involves adding differential privacy noise to the local state data; one distributed energy resource corresponds to one distributed node.
[0007] Optionally, the cloud server is specifically used for: The contribution weight of each local prediction model is determined based on the similarity between the local prediction model and the global prediction model of each edge client. Based on the contribution weights and model parameter update information of each local prediction model, parameter fusion is performed to obtain fused model parameters, and the global prediction model is updated using the fused model parameters.
[0008] Optionally, the multi-agent cooperative decision-making subsystem includes: A central network of commentators, deployed on cloud servers, is used to evaluate the value of joint actions of individual edge agents in joint states during the training phase. Multiple policy networks, each deployed on an edge agent corresponding to a distributed energy resource, are used for: during the training phase, generating scheduling instructions based on local observation states and historical action statistics of other agents, and updating the parameters of the policy networks based on the value assessment output by the central critic network; during the execution phase, independently running the corresponding policy network based on the local observation states of each edge agent to generate distributed scheduling instructions. The edge agent corresponds one-to-one with the distributed node.
[0009] Optionally, the blockchain evidence storage subsystem further includes a smart contract module, which is used for: The consistency between the scheduling instructions and the actual execution results of the scheduling instructions is verified. Based on the verification results and the pre-stored contribution weight list, the corresponding incentive items are calculated and the corresponding incentive vouchers are issued to each edge client.
[0010] Optionally, the reward function corresponding to the scheduling instruction output by the policy network includes an incentive term calculated based on blockchain-stored information; the incentive term is calculated and distributed through a smart contract in the blockchain-stored information subsystem.
[0011] Optionally, the reward function corresponding to the scheduling instruction output by the policy network is:
[0012] in, For edge intelligent agents During the period The reward function value, For edge intelligent agents During the period Electricity / maintenance costs; For virtual power plants in different time periods The power exchanged with the power grid, and the target value of that power exchange; For virtual power plants in time periods The local consumption rate of renewable energy; Issuance to edge intelligent agents by the blockchain evidence storage subsystem During the period Incentive items; These are the weighting coefficients for each reward; Edge agents The incentives obtained are:
[0013] in, This serves as the total incentive pool for this round of model updates; Let be the excitation function. For edge intelligent agents During the period The actual execution power, For edge intelligent agents During the period The power adjustment amount corresponding to the scheduling command; For edge intelligent agents During the period The verification result is 1 if the verification is successful, and 0 otherwise. For edge intelligent agents Long-term reputation score, based on edge intelligent agents Historical contributions and performance records are dynamically updated; , Edge agents Weighting coefficients for each incentive component.
[0014] Optionally, it also includes: The digital twin visualization subsystem is used to synchronize and display the global prediction model, the collaborative scheduling instruction set, on-chain evidence information and incentive settlement results in real time, and to trace the decision-making process of the specified scheduling instruction based on the real-time synchronized data.
[0015] Optionally, the digital twin visualization subsystem includes: A digital twin construction module is used to construct a virtual model that is synchronized with the virtual power plant in real time. The strategy inference module is used to simulate and compare the effects of scheduling strategies in different future scenarios based on a trained multi-agent reinforcement learning model. The decision tracing module is used to respond to query requests and display the federated learning prediction results on which the specified historical scheduling instruction depends, the decision logic of the corresponding agent when the scheduling instruction was generated, and the full lifecycle evidence record of the scheduling instruction on the blockchain.
[0016] On the other hand, the present invention also provides a trusted collaborative operation optimization method for virtual power plants, comprising: In a virtual power plant, each distributed node trains a collaborative model based on its own local state data in a privacy-preserving manner to obtain a global prediction model. Each distributed node generates a set of cooperative scheduling instructions and executes the corresponding scheduling instructions based on the power prediction data obtained by the global prediction model through multi-agent reinforcement learning. Each distributed node uploads its contribution information, collaborative scheduling instruction set, and instruction execution results to the blockchain network for on-chain storage. The blockchain network automatically executes incentive settlement based on the stored information and verifies the execution deviation of the scheduling instructions; In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.
[0017] Optionally, in the virtual power plant, each distributed node trains a collaborative model based on its own local state data in a privacy-preserving manner to obtain a global prediction model, including: Each distributed node performs privacy processing on its local state data and trains a local prediction model based on the privacy-processed data. The cloud server receives model parameter update information uploaded by each distributed node, updates the global prediction model by fusing the update information of each model parameter, and distributes the updated global prediction model to each distributed node. The privacy processing involves adding differential privacy noise to the local state data; one distributed energy resource corresponds to one distributed node.
[0018] Optionally, the step of updating the global prediction model through the fusion processing of update information for each model parameter includes: The contribution weight of each local prediction model is determined based on the similarity between the local prediction model and the global prediction model. Based on the contribution weights and model parameter update information of each local prediction model, parameter fusion is performed to obtain fused model parameters, and the global prediction model is updated using the fused model parameters.
[0019] Optionally, each distributed node generates a cooperative scheduling instruction set based on the power prediction data obtained through the global prediction model, using multi-agent reinforcement learning, including: The cloud server utilizes a central commentator network to evaluate the value of joint actions of each edge agent in a joint state during the training phase; During the training phase, each edge agent utilizes its local policy network to generate scheduling instructions based on its local observation state and the historical action statistics of other agents, and updates the parameters of the policy network based on the value assessment output by the central critic network. During the execution phase, each edge agent independently runs its corresponding policy network based on its local observation state to generate distributed scheduling instructions. The edge agent corresponds one-to-one with the distributed node.
[0020] Optionally, the blockchain network automatically executes incentive settlement based on the stored evidence information and verifies the execution deviation of the scheduling instructions, including: The blockchain network verifies the consistency between scheduling instructions and their actual execution results. Based on the verification results and the pre-stored list of contribution weights, it calculates the corresponding incentive items and issues the corresponding incentive certificates to each distributed node.
[0021] Optionally, the reward function corresponding to the scheduling instruction output by the strategy network includes an incentive term calculated based on blockchain-stored information; the incentive term is calculated and distributed through smart contracts in the blockchain network.
[0022] Optionally, the reward function corresponding to the scheduling instruction output by the policy network is:
[0023] in, For edge intelligent agents During the period The reward function value, For edge intelligent agents During the period Electricity / maintenance costs; For virtual power plants in different time periods The power exchanged with the power grid, and the target value of that power exchange; For virtual power plants in time periods The local consumption rate of renewable energy; Issuance to edge intelligent agents by the blockchain evidence storage subsystem During the period Incentive items; These are the weighting coefficients for each reward; Edge agents The incentives obtained are:
[0024] in, This serves as the total incentive pool for this round of model updates; Let be the excitation function. For edge intelligent agents During the period The actual execution power, For edge intelligent agents During the period The power adjustment amount corresponding to the scheduling command; For edge intelligent agents During the period The verification result is 1 if the verification is successful, and 0 otherwise. For edge intelligent agents Long-term reputation score, based on edge intelligent agents Historical contributions and performance records are dynamically updated; , Edge agents Weighting coefficients for each incentive component.
[0025] Optionally, it also includes: The global prediction model, the collaborative scheduling instruction set, on-chain evidence storage information, and incentive settlement results are synchronized and displayed in real time through the digital twin visualization subsystem, and the decision-making process of the specified scheduling instruction is traced based on the real-time synchronized data.
[0026] Optionally, the process of tracing the decision-making process for a specified scheduling instruction based on real-time synchronized data includes: Construct a virtual model that is synchronized with the virtual power plant in real time; Based on a trained multi-agent reinforcement learning model, the scheduling strategies in different future scenarios are simulated and their effects are compared. In response to a query request, the system displays the federated learning prediction results upon which the specified historical scheduling instruction is based, the decision logic of the corresponding agent when the scheduling instruction was generated, and the full lifecycle record of the scheduling instruction on the blockchain.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a trusted collaborative operation optimization system and method for virtual power plants. Through a federated learning subsystem, collaborative model training is performed based on local state data in a privacy-preserving manner, allowing distributed nodes to participate in model co-construction without uploading raw data. This fundamentally eliminates the privacy leakage risks associated with data centralization, addresses core user concerns, and solves the problem of user privacy breaches. A blockchain notarization subsystem stores model training contribution information, scheduling instruction sets, and execution results on-chain, and automatically executes incentive settlement based on the notarized information, constructing an immutable, traceable, objective record and automatic execution system. This feature transforms previously ambiguous contributions and behaviors into verifiable, programmable on-chain data, making the incentive settlement process transparent, fair, and automatic, thereby effectively incentivizing users to contribute data and adjust their capabilities.
[0028] This invention enhances the scientific rigor and foresight of decision-making by employing a multi-agent collaborative decision-making subsystem based on a high-precision global prediction model and generating a collaborative scheduling instruction set through reinforcement learning. By verifying the execution deviations of scheduling instructions across distributed nodes, reliable oversight of instruction execution is strengthened, ensuring the accurate implementation of optimization intentions. The resulting state transition data is used for online optimization of the local policy network, enabling the system to continuously learn and improve from actual operation, dynamically enhancing overall operational efficiency and invocation capabilities. Attached Figure Description
[0029] Figure 1 This is a structural block diagram of a trusted collaborative operation optimization system for a virtual power plant, as an example of the present invention. Figure 2 This is a system architecture diagram of a trusted collaborative operation optimization system for a virtual power plant, as an example of the present invention. Figure 3 This is a schematic diagram illustrating the key event recording and automatic incentive settlement process of a blockchain evidence storage subsystem, as an example of the present invention. Figure 4 A schematic diagram illustrating the process of recording model contributions to a blockchain in a federated learning subsystem, as an example of the present invention. Figure 5 A schematic diagram of the visualization monitoring and decision tracing interface of a digital twin visualization subsystem, as an example of the present invention; Figure 6 This is a diagram illustrating the system workflow and data interaction of an example of the present invention. Detailed Implementation
[0030] The current development of Virtual Power Plants (VPPs) faces the following challenges: Existing centralized aggregation models require participants to upload all raw data, infringing on user privacy and resulting in low participation. Traditional optimization methods are mostly offline static optimizations based on historical data, unable to respond in real-time to market price fluctuations, the randomness of renewable energy output and load, leading to lagging optimization results. Deep learning-based scheduling models lack interpretability; their decision-making logic is opaque, making it difficult to gain the trust of grid dispatching agencies and participants, and failing to meet the high reliability requirements of power systems. The lack of fair, transparent, and automated value contribution measurement and benefit distribution mechanisms makes it difficult to incentivize rational participants to contribute data and regulation capabilities, hindering the achievement of global optimization goals. The issuance, execution, and feedback processes of dispatching instructions lack tamper-proof and traceable records, making it impossible to define responsibility and effectively verify deviations in instruction execution.
[0031] While federated learning technology can collaboratively train models without sharing raw data, it can only perform prediction tasks and cannot make sequential decisions; moreover, it lacks incentive mechanisms to counteract speculative behavior by rational participants. Therefore, a comprehensive technical solution is urgently needed that can simultaneously guarantee privacy, achieve dynamic decision-making, establish credible incentives, and provide decision explanations. To address these issues, this invention proposes a trusted collaborative operation optimization system and method for virtual power plants. It involves a collaborative optimization system and method for virtual power plants integrating federated learning (FL), multi-agent deep reinforcement learning (MADRL), blockchain, and digital twins. This aims to solve the four core challenges in distributed energy aggregation: data privacy, collaborative decision-making, lack of trust, and decision-making black boxes, and to build a secure, efficient, trustworthy, and transparent next-generation intelligent operation platform for virtual power plants.
[0032] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0033] Example 1 The present invention provides a trusted collaborative operation optimization system for a virtual power plant, the schematic diagram of which is shown below. Figure 1 As shown, it includes: Federated learning subsystem 110 is used to train a collaborative model in a privacy-preserving manner based on the local state data of each distributed node in the virtual power plant to obtain a global prediction model. The multi-agent cooperative decision-making subsystem 120 is communicatively connected to the federated learning subsystem 110 and is used to generate a cooperative scheduling instruction set based on the power prediction data obtained by the global prediction model through multi-agent reinforcement learning so that each distributed node executes the corresponding scheduling instruction. The blockchain evidence storage subsystem 130 is communicatively connected to the federated learning subsystem 110 and the multi-agent collaborative decision-making subsystem 120, respectively; it is used to receive contribution information from the federated learning subsystem 110 during the collaborative model training process, the collaborative scheduling instruction set from the multi-agent collaborative decision-making subsystem 120, and the instruction execution results from each distributed node, and to store them on the blockchain; it automatically executes incentive settlement based on the evidence storage information; and it verifies the execution deviation of the scheduling instructions. In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.
[0034] In this example implementation, the system adopts a layered architecture of cloud-edge-device collaboration, as shown in the attached diagram. Figure 2As shown, the system comprises a physical layer, an edge agent layer, a platform layer, and an application layer. The physical layer includes various distributed energy resources (PV, WT, ES, CL, etc.). The edge agent layer equips each distributed energy resource (DER) with an edge agent, acting as a client of the system, integrating local model training, policy execution, lightweight data processing, and blockchain lightweight node functionality. The platform layer includes: a federated learning subsystem 110 primarily responsible for coordinating cross-client model training and secure aggregation, which may include a hierarchical aggregation server; a multi-agent collaborative decision-making subsystem 120, acting as the system's intelligent brain, generating globally optimal scheduling strategies based on MADRL, which may include a MADRL decision engine and an environment simulator; and a blockchain evidence storage subsystem 130 providing immutable log recording and automatic incentive settlement based on smart contracts, which may encompass a consortium blockchain network (Fabric), smart contracts, and incentive ledgers. The platform layer can also include a digital twin and visualization subsystem for building virtual images of the physical VPP, providing monitoring, simulation, and decision interpretation interfaces (such as a 3D visualization engine, strategy simulation sandbox, and decision traceability dashboard). The application layer includes various business applications for operators, participants, and regulators (such as market trading, demand response, and carbon asset management). The federated learning subsystem 110 addresses privacy and security concerns arising from data centralization in virtual power plant operations. It allows distributed nodes (e.g., agents deployed on photovoltaic, energy storage, and adjustable load sides) to participate in model training without transmitting raw local state data (such as real-time power, equipment status, and meteorological information). Specifically, each node uses local state data for computation, uploading only the update amount of model parameters (not the data itself). The server aggregates the received parameter updates to generate a high-quality global prediction. The model. This subsystem acts as the system's perceptual nerve, providing a holistic understanding of key factors such as renewable energy output and load demand within the virtual power plant without exposing raw data, thus providing a data foundation for subsequent decision-making. The multi-agent collaborative decision-making subsystem 120 receives global prediction data from the federated learning subsystem 110 and uses it as important environmental state input. Based on this, each distributed node within the virtual power plant is treated as an autonomous agent, constructing a multi-agent reinforcement learning framework. Each agent (i.e., distributed node) calculates and proposes scheduling action suggestions (such as adjusting power generation or changing charging / discharging states) through its internal policy network based on its locally observed specific state (combined with global prediction information). The server is responsible for coordinating these decentralized decisions, evaluating their joint effects, and ultimately generating a consistent set of collaborative scheduling instructions, which is then distributed to each node for execution.This process represents a leap from static, offline optimization to dynamic, online, and adaptive decision-making, enabling the system to respond in real time to complex dynamic environments such as electricity price fluctuations and random changes in renewable energy output, thereby significantly improving the overall operational efficiency and economic benefits of the virtual power plant. The blockchain evidence storage subsystem 130 is responsible for the immutable recording of key events during system operation. This includes recording the contribution information of each node in federated learning (quantifying its data value), recording each set of scheduling instructions issued by the multi-agent decision-making subsystem 120, and recording the final actual execution results of each node. All records are stored on the blockchain in the form of hash values, forming a traceable chain of evidence. Based on this on-chain evidence information, the smart contract automatically triggers the execution incentive settlement. The smart contract calculates the incentive that each node should receive based on its model contribution and the execution of scheduling instructions (especially the verification results of execution deviations) according to preset, transparent rules, and automatically distributes the incentives. This mechanism directly, transparently, and reliably links contribution and reward, effectively incentivizing nodes to actively participate and faithfully execute instructions.
[0035] It should be noted that after executing a scheduling instruction, each distributed node collects "state transition data" for this action. This includes the state before execution (containing local privacy-preserving observation data and historical action statistics of other agents), the action executed (DER scheduling instructions, such as power generation / charging power), the immediate feedback (rewards), and the new state after execution. This data is used to optimize and fine-tune the node's local policy network online, enabling the decision network to continuously learn and evolve to adapt to new scenarios or environmental changes.
[0036] In some implementations, the federated learning subsystem includes: Multiple edge clients, each deployed on a distributed energy resource side, are used to perform privacy processing on local state data locally and train local prediction models based on the privacy-processed data; The cloud server is used to receive model parameter update information uploaded by each edge client, and to update the global prediction model by fusing the update information of each model parameter. The updated global prediction model is then distributed to each edge client. The privacy processing involves adding differential privacy noise to the local state data; one distributed energy resource corresponds to one distributed node.
[0037] In this example implementation, each edge client is specifically deployed on a particular distributed energy resource side, such as the inverter controller of a photovoltaic power station, the energy management system of an energy storage system, or the automation control unit of a building. Each client first performs privacy processing on its local state data (such as power, electricity price, weather data, etc.) through differential privacy perturbation, that is, in the local state data... Before uploading to any computation process, differential privacy perturbation is performed to generate perturbed data:
[0038] In the formula, For edge clients The original high-dimensional data vector (such as power, electricity price, and meteorological data). To The data after adding Laplace noise, i.e., the data after privacy processing; For query functions The sensitivity (usually taken as the maximum absolute value of the data change). A privacy budget is allocated to control the strength of privacy protection; a smaller value indicates stronger privacy protection. Each client then independently trains a local prediction model using the privacy-enhanced data. This model focuses on learning the operational patterns of its local resource point (e.g., the relationship between photovoltaic output and sunlight and temperature). After training, the client sends the updated parameters of its local model to the cloud server. The cloud server acts as a coordinator and integrator. It receives model parameter updates from all participating clients and aggregates these dispersed model updates trained on different data distributions using a specific fusion algorithm (e.g., federated averaging), thereby generating an updated, more powerful global prediction model. Finally, the server distributes this new model to all clients, completing one learning cycle.
[0039] Local model training is performed on each distributed node, with each edge client maintaining two models: a local prediction model and a local policy network. For example, the local prediction model... It could be a lightweight CNN neural network used to learn static mappings of the local environment, such as... ,in, For node i, the meteorological data includes light intensity, ambient temperature, and wind speed. Contributing to photovoltaic forecasting. Local strategy network. As an actor network in multi-agent reinforcement learning, it can adopt a fully connected neural network structure, with local state as input. Output scheduling action (e.g., charging and discharging power).
[0040] Cloud servers are used for collaborative aggregation. To address the issue of non-independent and identically distributed data, personalized aggregation strategies based on clustering or similarity are employed. Global prediction model. The input consists of privacy-protected weather and equipment status data from each client, and the output is the global distributed energy output and load forecast results. In some implementations, the cloud server is specifically used for: The contribution weight of each local prediction model is determined based on the similarity between the local prediction model and the global prediction model of each edge client. Based on the contribution weights and model parameter update information of each local prediction model, parameter fusion is performed to obtain fused model parameters, and the global prediction model is updated using the fused model parameters.
[0041] In this example implementation, the cloud server assigns a dynamically calculated contribution weight to each client's local model update during aggregation. This weight is determined based on two key factors: first, the amount of local data used by the client in this training round; the larger the data volume, the more information the update is generally considered to contain; second, the similarity between the client's local power prediction model and the current global prediction model. The similarity function measures the consistency or correlation between the local and global prediction models in the parameter space or function space. The parameter update formula for the global prediction model is:
[0042] In the formula, For the first Model parameters of the global prediction model; For the first Model parameters of the global prediction model; For client / smart agent In the Model parameters of the local prediction model after rounds of training; For any client / agent In the The model parameters of the local prediction model after one round of training; N is the number of clients / agents; For the client The amount of training data in this round; Any client The amount of training data in this round; This is a similarity function between the local prediction model and the global prediction model, used to dynamically adjust the contribution weights. This makes aggregation more inclined to align with the global direction for clients.
[0043] In some embodiments, the multi-agent cooperative decision-making subsystem includes: A central network of commentators, deployed on cloud servers, is used to evaluate the value of joint actions of individual edge agents in joint states during the training phase. Multiple policy networks, each deployed on an edge agent corresponding to a distributed energy resource, are used for: during the training phase, generating scheduling instructions based on local observation states and historical action statistics of other agents, and updating the parameters of the policy networks based on the value assessment output by the central critic network; during the execution phase, independently running the corresponding policy network based on the local observation states of each edge agent to generate distributed scheduling instructions. The edge agent corresponds one-to-one with the distributed node.
[0044] In this example implementation, VPP scheduling is modeled as a decentralized, partially observable Markov decision process.
[0045] Each DER is treated as an edge agent, with state... Includes local privacy-protected observation data Summary information of historical actions of other intelligent agents (Statistical characteristics such as the average action value and action change rate of other agents in the previous multiple scheduling cycles are obtained through communication networks or attention mechanisms.) Action This represents the scheduling instructions (such as power generation / charging power) of the DER. The reward function guides agents to collaboratively achieve the global goal. This subsystem employs an advanced framework in reinforcement learning for multi-agent environments: centralized training and distributed execution. The system comprises a central critic network deployed on a cloud server and local policy networks deployed on each edge agent (i.e., distributed nodes). During the training phase, the system simulates or utilizes historical data for iterative learning. The central critic network acts as the global coach, evaluating the overall value of joint actions taken by acquiring the joint state information (i.e., the global perspective) of all edge agents. Each edge agent's policy network acts as the executor, generating a scheduling instruction based on its observed local state (e.g., local device power, energy storage state of charge, local prediction information) and the historical action statistics of other agents (excluding the current agent) (summary information obtained through communication or attention mechanisms). During the training phase, each policy network updates its parameters using methods such as policy gradients based on the value evaluation signal provided by the central critic network, learning how to make decisions more beneficial to the global goal. For example, during the training phase, the cloud server maintains a central critic network. (This central commentator network is a value network used in reinforcement learning.) The parameters of the central commentator network (including the weights and biases of each layer of the network) allow access to the joint state of all agents. and joint actions This is used to more accurately evaluate the value of actions. The policy network of each agent... By updating the policy gradient method and evaluating the value of joint actions through a central commentator network, the system guides the updates of the agent networks, making agents inclined to choose high-value actions. This solves the problem of poor coordination in multi-agent environments due to the lack of a global perspective. When training is complete and the system enters the actual execution phase, it switches to a decentralized execution mode. At this time, the central commentator network no longer directly participates in the generation of each decision. Each edge agent relies entirely on its own pre-trained local policy network, making independent and rapid decisions based on real-time local observations, generating distributed scheduling instructions, and achieving fast and decentralized decision-making responses. This architecture ensures both the ability to learn efficient cooperative policies using global information during the training phase and the speed, privacy, and robustness of decision-making during the execution phase (not relying on real-time communication with a central server), perfectly adapting to the distributed and real-time operational characteristics of virtual power plants.
[0046] In some implementations, the blockchain evidence storage subsystem further includes a smart contract module, which is used for: The consistency between the scheduling instructions and the actual execution results of the scheduling instructions is verified. Based on the verification results and the pre-stored contribution weight list, the corresponding incentive items are calculated and the corresponding incentive vouchers are issued to each edge client.
[0047] In this example implementation, the blockchain evidence storage subsystem provides a verifiable and non-repudiable chain of evidence for all key interactions. A consortium blockchain network is constructed using Hyperledger Fabric, with VPP operators, major participants, and regulatory bodies acting as consensus nodes (responsible for block generation and consensus reaching, and these are core participants). On-chain evidence storage of key events can include model contribution storage, i.e., the global model hash is recorded after each round of federated learning. List of participating clients and their contribution weights Uploading to the blockchain. The storage of scheduling instructions generates a set of scheduling instructions. Then, calculate its hash. And upload it to the blockchain along with the timestamp. The scheduling instructions are for client / smart agent N; execution result verification: after each client executes the instruction, it uploads the hash of the actual execution data (such as actual power) to the blockchain. The smart contract automatically compares the instruction hash with the execution result hash to verify the execution integrity and generates a verification result. Loyalty is determined by a preset deviation threshold, i.e., if , If a preset power deviation threshold is set, then the verification result will be... This indicates that the execution was successful; otherwise... A failure to meet the verification criteria indicates an unsatisfactory execution. Based on the verification results and other pre-stored information on the chain (such as a list of contribution weights for each node obtained from the federated learning process records), the smart contract module executes a transparent incentive calculation logic. This logic comprehensively considers factors such as the node's contribution to model training and the execution quality of scheduling instructions (verification results) to calculate the incentive value that the node should receive. Finally, the module automatically triggers the on-chain incentive certificate (which can be a token, points, or a credential for settlement with an external system) distribution process, distributing the incentive to the on-chain address corresponding to the node. This ensures the timeliness, accuracy, and non-repudiation of incentive settlement.
[0048] In some implementations, the reward function corresponding to the scheduling instruction output by the policy network includes an incentive term calculated based on blockchain-stored information; the incentive term is calculated and distributed through a smart contract in the blockchain-stored information subsystem.
[0049] In this example implementation, a key coupling link is established between the multi-agent decision-making subsystem and the blockchain evidence storage subsystem. The reward function used by the policy network of each agent (i.e., edge node) when generating scheduling instructions must include a special incentive term calculated based on blockchain evidence storage information. This means that when learning and making decisions, agents need to consider verifiable economic incentives from the blockchain system as part of their decision-making rewards. This incentive is not a pre-set fixed value, but a value certificate dynamically calculated and issued by smart contracts in the blockchain evidence storage subsystem based on the agent's real-time or recent on-chain behavior records (such as model contribution, historical instruction execution verification records, etc.). By internalizing this on-chain incentive as part of the agent's reinforcement learning algorithm, a deep closed loop of algorithm-level behavioral guidance and economic-level value feedback is achieved. To maximize its long-term cumulative rewards (including the economic incentive term), the agent will spontaneously optimize its behavior during the learning process, tending to make decisions that both satisfy the global optimization objective (reflected in other terms of the reward function) and obtain higher on-chain incentives. This fundamentally drives distributed nodes to become trusted, active, and valuable collaborative participants in the virtual power plant.
[0050] For example, the reward function corresponding to the scheduling instruction output by the policy network is:
[0051] in, For edge intelligent agents During the period The reward function value, For edge intelligent agents During the period Electricity / maintenance costs; For virtual power plants in different time periods The power exchanged with the power grid, and the target value of that power exchange; For virtual power plants in time periods The local consumption rate of renewable energy; Issuance to edge intelligent agents by the blockchain evidence storage subsystem During the period Incentive items; These are the weighting coefficients for each reward; Edge agents The incentives obtained are:
[0052] in, This serves as the total incentive pool for this round of model updates; Let be the excitation function. For edge intelligent agents During the period The actual execution power, For edge intelligent agents During the period The power adjustment amount corresponding to the scheduling command; For edge intelligent agents During the period The verification result is 1 if the verification is successful, and 0 otherwise. For edge intelligent agents Long-term reputation score, based on edge intelligent agents Historical contributions and performance records are dynamically updated; , Edge agents Weighting coefficients for each incentive component.
[0053] like Figure 3 The diagram illustrates the key event recording and automatic incentive settlement process of the blockchain evidence storage subsystem. Event 1's on-chain process involves generating a global prediction model hash and contribution list after model aggregation, generating transaction Tx1, connecting to consensus, and then on-chain. Event 2's on-chain process involves generating a scheduling instruction, triggering the corresponding instruction hash generation, generating transaction Tx2, connecting to consensus, and then on-chain. Event 3's on-chain process involves generating an execution result hash after local instruction execution, generating transaction Tx3, connecting to consensus, and then on-chain. The smart contract automatically triggers the on-chain evidence storage information, compares the instruction hash with the instruction execution result hash for consistency verification, and calculates the incentive value based on the on-chain contribution list and verification results; incentives are then automatically distributed based on the incentive value.
[0054] like Figure 4The diagram illustrates the process of recording model contributions to the blockchain in a federated learning subsystem. The cloud server of the federated learning subsystem exchanges model parameters with multiple edge clients. The cloud server outputs a global prediction model, and the clients return local model updates. Afterward, the cloud server outputs a data packet to be stored, which may include the global prediction model hash and a list of client contributions (such as the contribution weight of each client). This data packet is then inflated into a transaction and placed in a block along with other transactions. The block contains a block header (parent hash, timestamp, Merkle root, and transaction list). The block is then linked to a blockchain consisting of consecutive blocks, completing the on-chain process.
[0055] In some implementations, it also includes: The digital twin visualization subsystem is used to synchronize and display the global prediction model, the collaborative scheduling instruction set, on-chain evidence information and incentive settlement results in real time, and to trace the decision-making process of the specified scheduling instruction based on the real-time synchronized data.
[0056] In this example implementation, the digital twin visualization subsystem does not directly participate in control and optimization decisions but serves as a crucial human-computer interaction interface. This subsystem continuously obtains the latest global prediction model and its output results from the federated learning subsystem via a data interface, acquires the generated collaborative scheduling instruction set from the multi-agent collaborative decision-making subsystem, and obtains on-chain evidence storage information and incentive settlement results from the blockchain evidence storage subsystem. It then integrates and visualizes this multi-source, heterogeneous data on a unified virtualization platform. This allows operators to intuitively and comprehensively grasp the real-time operational panorama of the virtual power plant, the basis of AI decision-making, and the flow of value on a single interface, significantly improving operational efficiency and situational awareness. This subsystem provides powerful decision-making process traceability capabilities. Addressing concerns about the potential "black box" nature of AI decision-making, operators can query any specified scheduling instruction in the history. The system can, based on real-time synchronized full data, correlate and display all federated learning prediction data relied upon when the instruction was generated (e.g., prediction curves for future weather and output at that time), clearly revealing the data foundation of the decision. Simultaneously, it can display the action value estimation output (e.g., Q-value or action probability distribution) within the policy network of the relevant intelligent agent when the instruction is generated, visually explaining "why the intelligent model makes this decision." Furthermore, by connecting to a blockchain explorer, it can display the entire lifecycle of the instruction's notation on the blockchain with a single click, including its generation hash, issuance time, execution result hash, and verification status, forming a complete trust chain. This effectively solves the challenges of decision credibility and interpretability, meeting the stringent requirements of the power industry for high security and regulatory compliance.
[0057] In some implementations, the digital twin visualization subsystem includes: A digital twin construction module is used to construct a virtual model that is synchronized with the virtual power plant in real time. The strategy inference module is used to simulate and compare the effects of scheduling strategies in different future scenarios based on a trained multi-agent reinforcement learning model. The decision tracing module is used to respond to query requests and display the federated learning prediction results on which the specified historical scheduling instruction depends, the decision logic of the corresponding agent when the scheduling instruction was generated, and the full lifecycle evidence record of the scheduling instruction on the blockchain.
[0058] In this example implementation, the digital twin construction module is used to establish a 3D virtual model that is synchronized with the physical VPP in real time, integrating multi-dimensional models such as physical, electrical, and market aspects. The strategy deduction module integrates a pre-trained multi-agent reinforcement learning model, providing a secure "sandbox" environment. Operators can set various scenario parameters that may occur in the future in this module, such as extreme electricity prices, sudden drops in renewable energy, and sudden load surges. The system uses a pre-trained MADRL model to perform multi-step forward-looking deductions, and compares the deduction results with long-term costs, carbon emissions, and other indicators of different strategies using visualized curves (long-term costs can include the sum of the differences between grid power purchase costs, equipment operation and maintenance costs, and blockchain incentive benefits for each period within the deduction period; carbon emissions are the sum of the differences between grid power purchase carbon emissions and local renewable energy consumption carbon emission reductions for each period within the deduction period). When a historical dispatch instruction needs to be reviewed, the decision tracing module responds to the query request, quickly retrieves and organizes relevant information from the data warehouse, and displays it side-by-side on a dedicated dashboard interface. For any historical scheduling instruction, one-click tracing is possible. The tracing process is as follows: For data tracing, the federated learning prediction results upon which the decision was based are displayed. This is achieved by retrieving the input and output data of the corresponding prediction model through the global model hash stored on the blockchain. For decision logic tracing, the probability distribution or Q-value (evaluation value) of the policy network output of the relevant intelligent agent at the decision-making moment is displayed. This is achieved by retrieving the network parameters and state input records stored locally by the edge intelligent agent at the decision-making moment. For trust tracing, the instruction hash value is entered through the blockchain explorer API (Application Programming Interface) to query and display the entire on-chain evidence record of the instruction from generation and issuance to execution verification, including transaction hash, block height, timestamp, and verification result.
[0059] For example, such as Figure 5The image shows the visualization monitoring and decision tracing interface of the digital twin visualization subsystem. It includes the main view area on the left, displaying a 3D park map with colored icons representing photovoltaics, energy storage, charging piles, and buildings. Each icon has dynamic data labels such as real-time power and SOC (State of Charge). The upper right panel displays key indicators in dashboard and digital form, including total system cost (cumulative electricity cost + operation and maintenance cost - incentive revenue), real-time carbon emissions (calculated based on grid power purchase ratio and green electricity consumption rate), renewable energy consumption rate (renewable energy output / (renewable energy output + grid power purchase)), and network stability index (calculated based on power fluctuation variance and preset threshold normalization). The decision traceability panel in the lower right corner is a pop-up window that displays the results of scheduling instruction traceability. It can include dependent data (photovoltaic output curve for the next 15 minutes), decision logic (such as the estimated value of agent #203 (energy storage) action as discharging 1MW (Q value of 85) and charging 0.5MW (Q value of 60)) and blockchain evidence (such as instruction hash as 0xabc123, generation time as 10:00:00, and verification status as verified).
[0060] For example, the system workflow provided in this example is as follows: Initialization: Each DER registers, deploys edge agents, and joins the federation and blockchain network.
[0061] Federated Awareness: Initiating privacy-enhanced federated learning, each client uses local data to train prediction and policy models, the server performs personalized aggregation, updates global knowledge, and records contributions on the blockchain.
[0062] Collaborative decision-making: At each scheduling moment, each edge agent generates scheduling actions through its local policy network based on its local state and guidance from the global commentator network. The central server integrates these actions to form an instruction set, hashes the instructions, and distributes them after they are uploaded to the blockchain.
[0063] Trusted execution and feedback: Edge agents execute instructions and hash the execution results onto the blockchain. Blockchain smart contracts automatically verify execution deviations.
[0064] Online learning and incentives: After each edge agent executes the scheduling instruction, its state before execution is collected. Execution of actions Instant rewards and post-execution status Forming state transition data ( The data is then stored in a local experience pool for online fine-tuning of the strategy network. Simultaneously, the smart contract automatically settles and distributes incentives based on the on-chain evidence of contributions and execution records.
[0065] Visualized monitoring: Operation and maintenance personnel can monitor the entire process in real time through a digital twin interface, and can trace decisions and make manual interventions.
[0066] The following uses a park-level VPP (Virtual Power Plant) that includes photovoltaics, energy storage, charging piles, and commercial buildings as an example to illustrate the specific implementation of this invention. The system deployment is as follows: Physical layer: The park has rooftop photovoltaic (total capacity 2MW), energy storage power station (1MW / 2MWh), electric vehicle charging piles (50, V2G (reverse charging)), and adjustable load for commercial buildings (500kW).
[0067] Edge layer: Deploy edge agent software in each photovoltaic inverter, energy storage converter, charging pile controller, and building energy management system (BMS).
[0068] Platform layer: Deploy the VPP-TrustOS central platform in the park's cloud data center, which includes cloud servers, MADRL training and decision engine, Hyperledger Fabric blockchain network (nodes include: park operator, energy storage owner, charging pile operator, third-party auditing agency), and digital twin rendering server.
[0069] The key technical parameters are configured as follows: Federated learning: The training cycle is 15 minutes. The FedProx algorithm is used to handle device heterogeneity, and the training process is divided into two stages: local training and global aggregation.
[0070] (1) Local training phase: Each client updates its local model iteratively by minimizing the loss function based on local data. The loss function is as follows:
[0071] in, Let i be the loss function of the local prediction model corresponding to the client / agent i. For the model parameters of the local prediction model corresponding to client / agent i; The loss function for prediction (such as mean squared error or average absolute error, etc.). For client / agent i, the input samples (such as local state data) correspond to the local prediction model. The number of local samples for client / agent i. for Labels (such as actual output during the local forecast period). For the mapping function of the local prediction model corresponding to client / agent i, These are the proximal term coefficients, used to limit the deviation between the local prediction model and the global prediction model. These are the model parameters for the global prediction model.
[0072] (2) Global Aggregation Stage: The server collects the local model parameters from each client, aggregates them according to contribution weights to generate a new global prediction model, and distributes it to the clients. The local training loss function includes a proximate term. MADRL: The state space includes real-time power, SOC, electricity price forecast, and weather forecast for each DER. The action space contains continuous power values. The MADDPG algorithm is used, and the commentator network is a fully connected network.
[0073] Blockchain: Utilizing Kafka sorting service, transaction finality time (the time from transaction submission to on-chain confirmation) is less than 2 seconds. Incentive smart contracts automatically settle accounts every 24 hours.
[0074] Digital Twin: Utilizes WebGL technology to achieve 3D visualization on the browser side, with a data interface of WebSocket, enabling sub-second refresh rates.
[0075] The proposed optimization system uses the output of the federated learning subsystem as the state input of the multi-agent collaborative decision-making subsystem. The output instructions of the multi-agent collaborative decision-making subsystem are sent to the blockchain evidence storage subsystem for storage before execution. The execution results are fed back to the multi-agent collaborative decision-making subsystem for online learning, simultaneously triggering incentive settlement in the blockchain evidence storage subsystem. All processes can be monitored and traced by the digital twin visualization subsystem. This achieves a model-driven, data-stationary system with verifiable privacy and security. Through triple protection of local differential privacy, federated learning, and blockchain evidence storage, the system eliminates the risk of original data leakage while leveraging blockchain to ensure the fairness and auditability of the federated learning process itself, fundamentally establishing trust among participants. It achieves adaptive collaborative decision-making and globally optimal dynamic response. Employing the MADRL CTDE framework, the system possesses online learning and dynamic optimization capabilities, enabling it to proactively explore optimal strategies, effectively address source-load dual uncertainties, and significantly improve overall energy efficiency. Simulations show that compared to traditional model predictive control, the total system operating cost is reduced by 15%-25%. Blockchain-based smart contracts quantify model contributions and scheduling responses into automatically settled incentives, achieving a reward-for-effort and performance-based incentive mechanism that greatly enhances user participation and system regulation capabilities. It breaks down the "black box" of AI (Artificial Intelligence), making the decision-making process explainable. Through digital twins and decision traceability dashboards, complex AI decision-making logic is presented visually, enabling operators and regulators to understand, trust, and effectively supervise AI decisions, meeting the power industry's high requirements for safety and transparency. It provides a complete, high-barrier standardized solution: four deeply coupled subsystems form a complete technical closed loop from perception, decision-making, execution to auditing, constituting a highly reliable and highly available VPP operating system for new power systems. This technology is difficult to replicate and has high commercial value.
[0076] Example 2 Based on the same inventive concept, this invention also provides a reliable collaborative operation optimization method for a virtual power plant, comprising: In a virtual power plant, each distributed node trains a collaborative model based on its own local state data in a privacy-preserving manner to obtain a global prediction model. Each distributed node generates a set of cooperative scheduling instructions and executes the corresponding scheduling instructions based on the power prediction data obtained by the global prediction model through multi-agent reinforcement learning. Each distributed node uploads its contribution information, collaborative scheduling instruction set, and instruction execution results to the blockchain network for on-chain storage. The blockchain network automatically executes incentive settlement based on the stored information and verifies the execution deviation of the scheduling instructions; In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.
[0077] In one possible implementation, each distributed node in the virtual power plant trains a collaborative model based on its own local state data in a privacy-preserving manner to obtain a global prediction model, including: Each distributed node performs privacy processing on its local state data and trains a local prediction model based on the privacy-processed data. The cloud server receives model parameter update information uploaded by each distributed node, updates the global prediction model by fusing the update information of each model parameter, and distributes the updated global prediction model to each distributed node. The privacy processing involves adding differential privacy noise to the local state data; one distributed energy resource corresponds to one distributed node.
[0078] In one possible implementation, the step of updating the global prediction model through the fusion processing of update information for each model parameter includes: The contribution weight of each local prediction model is determined based on the similarity between the local prediction model and the global prediction model. Based on the contribution weights and model parameter update information of each local prediction model, parameter fusion is performed to obtain fused model parameters, and the global prediction model is updated using the fused model parameters.
[0079] In one possible implementation, each distributed node generates a cooperative scheduling instruction set based on power prediction data obtained from the global prediction model, through multi-agent reinforcement learning, including: The cloud server utilizes a central commentator network to evaluate the value of joint actions of each edge agent in a joint state during the training phase; During the training phase, each edge agent utilizes its local policy network to generate scheduling instructions based on its local observation state and the historical action statistics of other agents, and updates the parameters of the policy network based on the value assessment output by the central critic network. During the execution phase, each edge agent independently runs its corresponding policy network based on its local observation state to generate distributed scheduling instructions. The edge agent corresponds one-to-one with the distributed node.
[0080] In one possible implementation, the blockchain network automatically executes incentive settlement based on the stored evidence information and verifies the execution deviation of the scheduling instructions, including: The blockchain network verifies the consistency between scheduling instructions and their actual execution results. Based on the verification results and the pre-stored list of contribution weights, it calculates the corresponding incentive items and issues the corresponding incentive certificates to each distributed node.
[0081] In one possible implementation, the reward function corresponding to the scheduling instruction output by the policy network includes an incentive term calculated based on blockchain-stored information; the incentive term is calculated and distributed through smart contracts in the blockchain network.
[0082] In one possible implementation, the reward function corresponding to the scheduling instruction output by the policy network is:
[0083] in, For edge intelligent agents During the period The reward function value, For edge intelligent agents During the period Electricity / maintenance costs; For virtual power plants in different time periods The power exchanged with the power grid, and the target value of that power exchange; For virtual power plants in time periods The local consumption rate of renewable energy; Issuance to edge intelligent agents by the blockchain evidence storage subsystem During the period Incentive items; These are the weighting coefficients for each reward; Edge agents The incentives obtained are:
[0084] in, This serves as the total incentive pool for this round of model updates; Let be the excitation function. For edge intelligent agents During the period The actual execution power, For edge intelligent agents During the period The power adjustment amount corresponding to the scheduling command; For edge intelligent agents During the period The verification result is 1 if the verification is successful, and 0 otherwise. For edge intelligent agents Long-term reputation score, based on edge intelligent agents Historical contributions and performance records are dynamically updated; , Edge agents Weighting coefficients for each incentive component.
[0085] In one possible implementation, it also includes: The global prediction model, the collaborative scheduling instruction set, on-chain evidence storage information, and incentive settlement results are synchronized and displayed in real time through the digital twin visualization subsystem, and the decision-making process of the specified scheduling instruction is traced based on the real-time synchronized data.
[0086] In one possible implementation, tracing the decision-making process for a specified scheduling instruction based on real-time synchronized data includes: Construct a virtual model that is synchronized with the virtual power plant in real time; Based on a trained multi-agent reinforcement learning model, the scheduling strategies in different future scenarios are simulated and their effects are compared. In response to a query request, the system displays the federated learning prediction results upon which the specified historical scheduling instruction is based, the decision logic of the corresponding agent when the scheduling instruction was generated, and the full lifecycle record of the scheduling instruction on the blockchain.
[0087] For example, such as Figure 6 As shown, a time-series flowchart illustrates the system's operational steps and interactions between components throughout a complete scheduling cycle. Key participants include the federated server, the MADRL decision center, the blockchain network, edge agents (clients), and the digital twin. The process unfolds chronologically from top to bottom: Step 1 (Federated Training): The edge agent outputs local model updates (the parameter updates of the local prediction model and policy model) to the federated server. After processing, the federated server calculates contribution weights based on a personalized federated aggregation algorithm, combining the data volume of each client and model similarity, aggregates to generate a global model, outputs the global model and contribution records to the blockchain network, and broadcasts the updated model to the edge agent.
[0088] Step 2 (Collaborative Decision-Making): The edge agent outputs its local state to the MADRL decision center, that is, it outputs the local privacy-preserving device state and environmental observation data. After calculation by the MADRL decision center (inputting the joint state of all agents, evaluating the value of actions through a global commentator network, guiding the agent actor network to generate actions, and integrating them into a scheduling instruction set), the scheduling instructions are output to the blockchain network frame for storage and simultaneously distributed to the edge agents.
[0089] Step 3 (Trusted Execution): After the edge agent executes the action locally, it outputs the execution result hash to the blockchain network for verification.
[0090] Step 4 (Online Learning and Incentives): The blockchain network triggers a smart contract, automatically calculating and distributing incentives to each edge agent. Simultaneously, the edge agents feed back the state transition data generated during execution to the MADRL decision center for online learning.
[0091] Step 5 (Visual Synchronization): All key data (status, instructions, and evidence records) throughout the process are synchronized to the digital twin in real time for display.
[0092] The following specific embodiment illustrates the application process of the optimization method of the present invention.
[0093] 1) One afternoon, the output of photovoltaic power suddenly dropped. The federated learning subsystem predicted the downward trend of output 10 minutes in advance based on local meteorological data from each photovoltaic site.
[0094] 2) Upon receiving this prediction, the MADRL decision subsystem, combined with the real-time electricity price (during peak hours), evaluates the value of the joint actions of each agent through a centrally trained global commentator network. This guides the energy storage agent to output a discharge action, the charging pile agent to output a V2G discharge action, and the building agent to output a non-critical load reduction action. Through collaborative calculation among the agents, the optimal strategy is determined: prioritize the discharge of energy storage and start the V2G discharge mode of the charging pile, while slightly reducing the non-critical load of the building.
[0095] 3) The scheduling instruction set is generated, and its hash value is immediately stored on the blockchain. The instructions are then issued to each edge agent.
[0096] 4) Each edge agent controls the local device to execute commands and hashes the actual output / load data onto the blockchain. The blockchain smart contract verifies the execution and records a successful collaborative response.
[0097] 5) The digital twin interface displays the dynamic process of photovoltaic output curve decline, energy storage SOC decline and V2G discharge power increase in real time, and highlights the cost savings brought about by this automatic decision (cost savings = peak period grid purchase price × (grid purchase power before dispatch - grid purchase power after dispatch) × dispatch duration - energy storage discharge loss cost - charging pile V2G response cost).
[0098] 6) At the time of settlement, the smart contract automatically distributes digital rewards to each participant from the incentive pool based on the model-predicted contribution and the actual adjusted contribution of each participant in this event.
[0099] This invention deploys a local policy network (Actor) on a distributed client and a global value evaluation network (Critic) on a central server, achieving collaborative optimization through a framework of centralized training and distributed execution. Furthermore, the parameters of the client's local policy network are updated using a federated averaging algorithm with personalized weights, determined by the amount of local data on the client and the similarity between its model and the global model. In the agent's reward function, in addition to cost, power smoothing, and green energy consumption terms, a blockchain-based incentive term is added (verifiable incentive certificates calculated and issued via on-chain smart contracts based on the agent's historical model contribution and scheduling instruction execution loyalty), thus achieving a closed-loop unification of algorithmic and economic incentives. Through a high-fidelity VPP model and decision traceability dashboard, in response to a query operation for a historical scheduling instruction, the federated learning prediction data relied upon when the instruction was generated, the action value estimation output of the relevant agent's policy network, and the hash and verification results of the instruction's full lifecycle record in the blockchain network are displayed in a correlated and side-by-side manner.
[0100] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.
Claims
1. A trusted collaborative operation optimization system for a virtual power plant, characterized in that, include: The federated learning subsystem is used to train a collaborative model in a privacy-preserving manner based on the local state data of each distributed node in the virtual power plant, so as to obtain a global prediction model. A multi-agent collaborative decision-making subsystem, which is communicatively connected to the federated learning subsystem, is used to generate a collaborative scheduling instruction set based on the power prediction data obtained by the global prediction model, so that each distributed node executes the corresponding scheduling instruction; The blockchain evidence storage subsystem is communicatively connected to both the federated learning subsystem and the multi-agent collaborative decision-making subsystem. It receives contribution information from the federated learning subsystem during the collaborative model training process, the collaborative scheduling instruction set from the multi-agent collaborative decision-making subsystem, and the instruction execution results from each distributed node, and stores these on-chain. Based on the evidence storage information, it automatically executes incentive settlement and verifies the execution deviation of the scheduling instructions. In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.
2. The system according to claim 1, characterized in that, The federated learning subsystem includes: Multiple edge clients, each deployed on a distributed energy resource side, are used to perform privacy processing on local state data and train local prediction models based on the privacy-processed data; The cloud server is used to receive model parameter update information uploaded by each edge client, and to update the global prediction model by fusing the update information of each model parameter. The updated global prediction model is then distributed to each edge client. The privacy processing involves adding differential privacy noise to the local state data; one distributed energy resource corresponds to one distributed node.
3. The system according to claim 2, characterized in that, The cloud server is specifically used for: The contribution weight of each local prediction model is determined based on the similarity between the local prediction model and the global prediction model of each edge client. Based on the contribution weights and model parameter update information of each local prediction model, parameter fusion is performed to obtain fused model parameters, and the global prediction model is updated using the fused model parameters.
4. The system according to claim 2, characterized in that, The multi-agent cooperative decision-making subsystem includes: A central network of commentators, deployed on cloud servers, is used to evaluate the value of joint actions of individual edge agents in joint states during the training phase. Multiple policy networks, each deployed on an edge agent corresponding to a distributed energy resource, are used for: during the training phase, generating scheduling instructions based on local observation states and historical action statistics of other agents, and updating the parameters of the policy networks based on the value assessment output by the central critic network; during the execution phase, independently running the corresponding policy network based on the local observation states of each edge agent to generate distributed scheduling instructions. The edge agent corresponds one-to-one with the distributed node.
5. The system according to claim 1, characterized in that, The blockchain-based evidence storage subsystem also includes a smart contract module, which is used for: The consistency between the scheduling instructions and the actual execution results of the scheduling instructions is verified. Based on the verification results and the pre-stored contribution weight list, the corresponding incentive items are calculated and the corresponding incentive vouchers are issued to each edge client.
6. The system according to claim 5, characterized in that, The reward function corresponding to the scheduling instruction output by the strategy network includes an incentive term calculated based on blockchain-stored information; the incentive term is calculated and distributed through a smart contract in the blockchain-stored information subsystem.
7. The system according to claim 6, characterized in that, The reward function corresponding to the scheduling instruction output by the policy network is: in, For edge intelligent agents During the period The reward function value, For edge intelligent agents During the period Electricity / maintenance costs; For virtual power plants in different time periods The power exchanged with the power grid, and the target value of that power exchange; For virtual power plants in time periods The local consumption rate of renewable energy; Issuance to edge intelligent agents by the blockchain evidence storage subsystem During the period Incentive items; These are the weighting coefficients for each reward; Edge agents The incentives obtained are: in, This serves as the total incentive pool for this round of model updates; Let be the excitation function. For edge intelligent agents During the period The actual execution power, For edge intelligent agents During the period The power adjustment amount corresponding to the scheduling command; For edge intelligent agents During the period The verification result is 1 if the verification is successful, and 0 otherwise. For edge intelligent agents Long-term reputation score, based on edge intelligent agents Historical contributions and performance records are dynamically updated; , Edge agents Weighting coefficients for each incentive component.
8. The system according to claim 1, characterized in that, Also includes: The digital twin visualization subsystem is used to synchronize and display the global prediction model, the collaborative scheduling instruction set, on-chain evidence information and incentive settlement results in real time, and to trace the decision-making process of the specified scheduling instruction based on the real-time synchronized data.
9. The system according to claim 8, characterized in that, The digital twin visualization subsystem includes: A digital twin construction module is used to construct a virtual model that is synchronized with the virtual power plant in real time. The strategy inference module is used to simulate and compare the effects of scheduling strategies in different future scenarios based on a trained multi-agent reinforcement learning model. The decision tracing module is used to respond to query requests and display the federated learning prediction results on which the specified historical scheduling instruction depends, the decision logic of the corresponding agent when the scheduling instruction was generated, and the full lifecycle evidence record of the scheduling instruction on the blockchain.
10. A reliable collaborative operation optimization method for a virtual power plant, characterized in that, The method includes: In a virtual power plant, each distributed node trains a collaborative model based on its own local state data in a privacy-preserving manner to obtain a global prediction model. Each distributed node generates a set of cooperative scheduling instructions and executes the corresponding scheduling instructions based on the power prediction data obtained by the global prediction model through multi-agent reinforcement learning. Each distributed node uploads its contribution information, collaborative scheduling instruction set, and instruction execution results to the blockchain network for on-chain storage. The blockchain network automatically executes incentive settlement based on the stored information and verifies the execution deviation of the scheduling instructions; In this process, after executing the scheduling instructions, each distributed node uses the resulting state transition data for online optimization of the local policy network in multi-agent reinforcement learning.