Virtual power plant privacy protection data fusion method and system based on federated learning

By employing a three-tiered network architecture based on federated learning, along with methods such as feature filtering, encrypted transmission, and dynamic aggregation, the privacy leakage and model bias issues in virtual power plant data fusion are resolved, achieving efficient and secure data fusion and decision support.

CN121167786AActive Publication Date: 2025-12-19HUBEI UNIV OF EDUCATION

Patent Information

Application Number
CN202511697061.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2025-12-19
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing virtual power plant data fusion technologies suffer from privacy risks, high bandwidth consumption due to large data transmission volumes, large processing latency, inability of traditional feature selection strategies to adapt to dynamic operating states, and large deviations in model training results, all of which affect the efficient operation and decision support of virtual power plants.

Method used

A three-tier network architecture based on federated learning is adopted. Through a hierarchical collaboration mode of edge nodes, regional servers, and cloud platforms, a value-risk assessment model is constructed for feature screening. Paillier homomorphic encryption and sensitivity sharding are used to transmit parameters. Parameter aggregation is combined with time-series freshness weighting, and the model is iteratively optimized.

Benefits of technology

It achieves efficient data fusion under the premise of privacy protection, improves data availability and model accuracy, meets the real-time requirements of virtual power plants, and ensures the efficient operation of core tasks and decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167786A_ABST
    Figure CN121167786A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual power plant privacy protection data fusion method and system based on federated learning. The method comprises the following steps: deploying a three-level federated learning network; a value-risk assessment model is constructed, quantification of the value coefficient of the virtual power plant core task by the features is completed, and a feature identifiability entropy value is calculated as a risk coefficient; meanwhile, a load linkage temperature coefficient is introduced to complete feature screening; the edge nodes generate parameter update values based on the distillation characteristics; during transmission, the parameters are encrypted, fragment segmentation is completed, the parameters are uploaded, and then fault-tolerant verification is carried out; region aggregation is completed based on a time sequence freshness weighting method; at the cloud, abnormal parameters are filtered to complete global aggregation to generate an optimization model; and after the cloud issues the global model parameters to the edge nodes through the regional server, each node tests the model performance by using an independent local verification set to complete feedback iteration until the model converges. The problem of consideration of privacy protection and operation efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of virtual power plants, and more particularly relates to a virtual power plant privacy protection data fusion method and system based on federated learning. BACKGROUND

[0002] In the field of virtual power plants, data fusion is a key link to support core tasks such as load forecasting and energy scheduling. However, the current data fusion work faces many real problems, which seriously restricts the efficient and safe operation of virtual power plants.

[0003] On the one hand, traditional data fusion mostly adopts a centralized architecture, which requires each edge node (such as user-side terminals and distributed power stations) to upload raw data to the cloud or regional platform for unified processing. However, these raw data contain a large amount of sensitive information, such as user electricity usage habits and enterprise production energy consumption patterns, which have a high risk of privacy leakage during cross-node transmission and centralized storage. Moreover, as the number of nodes participating in the virtual power plant continues to increase, the amount of data transmission increases dramatically, causing excessive occupation of network bandwidth and significant increase in data processing delay, which makes it difficult to meet the strict real-time requirements of virtual power plants.

[0004] On the other hand, existing data fusion technologies have obvious deficiencies in feature selection and parameter processing. The operating state of a virtual power plant is dynamically changing, and the demand for data features varies greatly at different times. However, traditional methods often use fixed feature retention strategies and cannot adjust the feature selection criteria according to real-time load levels and energy supply and demand balance conditions, which can easily cause redundant features to interfere with model training or key features to be lost, greatly affecting the usability of the fused data. At the same time, during the parameter transmission and aggregation process, there is a lack of mechanisms that can effectively protect privacy while ensuring data integrity. Either no encryption is performed, which poses a risk of privacy leakage, or encryption is performed without supporting fault tolerance and dynamic weighting strategies, resulting in large deviations in the aggregation results and making it difficult to provide strong support for accurate decision-making in virtual power plants.

[0005] The existence of these problems not only hinders the effective use of multi-source data by virtual power plants, but also reduces the willingness of each participating subject to cooperate due to privacy issues, thereby affecting the improvement of the overall operational efficiency of virtual power plants, so there is an urgent need for a method that can achieve efficient data fusion while protecting privacy. SUMMARY

[0006] This invention aims to address key issues in current virtual power plant data fusion: firstly, it avoids the privacy leakage risk of cross-node transmission and storage of raw sensitive data under a centralized architecture, alleviates the problems of high bandwidth consumption and large processing latency caused by large data volumes, and meets real-time requirements; secondly, it breaks through the limitations of traditional fixed feature screening strategies, adapts to the dynamic operating status of virtual power plants, and reduces redundant feature interference and key feature loss; ultimately, by optimizing the data fusion effect, it improves data availability, assists virtual power plants in making accurate decisions, promotes the stable and efficient operation of energy systems, and solves the dilemma of balancing privacy protection and operational efficiency.

[0007] To address the aforementioned deficiencies or improvement needs of existing technologies, as a first aspect of this invention, the present invention provides a privacy-preserving data fusion method for virtual power plants based on federated learning, comprising: S1. Deploy a three-tiered federated learning network consisting of edge nodes, regional servers, and a cloud platform; S2. Construct a value-risk assessment model and determine the value coefficients of features for the core tasks of the virtual power plant. The quantification and calculation of the feature identifiability entropy value are used as a risk coefficient. to form Coordinates; Introducing a load-linked temperature coefficient that is tied to the real-time operating status of the virtual power plant. and complete feature filtering; S3. Edge nodes generate parameter update values ​​based on distillation features; during transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 segments based on their sensitivity, uploaded in no less than 3 channels, and then fault tolerance verification is performed; and regional aggregation is performed based on the time-series freshness weighting method; in the cloud, abnormal parameters are filtered and global aggregation is performed to generate an optimized model. S4. After the cloud distributes the global model parameters to the edge nodes via the regional server, each node tests the model performance using an independent local validation set. If the prediction error exceeds a preset threshold, the temperature coefficient of dynamic feature distillation is automatically readjusted. Then, local training and parameter uploading are restarted; if the performance meets the requirements, the next iteration begins until the model converges.

[0008] Furthermore, the three-level federated learning network in S1 is specifically as follows: The edge nodes correspond to the data holders in the virtual power plant. They use their own data to train models locally, generating only model parameter update information while keeping the original data completely local. The regional server is responsible for collecting and initially processing the model parameter updates of edge nodes within its region, and simultaneously performing parameter aggregation and verification within the region. The cloud platform aggregates parameter updates uploaded by servers in various regions, performs global model aggregation and optimization, generates the final global model, and then distributes the optimized model parameters to servers in various regions and edge nodes.

[0009] Furthermore, the value coefficient in S2 The calculation method is as follows: Quantifying features using Shapley values Contribution to virtual power plant models For the total number of features, For without Feature subset, For subset Model performance: , in, Features to be calculated The value coefficient; For subset The number of features contained therein; When only a subset is used When considering the characteristics of the virtual power plant core task model, the performance indicators exhibited by the model are as follows: To evaluate the features Add to subset Then, use These characteristics are the model's performance metrics; Total characteristic number factorial; For subset The factorial of the number of features; Total characteristic number Subtract subset size The factorial after subtracting 1; simultaneously, normalize the Shapley values ​​of all features, and... Map to the interval [0,1].

[0010] Furthermore, the risk coefficient in S2 The calculation method is as follows: , in, Features to be calculated The risk factor; Features The number of generalized categories to which it belongs is used to reflect the impact of the complexity of the classification scenario in which the feature is located on the risk; Features The number of discrete value categories of a feature; if it is a continuous feature, it needs to be discretized first. For the number of discrete categories, reflecting the diversity of the feature's own value; For the feature Index of discrete value; For the feature Take the Probability of value; At the same time, normalize the Value of all features, and Map to the interval [0, 1].

[0011] Further, the specific process of feature screening in S2 is: Let the feature set be For any feature , the necessary and sufficient condition for whether it is retained is: , Where, The identifiable risk coefficient of the feature ; The value coefficient of the feature To the core task of the virtual power plant; The load linkage temperature coefficient.

[0012] Further, the calculation method of sensitivity in S3 is: Let the evaluation function of the virtual power plant core task be When the parameter update value is , the task evaluation result is ; When the parameter has a small perturbation , the evaluation result is ; The sensitivity of parameter is defined as: , Where, The absolute change of the core task evaluation result before and after the parameter perturbation, reflecting the influence amplitude of the parameter on the task result; The perturbation amplitude of the parameter; For the convenience of subsequent slice quantity decision, the sensitivity is normalized to the interval [0, 1].

[0013] Further, the time series freshness weighting method in S3 is: When the edge node uploads the parameter update value after encryption and fault tolerance check to the regional server, the regional server marks a time stamp For each received parameter update value Synchronously extract three kinds of time sequence features of the parameter update value: absolute time difference , cycle correlation degree And update stability​ ; and map the three types of time sequence features to the interval [0, 1], denoted as , eliminating dimensional differences; Then, a collaborative weighting formula of “basic time decay + scene correlation enhancement + stable reliability addition” is constructed to avoid single dimension dominant weighting: , wherein, is the basic decay coefficient, controlling the basic influence of the time dimension on the weight; are the weight coefficients of correlation degree and stability, respectively; The weighted aggregation of the regional parameter update value is completed to obtain the regional-level parameter : , wherein, is the regional-level parameter update value obtained after regional aggregation; is the weight of the th edge node parameter update value; is the parameter update value generated locally by the th edge node; is the total number of edge nodes participating in this regional aggregation.

[0014] Further, the specific calculation method of the three types of time sequence features is: , , , wherein, is the absolute time difference of the parameter update value; is the periodic correlation degree of the parameter update value; is the update stability of the parameter update value; is the aggregation moment; is the parameter upload timestamp; is the duration of the load fluctuation period; is the starting moment of the load period; is the parameter update of the th edge node in the current period; is the parameter update value of the th edge node in the last period.

[0015] As a second aspect of the present application, the present application provides a virtual power plant privacy protection data fusion system based on federated learning, comprising: A learning network deployment unit is used to deploy a three-level federated learning network of edge node-regional server-cloud platform; The feature selection unit is used to build a value-risk assessment model and determine the value coefficient of features for the core tasks of the virtual power plant. The quantification and calculation of the feature identifiability entropy value are used as a risk coefficient. to form Coordinates; Introducing a load-linked temperature coefficient that is tied to the real-time operating status of the virtual power plant. And complete feature filtering; The global optimization unit is used by edge nodes to generate parameter update values ​​based on distillation features. During transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 segments based on their sensitivity, and uploaded through no less than 3 channels before fault tolerance verification is performed. Regional aggregation is completed based on the time-series freshness weighting method. In the cloud, abnormal parameters are filtered, and global aggregation is completed to generate an optimized model. The iterative feedback unit is used to distribute global model parameters from the cloud to edge nodes via regional servers. Each node then tests the model performance using an independent local validation set. If the prediction error exceeds a preset threshold, the temperature coefficient of dynamic feature distillation is automatically readjusted. Then, local training and parameter uploading are restarted; if the performance meets the requirements, the next iteration begins until the model converges.

[0016] As a third aspect of the invention, the invention provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of any step of the described federated learning-based virtual power plant privacy-preserving data fusion method.

[0017] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. The present invention provides a privacy-preserving data fusion method for virtual power plants based on federated learning. This method deploys a three-tiered federated learning network—edge nodes, regional servers, and a cloud platform—to construct a hierarchical collaborative model of "local training, regional aggregation, and cloud optimization." Edge nodes train and generate parameter update values ​​based solely on local data, without directly uploading the original data, thus reducing the risk of privacy leakage at the data source. Regional servers receive parameters from the edge nodes and perform initial aggregation, reducing the pressure on the cloud platform to directly process massive amounts of node data. The cloud platform then optimizes the global model based on the regional aggregation results and distributes the optimized model to the edge nodes for iteration. This architecture avoids the cross-node transmission of original privacy data and improves data fusion efficiency through hierarchical collaboration, ensuring that model training for core virtual power plant tasks (such as load forecasting and energy dispatch) proceeds in an orderly manner while protecting privacy.

[0018] 2. The federated learning-based virtual power plant privacy protection data fusion method of the present application quantifies the feature value coefficient and the risk coefficient by constructing a "value-risk" evaluation model, and introduces a load linkage temperature coefficient to complete feature screening. Among them, the value coefficient calculates the marginal contribution of the feature to the core task through the Shapley value, and the risk coefficient quantifies the identifiable risk of the feature based on information entropy, and the two form a corresponding coordinate; the load linkage temperature coefficient binds the real-time running state of the virtual power plant, and screens the features whose value and risk ratio meet the requirements. This screening mechanism ensures that the retained features not only have high value to support the core task, but also control the identifiable risk within a reasonable range, improving the quality of model input while further reducing the possibility of privacy leakage caused by feature correlation.

[0019] 3. The federated learning-based virtual power plant privacy protection data fusion method of the present application processes the parameter update value generated by the edge node through Paillier homomorphic encryption, and then adaptively splits it into 3-5 fragments based on parameter sensitivity, uploads it to the regional server through not less than 3 channels and completes fault tolerance verification, the regional server aggregates the parameters using the time sequence freshness weighting method, and the cloud filters the abnormal parameters and completes global aggregation. Paillier homomorphic encryption ensures that the parameters cannot be decrypted during transmission, multi-channel fragmentation and fault tolerance verification improve the security and integrity of parameter transmission, time sequence freshness weighting makes the regional aggregation result fit the real-time running demand, and cloud abnormal filtering ensures the quality of the global model. This process realizes privacy protection in the whole life cycle of the parameter, while ensuring the effectiveness and stability of the model after data fusion. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 The figure is a federated learning-based virtual power plant privacy protection data fusion method of an embodiment of the present application; Figure 2 The figure is an architecture level schematic diagram of an embodiment of the present application; Figure 3 The figure is a feature screening schematic diagram of an embodiment of the present application; Figure 4 The figure is a system unit diagram of an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0022] Embodiment 1 Please refer to Figure 1The application provides a virtual power plant privacy protection data fusion method based on federated learning, comprising: S1. Deploying an edge node-area server-cloud platform three-level federated learning network; S2. Building a value-risk assessment model, completing the quantification of the value coefficient of the characteristics to the core task of the virtual power plant, and calculating the entropy value of the feature identifiable as the risk coefficient , to form coordinates; introducing a load linkage temperature coefficient bound to the real-time operation state of the virtual power plant , and completing feature screening; S3. The edge node generates parameter update values based on distilled features; during transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 fragments based on their sensitivity, uploaded through not less than 3 channels, and then fault tolerance verification is completed; and based on the time sequence freshness weighting method, regional aggregation is completed; in the cloud, abnormal parameters are filtered, and a global aggregation is generated to generate an optimization model; S4. After the cloud endows the global model parameters to the edge node through the area server, each node tests the model performance with an independent local validation set, if the prediction error exceeds the preset threshold, the temperature coefficient of the dynamic feature distillation is automatically adjusted , and local training and parameter uploading are redeveloped; if the performance meets the standard, it enters the next round of iteration until the model converges.

[0023] This embodiment 1 further expands the above steps.

[0024] (1) Learning network deployment In the traditional centralized data processing mode of the virtual power plant, the privacy leakage risk and massive data processing pressure brought by cross-node transmission of raw data have become a key bottleneck restricting multi-agent collaboration. Therefore, referring to Figure 2 , this method first deploys an edge node-area server-cloud platform three-level federated learning network, and builds a hierarchical collaboration framework, laying a foundation for privacy protection and efficient collaboration in the subsequent data fusion process.

[0025] ​The edge nodes correspond to each data holding subject in the virtual power plant, including user side terminals, distributed power stations and the like. These nodes locally train a model using their own data, only generate model parameter update information, and the original data is completely retained locally, thereby avoiding sensitive data leakage from the source and reducing the amount of cross-node data transmission. The regional server undertakes the parameter update of the edge nodes in the region, is responsible for preliminary processing, regional parameter aggregation and verification, reduces the load of directly processing massive node data in the cloud, and provides regional level guarantee for the accuracy of the parameters. The cloud platform aggregates the parameter updates of the regional servers, performs global model aggregation and optimization, and generates the final global model which is distributed to each level node.

[0026] This hierarchical architecture solves the privacy and efficiency problems of traditional centralized architecture through the mode of "data not moving and model moving", and provides clear execution subjects and collaboration paths for subsequent feature selection, parameter encryption transmission and dynamic aggregation through clear hierarchical division of labor.

[0027] (2) Feature selection Please refer to Figure 3 After the three-level federated learning network is established and the data collaboration boundaries of each level are clear, in order to solve the problem that the traditional feature selection strategy cannot adapt to the dynamic running state of the virtual power plant and is prone to cause interference of redundant features or loss of key features, accurate feature selection is needed to provide high-quality input for subsequent parameter training and transmission, so the feature processing link is entered, and the specific process and role are as follows: This link first constructs a value-risk evaluation model to quantify the feature attributes from the "utility" and "security" dimensions. The value coefficient is calculated by Shapley value: considering the possible combination of a certain feature with all other features, the difference in model performance before and after adding the feature is determined to determine the marginal contribution of the feature to the core task, and after normalization, the importance of the feature in supporting load prediction, energy scheduling and other tasks is directly reflected, avoiding the deletion of key high-value features. The risk coefficient is measured by the feature identifiable entropy: combining the complexity of the feature category and the diversity of its own value, the probability distribution of the feature value is analyzed to quantify the risk of identifying sensitive information (such as user electricity usage habits) from the feature, and after normalization, the risk evaluation standard is unified to prevent high-risk features from causing privacy leakage.

[0028] In a preferred embodiment, the value coefficient is calculated as follows: The Shapley value is used to quantify the contribution of the feature to the virtual power plant model, is the total number of features, is a subset of features that does not contain is the subset is the subset Model performance: , wherein, is the value coefficient of the feature to be calculated; is the number of features contained in the subset ; and is the performance index of the virtual power plant core task model when only using the features in the subset . is the performance index of the model when using the features after adding the feature to be evaluated to the subset . is the factorial of the total number of features . is the factorial of the number of features in the subset . is the factorial of the total number of features minus the size of the subset minus 1; and the Shapley values of all features are normalized, and is mapped to the interval [0, 1].

[0029] In a preferred embodiment, the calculation method of the risk coefficient is: , wherein, is the risk coefficient of the feature to be calculated; is the number of general categories to which the feature belongs, used to reflect the influence of the complexity of the classification scenario where the feature is located on the risk; is the number of discrete value categories of the feature itself, if it is a continuous feature, it needs to be discretized first, is the number of discrete categories, reflecting the diversity of the value of the feature itself; is the index of the discrete value of the feature ; is the probability of the feature taking the th value; and the values of all features are normalized, and is mapped to the interval [0, 1].

[0030] After forming the two-dimensional coordinates of the features based on the value coefficient and the risk coefficient, a load linkage temperature coefficient As a screening threshold, the coefficient will be dynamically adjusted according to the real-time load level and energy supply and demand balance, avoiding the defects that fixed threshold cannot adapt to scenario changes.

[0031] In the screening, it is judged whether the ratio of the feature risk coefficient and the value coefficient reaches the threshold value. If it meets the condition, it is retained, otherwise it is rejected. In a preferred embodiment, the specific process of feature screening is as follows: Let the feature set be For any feature The necessary and sufficient condition that it needs to meet is: , Wherein, is the identifiable risk coefficient of feature ; is the value coefficient of feature to the core task of virtual power plant; is the load linkage temperature coefficient.

[0032] The core role of this screening process is: through value evaluation, key features supporting the task are retained, through risk control, the possibility of privacy leakage is reduced, and at the same time, a dynamic threshold is adapted to the running state, providing "high value-low risk" high-quality feature input for subsequent local training of edge nodes, improving the quality of parameter update from the source, and ensuring the effectiveness of subsequent regional aggregation and cloud optimization.

[0033] (3) Global optimization After completing feature screening and obtaining "high value-low risk" distillation features that adapt to the running state of virtual power plant, the edge node carries out local training based on these features and generates parameter update values, entering the parameter transmission and aggregation link. This link is the key link connecting the front-end feature processing and the final global model generation. Through the construction of "encrypted transmission-layered aggregation" closed loop process, it not only solves the problem that privacy protection and data integrity are difficult to be considered in traditional parameter processing, but also improves the aggregation accuracy through dynamic weighting, providing reliable support for global model optimization.

[0034] In the parameter transmission stage, Paillier homomorphic encryption is first used to process the parameter update values, preventing information leakage in the transmission process from the root. At the same time, adaptive fragmentation is carried out according to the parameter sensitivity-the sensitivity is determined by analyzing the influence of small changes in the parameter on the core task results, the greater the influence, the more the number of fragments (3-5 pieces), and through no less than 3 channels for uploading respectively, cooperating with the subsequent fault tolerance verification mechanism, ensuring that the parameter fragments are complete and have not been tampered with, which not only enhances the transmission security, but also avoids data loss caused by single channel failure.

[0035] In a preferred embodiment, the calculation method of sensitivity is as follows: Let the evaluation function of the core task of the virtual power plant be: When the parameter update value is At that time, the task evaluation result was When the parameters have small perturbations At that time, the evaluation result was ; parameter sensitivity Defined as: , in, This represents the absolute change in the core task evaluation results before and after parameter perturbation, reflecting the magnitude of the parameter's impact on the task results. The value represents the perturbation amplitude of the parameter; to facilitate subsequent decision-making on the number of slices, the sensitivity is normalized to the [0,1] interval.

[0036] After receiving the parameters, the regional server uses a time-series freshness-weighted method to complete regional aggregation. This method first extracts three types of time-series features of the parameters: absolute time difference reflects the time distance between the parameter and the aggregation time; periodic correlation reflects the matching degree between the parameter upload time and the key stages of the load fluctuation cycle; and update stability measures the smoothness of the change of the current parameter compared with the parameter of the previous period. After normalization, the weight of each parameter is calculated by integrating the collaborative rules of time decay, scene correlation, and stability reliability, avoiding the dominance of a single dimension. The regional parameters obtained by weighted aggregation can better fit the real-time operation requirements and reduce the interference of random fluctuations on the results.

[0037] In a preferred embodiment, the time-series freshness weighting method is specifically as follows: When the edge node uploads the encrypted and fault-tolerant updated parameter values ​​to the regional server, the regional server processes each received parameter update value. Mark a timestamp Three types of time-series features of synchronously extracted parameter update values: absolute time difference Periodic correlation and update stability The three types of time series features are mapped to the [0,1] interval, denoted as... Eliminate dimensional differences; Furthermore, a collaborative weighted formula of "basic time decay + scene association enhancement + stability credibility bonus" is constructed to avoid a single dimension dominating the weighting: , in, The basic decay coefficient controls the fundamental impact of the time dimension on the weights. These are the weighting coefficients for correlation and stability, respectively. Complete the weighted aggregation of parameter update values ​​within the region to obtain the region-level parameters. : , wherein, is the regional-level parameter update value obtained after regional aggregation; is the weight of the parameter update value of the i-th edge node; is the parameter update value of the i-th edge node generated locally; is the parameter update value of the i-th edge node generated locally; is the total number of edge nodes participating in this regional aggregation.

[0038] In the preferred embodiment, the specific calculation methods of the three types of time sequence features are as follows: , , , wherein, is the absolute time difference of the parameter update value; is the period correlation degree of the parameter update value; is the update stability of the parameter update value; is the aggregation time; is the parameter upload timestamp; is the duration of the load fluctuation period; is the start time of the load period; is the parameter update of the i-th edge node in the current period; is the parameter update value of the i-th edge node in the previous period. After the cloud aggregates the parameters of each region, it first filters the obviously abnormal parameters, and then performs global aggregation to generate an optimized model. This process ensures the privacy and security of the parameters throughout their life cycle through encryption and sharding, and improves the quality of model aggregation through time sequence weighting and hierarchical aggregation, ultimately achieving precise optimization of the global model of the virtual power plant, and providing strong support for core tasks such as load prediction and energy scheduling.

[0039] (4) Iterative feedback After the cloud completes global aggregation and generates optimized global model parameters, it first distributes these parameters to regional servers, and then distributes them to all edge nodes within the jurisdiction of the regional servers, starting the model performance verification and iterative optimization process. This process is a key closed loop that ensures the global model continuously adapts to the actual operational needs of the virtual power plant.

[0040]

[0041] ​​​After receiving the global model parameters, each edge node uses its own independent local validation set to test the model's performance. The validation set data is independent of the training data and can objectively reflect the model's prediction performance in real-world scenarios. The core of the test is to determine whether the model's prediction error exceeds a preset threshold. If the error exceeds the threshold, it indicates a mismatch between the current model and the node's local data distribution or operating state. The node will automatically readjust the temperature coefficient of dynamic feature distillation to optimize the feature distillation effect, thereby correcting the model's training direction. After the adjustment is complete, the local training process is restarted, new parameter update values ​​are generated, and uploaded to the regional server according to the previous process.

[0042] If the model prediction error meets the preset threshold, it means that the current model performance meets the local requirements of the node. The node then enters the next iteration cycle, waiting for the next global model parameter distribution and verification from the cloud. All edge nodes perform performance verification and iteration operations according to this logic until the model prediction error of all nodes stabilizes within the threshold and the model performance no longer significantly improves, i.e., the model converges. This ensures that the final global model can achieve high-quality operation on all edge nodes and continuously support the core tasks of the virtual power plant.

[0043] Example 2 Please refer to Figure 4 This embodiment 2 provides a privacy-preserving data fusion system for virtual power plants based on federated learning, including: The learning network deployment unit is used to deploy a three-tier federated learning network consisting of edge nodes, regional servers, and a cloud platform. The feature selection unit is used to build a value-risk assessment model and determine the value coefficient of features for the core tasks of the virtual power plant. The quantification and calculation of the feature identifiability entropy value are used as a risk coefficient. to form Coordinates; Introducing a load-linked temperature coefficient that is tied to the real-time operating status of the virtual power plant. and complete feature filtering; The global optimization unit is used by edge nodes to generate parameter update values ​​based on distillation features. During transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 segments based on their sensitivity, and uploaded through no less than 3 channels before fault tolerance verification is performed. Regional aggregation is completed based on the time-series freshness weighting method. In the cloud, abnormal parameters are filtered, and global aggregation is completed to generate an optimized model. The iterative feedback unit is used to distribute global model parameters from the cloud to edge nodes via regional servers. Each node then tests the model performance using an independent local validation set. If the prediction error exceeds a preset threshold, the temperature coefficient of dynamic feature distillation is automatically readjusted. And re-develop local training and parameter upload; if the performance is up to standard, enter the next round of iteration until the model converges.

[0044] Embodiment 3 The embodiment 3 also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement any step of the method for federated learning-based virtual power plant privacy protection data fusion.

[0045] The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0046] For the computer readable storage medium provided in the present application, refer to the above method embodiments, which will not be repeated here.

[0047] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A virtual power plant privacy protection data fusion method based on federated learning, characterized in that, Comprise: S1. Deploying a three-level federated learning network of edge node-area server-cloud platform; S2. Construct a value-risk assessment model and determine the value coefficients of features for the core tasks of the virtual power plant. The quantification and calculation of the feature identifiability entropy value are used as a risk coefficient. to form Coordinates; Introducing a load-linked temperature coefficient that is tied to the real-time operating status of the virtual power plant. and complete feature filtering; S3. The edge node generates parameter update values based on distilled features; during transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 segments based on their sensitivity, uploaded through not less than 3 channels, and then fault tolerance verification is completed; and based on the time sequence freshness weighting method, regional aggregation is completed; in the cloud, abnormal parameters are filtered, and global aggregation is completed to generate an optimized model; S4. After the cloud end distributes the global model parameters to the edge nodes through the regional server, each node tests the model performance with the independent local validation set. If the prediction error exceeds the preset threshold, the temperature coefficient of dynamic feature distillation is automatically adjusted and the local training and parameter uploading are restarted. If the performance meets the standard, the next round of iteration is entered, and the model converges.

2. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The three-level federated learning network in S1 is specifically: The edge node corresponds to each data holder in the virtual power plant, and the model training is carried out locally using its own data, only the model parameter update information is generated, and the original data is completely retained locally; The area server is responsible for collecting and preliminarily processing the model parameter updates of the edge nodes in the area, and performing parameter aggregation and verification in the area; The cloud platform aggregates the parameter updates uploaded by each area server, performs global model aggregation and optimization, generates the final global model, and distributes the optimized model parameters to each area server and edge node.

3. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The value coefficient in S2 The calculation method of the value coefficient is: Quantifying features with Shapley values Contribution to the virtual power plant model, for the total number of features, for the feature subset without , and for the model performance of the subset . , in, Features to be calculated The value coefficient; For subset The number of features contained therein; When only a subset is used When considering the characteristics of the virtual power plant core task model, the performance indicators exhibited by the model are as follows: To evaluate the features Add to subset Then, use These characteristics are the model's performance metrics; Total characteristic number factorial; For subset The factorial of the number of features; Total characteristic number Subtract subset size The factorial after subtracting 1; simultaneously, normalize the Shapley values ​​of all features, and... Map to the interval [0,1].

4. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The risk coefficient in S2 The calculation method is: , in, Features to be calculated The risk factor; Features The number of generalized categories to which it belongs is used to reflect the impact of the complexity of the classification scenario in which the feature is located on the risk; Features The number of discrete value categories of a feature; if it is a continuous feature, it needs to be discretized first. The number of categories after discretization reflects the diversity of the feature's own values; Features Index of discrete values; Features Take the first The probability of the value; and simultaneously for all features Normalize the value, Map to the interval [0,1].

5. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The specific process of feature selection in S2 is: Let the set of features be For any feature A sufficient and necessary condition for it to be retained is that , wherein, is a characteristic identifiability risk coefficient; is a characteristic value coefficient for the virtual power plant core task; is a load linkage temperature coefficient.

6. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The calculation method of sensitivity in S3 is: The evaluation function of the virtual power plant core task is When the parameter update value is , the task evaluation result is ; When the parameters are slightly perturbed the evaluation results are ; Parameters sensitivity of the parameters is defined as: , wherein, is the absolute change of the core task evaluation result before and after the parameter disturbance, reflecting the influence amplitude of the parameter on the task result; is the disturbance amplitude of the parameter; for the convenience of subsequent slice quantity decision, the sensitivity is normalized to the interval [0, 1].

7. The virtual power plant privacy protection data fusion method based on federated learning according to claim 1, characterized in that, The time sequence freshness weighting method in S3 is specifically: When the edge node uploads the encrypted, fault-tolerant checked parameter update value to the regional server, the regional server marks a timestamp for each received parameter update value Mark a timestamp Synchronously extract three types of timing features of the parameter update value: absolute time difference , cycle correlation degree and update stability ; And the three kinds of time sequence characteristics are mapped to [0, 1] interval, denoted as , eliminating dimensional differences; A collaborative weighting formula of "basic time decay + scene correlation enhancement + stable credibility bonus" is reconstructed to avoid single dimension dominance weighting: , wherein, is a base attenuation coefficient, controlling the base influence of the time dimension on the weight; are weight coefficients of the correlation degree and stability, respectively; The weighted aggregation of the region-in parameter update values is completed to obtain a region-level parameter : , wherein, is the regional level parameter update value obtained after regional aggregation; is the weight of the parameter update value of the i-th edge node; is the parameter update value of the i-th edge node generated locally; is the parameter update value of the i-th edge node generated locally; is the parameter update value of the i-th edge node generated locally; is the total number of edge nodes participating in this regional aggregation.

8. The virtual power plant privacy protection data fusion method based on federated learning according to claim 7, characterized in that, The specific calculation method of the three types of time sequence features is: , , , in, The absolute time difference between parameter update values; The periodic correlation of parameter update values; For the stability of parameter update values; This is the aggregation moment; Upload timestamps for parameters; The duration of the load fluctuation cycle; This is the start time of the load cycle; For the first The parameter update of each edge node in the current period; For the first The parameter update values ​​of each edge node in the previous cycle.

9. A virtual power plant privacy protection data fusion system based on federated learning, characterized in that, Comprise: A learning network deployment unit for deploying a three-level federated learning network of edge node-area server-cloud platform; The feature selection unit is used to build a value-risk assessment model and determine the value coefficient of features for the core tasks of the virtual power plant. The quantification and calculation of the feature identifiability entropy value are used as a risk coefficient. to form Coordinates; Introducing a load-linked temperature coefficient that is tied to the real-time operating status of the virtual power plant. and complete feature filtering; A global optimization unit for the edge node to generate parameter update values based on distilled features; during transmission, the parameters are first homomorphically encrypted by Paillier, then adaptively split into 3-5 segments based on their sensitivity, uploaded through not less than 3 channels, and then fault tolerance verification is completed; and based on the time sequence freshness weighting method, regional aggregation is completed; in the cloud, abnormal parameters are filtered, and global aggregation is completed to generate an optimized model; An iterative feedback unit is configured to, after the cloud-side distributes the global model parameters to the edge nodes through the regional servers, each node tests the model performance using an independent local validation set, and if the prediction error exceeds a preset threshold, automatically readjusts the temperature coefficient of the dynamic feature distillation and re-starts local training and parameter uploading. If the performance meets the standard, the next round of iteration is entered, and the model converges.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to perform the federated learning-based virtual power plant privacy protection data fusion method of any one of claims 1-8.

Citation Information

Patent Citations

  • Virtual power plant data interaction method and device based on federated learning and medium

    CN118747538A

  • Virtual power plant cloud edge collaborative optimization method and system based on improved federated learning

    CN119849654A

  • Big data privacy protection modeling method and system based on federated learning and block chain

    CN120951375A

  • Wireless service traffic prediction method based on weighted federated learning

    WO2021169577A1

Cited By

  • Streaming data privacy protection method and device based on secret state calculation

    CN122137692A