Privacy protection federal basic model continuous learning method in multi-source new energy main body collaborative scheduling scene
By introducing supervision nodes and privacy protection technologies into the renewable energy power system, the privacy leakage and communication overhead problems of renewable energy power terminal devices are solved, a balance between privacy protection and efficient training is achieved, and the security of the renewable energy power system and the efficiency of model training are improved.
Patent Information
- Application Number
- CN202510778539.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional machine learning technology cannot meet the privacy protection and efficient training requirements of massive data in renewable energy power systems. In particular, there are problems of privacy leakage and excessive communication overhead during data transmission and model training of renewable energy power terminal devices.
A privacy-preserving federated learning method is adopted in the scenario of collaborative scheduling of multiple new energy entities. By introducing regulatory nodes to dynamically evaluate the contribution of the power terminal equipment of new energy entities, differential privacy perturbation, one-time masking and noising, homomorphic encryption and streaming hierarchical clustering technology are combined to achieve data security transmission and model optimization.
While reducing communication overhead, the level of privacy protection is improved, and a balance between privacy and utility is achieved. Through technical means, the needs of privacy protection and efficient training are achieved, and the security of the new energy main power system and the training efficiency of the model are improved.
Smart Images

Figure CN120692062A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer science and information technology, and specifically relates to a federated learning method with differential privacy protection and homomorphic encryption protection, as well as a distributed computing technology involving cluster training, secure aggregation and continuous learning of new energy main power terminal equipment. Background Art
[0002] Global warming is a major global issue that needs to be addressed. Accelerating the transition away from fossil fuels and improving carbon emissions management are crucial. A power system dominated by renewable energy is a key initiative. Aiming to achieve low-carbon, clean, safe, controllable, flexible, and efficient performance, this system is committed to building an intelligent, diversified, and modern power ecosystem. With the rapid development of artificial intelligence (AI), it is playing an increasingly important role in renewable energy-based power systems. However, it requires massive amounts of data to provide diverse and rich material for model training. The large number of devices connected to the power grid generates massive amounts of data. In this context, traditional machine learning techniques still cannot meet the requirements of privacy protection and efficient training. Summary of the Invention
[0003] The purpose of this invention is to propose a new learning method based on the federated learning model, which introduces a supervision node to dynamically evaluate the contribution of the new energy subject power terminal equipment to the global index, thereby selectively executing the new energy subject power terminal equipment rejection strategy or reducing the global budget, greatly reducing the communication overhead, and further improving the level of privacy protection, achieving a good balance between privacy and utility.
[0004] The present invention is implemented by the following technical means: a privacy-preserving federated basic model continuous learning method in a multi-source new energy subject collaborative scheduling scenario, which relies on a new energy subject client and a distribution management server and includes the following steps:
[0005] The phases of data collection and privacy protection of new energy power terminal equipment, regional aggregation and parameter upload, dynamic monitoring of privacy utility, and decryption and update of distribution management system servers are executed in a loop until the final training is completed.
[0006] The data collection and privacy protection stage of the new energy subject power terminal device includes the following steps: the new energy subject power terminal device collects the power consumption data of devices such as smart meters and the power generation data of devices such as distributed photovoltaic inverters within a certain time interval, and executes the embedded function to obtain task signature data;
[0007] Perform differential privacy perturbation on the task signature data and then upload it to the regional aggregation node;
[0008] The power terminal equipment of the new energy subject uses local data to fine-tune the shared global model, and uses a one-time mask to add noise to the fine-tuning parameters, and uploads the noised results to the regional aggregation node;
[0009] The regional aggregation and parameter uploading phase specifically includes the following steps:
[0010] Perform streaming hierarchical clustering on the collected task signature data and keep the clusters stable during the cooling time;
[0011] The regional aggregation node collects the masks uploaded by the clustered new energy main power terminal devices, sums them, and then cancels the masks to obtain cluster-level updates. Finally, it is homomorphically encrypted and uploaded to the distribution management system server;
[0012] The privacy utility dynamic monitoring stage specifically includes the following steps:
[0013] The supervisory node checks the global indicators for each round and calculates the gain of the power terminal equipment of a single new energy subject. It then determines the gain of the power terminal equipment of a single new energy subject to determine whether it can participate in subsequent training. It also scores the power terminal equipment of a single new energy subject that can participate in subsequent training and determines the priority of the equipment's participation.
[0014] The decryption and update phase of the power distribution management system server specifically includes the following steps:
[0015] The distribution management system server calculates the weights and then distributes them. Based on the weights, the distribution management system server performs weighted aggregation on the cluster-level updates encrypted from the regional aggregation nodes in the ciphertext domain to obtain the aggregated global update. Finally, the distribution management system server decrypts the aggregated global update and updates the federated learning basic model.
[0016] The task signature data is specifically:
[0017]
[0018] s i represents the encoded feature vector used to represent the task signature after the embedded function is executed, f enc is an encoding function that maps the original input features into a vector of fixed dimension, FFT(p i ) represents the original power time series characteristic data p i Perform fast Fourier transform to extract the frequency domain features of the time series data.
[0019] The specific formula for performing differential privacy perturbation on task signature data is as follows:
[0020]
[0021] in is the result after adding noise, which satisfies ((ε,δ))-DP, s i It is the signature data of the produced task. The mean is 0, and the covariance matrix is Gaussian noise, I d is the d-dimensional identity matrix, σ 2 is the variance of the Gaussian distribution, indicating the intensity of the noise, δ indicates the allowed failure probability, and ε is the privacy budget, which measures the degree of privacy leakage.
[0022] The calculation formula for the fine-tuning increment of the global shared model is as follows:
[0023]
[0024] Where W is the global shared model, It is the time window data of the new energy main power terminal device i. LoRA is a lightweight parameter structure used for local tuning. ΔA i It is the model fine-tuning increment learned locally by the new energy entity power terminal device i;
[0025] The specific formula for adding noise using a one-time mask is as follows
[0026]
[0027] Among them (r i,out ,r i,in ) is the coordinated access mask pair between the new energy main power terminal equipment, satisfying the mutual offset relationship, is the noise increment result;
[0028] A disposable mask satisfies:
[0029]
[0030] Among them br i,out ,br i,in Indicates that the mask vector is bitwise added or subtracted according to each dimension component, (r i,out ,r i,in )The mask follows a uniform distribution and the range is (-λ,λ).
[0031] The cluster merging criterion for performing streaming hierarchical clustering on the task signature data is:
[0032]
[0033] in is the task signature vector of the two new energy main power terminal devices, τ is the preset similarity threshold, d cosUsed to represent the similarity of the task signature vectors of the power terminal equipment of two new energy entities;
[0034] The intra-cluster similarity is determined by Euclidean distance. When the variance is greater than a specified threshold, the cluster is split. The specific formula is as follows:
[0035]
[0036] in For a cluster, is the task signature vector of each new energy main power terminal device in the cluster, is the cluster center, i.e., the mean vector, and η is the maximum tolerable threshold of the variance.
[0037] The cluster-level update is the model fine-tuning increment after noise addition uploaded by each new energy main power terminal device in the cluster. The sum of
[0038]
[0039] The homomorphic encryption is CKKS encryption.
[0040] The formula for the comprehensive score of power terminal equipment of a single new energy entity is:
[0041]
[0042] Among them, S i reflects the comprehensive score, μ is the weight parameter, f b It is the benchmark dispatch frequency of the power terminal equipment of the single new energy subject, which is set by the supervision node. sum is the historical dispatch frequency of the power terminal equipment of the single new energy entity, Q i Reflects the data quality score, ∈ i is the remaining privacy budget of the single new energy entity's power terminal equipment, ∈ max is the maximum privacy budget allowed by the supervisory node, γ is a preset parameter, Reflects the scheduling frequency, Reflects the privacy level.
[0043] The specific formula for weighted aggregation is as follows:
[0044]
[0045] where ω k is the weight, is the global update after aggregation,
[0046]
[0047] where η k is the degree of conflict, β is the coefficient that controls the steepness of the weight,
[0048]
[0049] in is the anchor soft label vector of the K-th cluster, representing the output of the cluster model for some standard samples, is the average soft label vector of all clusters, representing the common direction.
[0050] The present invention designs a clustering method for new energy main body power terminal equipment, performs streaming hierarchical clustering on the original task signature data uploaded by the new energy main body power terminal equipment, and continuously adjusts the clustering center to update the local task objectives towards the global optimal direction; provides multiple privacy protection technologies during data transmission, including differential privacy perturbation, one-time mask noise and homomorphic encryption technology to ensure data privacy security; at the same time, introduces a supervision audit node, which will dynamically adjust the status of the new energy main body power terminal equipment participating in the training and the global budget according to the contribution level of the new energy main body power terminal equipment to the global indicators, while reducing communication overhead and achieving a good balance between privacy and model utility. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of the continuous learning method of the privacy-preserving federated basic model in the scenario of collaborative scheduling of multiple new energy entities proposed by the present invention.
[0052] Figure 2 This is a cross-scenario scheduling flowchart provided by the present invention.
[0053] Figure 3 This is a framework diagram of the continuous learning method of the privacy-preserving federated basic model in the scenario of collaborative scheduling of multiple new energy entities proposed in the present invention. DETAILED DESCRIPTION
[0054] The present invention will be further described below in conjunction with the accompanying drawings:
[0055] The present invention proposes a privacy-preserving federated basic model continuous learning method in a multi-source new energy subject collaborative scheduling scenario. The method consists of a system architecture consisting of four parts: new energy subject power terminal equipment, regional aggregation nodes, distribution management system servers and supervision nodes. The new energy subject power terminal equipment is responsible for collecting power data and reporting the data to the regional aggregation node after encryption. The power data includes electricity consumption and power generation data such as distributed photovoltaic inverters, energy storage EMS, smart meters or charging piles. The regional aggregation node is responsible for aggregating and reporting the power data. The supervision and audit node is responsible for real-time monitoring of the gain and privacy budget dynamic allocation of new energy subject power terminal equipment, maintaining the stability of global model performance while ensuring data privacy compliance. The distribution management system server is responsible for final aggregation, model update, and power forecasting and scheduling optimization.
[0056] The learning method described in this paper consists of four phases: data collection and privacy protection for renewable energy power terminal devices, regional aggregation and parameter upload, dynamic monitoring of privacy utility, and decryption and update of the distribution management system server. Each phase is executed repeatedly until the training is complete.
[0057] The data collection and privacy protection phase of the new energy subject power terminal equipment includes the following steps:
[0058] Step 1-1: The new energy main power terminal device collects the electricity consumption data of smart meters and other devices and the power generation data of distributed photovoltaic inverters and other devices within a certain time interval, and executes the embedded function to obtain the task signature data. The specific steps include the following:
[0059] Step 1-1-1: The main power terminal equipment of new energy sources collects raw data. The main power terminal equipment of new energy sources (c i ) Local time window The original data in is:
[0060]
[0061] where x i,t Refers to the data record of the new energy main power terminal equipment in a certain period of time (1~T), which includes three parts: pi,t , m i,t and d i , respectively representing the original power time series characteristics (such as power, current, voltage, etc.), model context information (such as prediction task objectives, control parameters, etc.), and new energy main power terminal equipment attribute information (such as equipment type, geographical location, etc.). It refers to the local time window data set formed by collecting all data records of the new energy entity power terminal device i within a certain period of time.
[0062] Step 1-1-2: Execute the embedding function to obtain the following formula:
[0063]
[0064] where s i Represents the encoded feature vector obtained after the embedding function is executed, which is used to represent the task signature. enc Is an encoding function that maps the original input features into a vector of fixed dimension. FFT(p i ) represents the original power time series characteristic data p i Perform fast Fourier transform to extract the frequency domain features of the time series data.
[0065] Step 1-2: Perform differential privacy perturbation on the task signature data and then upload it to the regional aggregation node to ensure data security during transmission. The specific steps include the following:
[0066] Step 1-2-1: The new energy entity's power terminal device performs differential privacy perturbation on the task signature. The specific formula is as follows:
[0067]
[0068] in is the result after adding noise, which satisfies ((ε,δ))-DP, s i It is the signature data of the produced task. The mean is 0, and the covariance matrix is Gaussian noise, I d is the identity matrix of dimension d. 2 δ is the variance of the Gaussian distribution, which indicates the intensity of noise. A larger δ indicates stronger privacy protection. δ represents the acceptable failure probability, indicating that privacy protection may fail under very small probabilities. It is generally taken to be a very small value. ε is the privacy budget, which measures the degree of privacy leakage. A smaller ε indicates stronger privacy.
[0069] Step 1-2-2: Upload the task signature data after differential privacy perturbation to the regional aggregation node h k .
[0070] Steps 1-3: To achieve continuous learning, the new energy main power terminal equipment uses local data to fine-tune the shared global model, uses a one-time mask to add noise to the fine-tuning parameters, and uploads the noised results to the regional aggregation node. The specific steps include the following:
[0071] Step 1-3-1: Fine-tune the low-rank adapter / prompt parameters of the new energy main power terminal equipment. The specific formula is as follows:
[0072]
[0073] Where W is the global shared model, It is the time window data of the new energy main power terminal device i. LoRA is a lightweight parameter structure used for local tuning. Using LoRA can reduce the number of communication parameters, thereby reducing communication overhead, and is especially suitable for edge devices. i It is the model fine-tuning increment learned locally by the new energy entity power terminal device i.
[0074] Step 1-3-2: ΔA i Use a one-time mask (r i,out ,r i,in ) forms the noise increment, and the specific formula is as follows:
[0075]
[0076] Among them (r i,out ,r i,in ) is the coordinated access mask pair between the new energy main power terminal equipment, satisfying the mutual offset relationship, is the noise increment result.
[0077] A disposable mask satisfies:
[0078]
[0079] Among them br i,out ,br i,in Indicates that the mask vector is bitwise added or subtracted according to each dimension component, (r i,out ,r i,in )The mask follows a uniform distribution and the range is (-λ,λ).
[0080] Each round will renegotiate the mask key pair (r i,out ,r i,in ), which ensures that the aggregation node cannot reconstruct ΔA i , and only aggregate results can be obtained.
[0081] Finally Upload to regional aggregation node h k .
[0082] The regional aggregation and parameter uploading phase mainly includes the following steps:
[0083] Step 2-1: Perform streaming hierarchical clustering on the collected task signature data and keep the clusters stable during the cooling time. Specifically, dynamically cluster the task signature stream data, allowing similar tasks to form clusters, set a cooling time to keep the clusters stable, and then obtain the cluster set after clustering. After the cluster set is obtained, each new energy main power terminal device in the cluster uploads the model fine-tuning increment after masking and noise to the regional aggregation node h k , h k The model fine-tuning increments of all new energy main power terminal devices in the cluster are aggregated to obtain cluster-level updates, and the distribution management system server aggregates the cluster-level updates to obtain global updates, thereby updating the global model.
[0084] Step 2-1-1: The regional aggregation node groups the new energy main power terminal devices with similar tasks into one cluster using a streaming hierarchical clustering method. The merging criteria of the streaming hierarchical clustering method are:
[0085]
[0086] in is the task signature vector of the two new energy main power terminal devices, τ is the preset similarity threshold, d cos Used to represent the similarity of the task signature vectors of the power terminal equipment of two new energy entities.
[0087] Step 2-1-2: In subsequent training, the cluster centers are constantly adjusted. If the tasks of some new energy main power terminal devices in a cluster are significantly different from those of other new energy main power terminal devices, they need to be split out. The criterion for splitting is the intra-class variance:
[0088]
[0089] in For a cluster, is the task signature vector of each new energy main power terminal device in the cluster, is the cluster center, i.e., the mean vector, and η is the maximum tolerable threshold of the variance.
[0090] The present invention determines intra-cluster similarity by calculating the Euclidean distance. When the task signature vectors of the renewable energy-based power terminal devices in a cluster deviate from the central vector, i.e., the variance exceeds the threshold η, it indicates that the tasks of the renewable energy-based power terminal devices in this cluster are too different and need to be split.
[0091] Step 2-2: To achieve secure aggregation and continuous learning, the regional aggregation node collects the masks uploaded by the clustered new energy main power terminal devices, sums them, and then cancels the masks to obtain cluster-level updates. Finally, it is homomorphically encrypted and uploaded to the distribution management system server. The specific steps are as follows:
[0092] Step 2-2-1: Each new energy main power terminal device in the cluster uploads the model fine-tuning increment after adding noise To sum it up, the formula is as follows:
[0093]
[0094] During the summation process, the mask pair (r i,out ,r i,in ) to offset and obtain cluster-level update.
[0095] Step 2-2-2: There is also a differentially private synthetic data generator It is mainly used to generate differential private synthetic samples The sample and historical model are used to perform knowledge distillation to suppress catastrophic forgetting. In this way, when the new model learns new data, its performance on the synthesized old data remains consistent, retaining the predictive power of the old knowledge.
[0096] Step 2-2-3: Regional aggregation node h k Will Perform CKKS encryption. The specific formula is as follows:
[0097]
[0098] CKKS is a homomorphic encryption algorithm that allows encryption in a ciphertext state without decryption, which can perform complex calculations while ensuring data privacy.
[0099] Step 2-2-4: Encrypt the result Upload to the power distribution management system server.
[0100] The privacy utility dynamic monitoring phase mainly includes the following steps:
[0101] Step 3-1: The supervisory node checks the global indicators in each round. When the individual gain of a new energy entity's power terminal device is insufficient, it will be denied participation in subsequent training or the global budget will be adjusted to prevent its information leakage. The specific steps are as follows:
[0102] Step 3-1-1: Supervisory Node (R g ) for each round of global indicators (l p ,l u )examine:
[0103] C1≤l p +C2l u
[0104] where l p is the privacy loss indicator of the model, represented by privacy budget ∈, l u This is the utility loss indicator for the model. If the model accuracy decreases, the loss increases, and the model performance deteriorates. C1 and C2 are two weight parameters used to balance privacy and model utility.
[0105] Step 3-1-2: The supervisory node calculates the gain of the power terminal equipment of a single new energy entity, which is expressed as:
[0106]
[0107] where ΔAcc i It indicates the improvement of the accuracy of the model on the validation set before and after the participation of the new energy subject power terminal equipment i, that is, the positive contribution of the new energy subject power terminal equipment, represents the privacy budget consumption of the new energy subject’s power terminal device i, and ψ is the set threshold.
[0108] Step 3-1-3: The supervisory node judges the gain of the power terminal equipment of a single new energy subject. When the gain of the power terminal equipment of a single new energy subject is g i When the value is less than the set threshold, the supervisory node issues a rejection instruction to restrict the new energy main power terminal equipment from participating in subsequent training, or adjust the global budget ε tot , by reducing the global privacy budget, the gain of the power terminal equipment of the new energy subject can be improved. The supervisory node describes the global privacy-utility multi-objective optimization problem as follows:
[0109]
[0110] Where α∈[0,1] is dynamically adjusted according to real-time power dispatch risks and regulatory requirements. tot is the upper limit of the global privacy budget.
[0111] The present invention is implemented using Sequential Quadratic Programming (SQP) and a step-by-step elimination strategy for new energy subject power terminal devices. SQP is a method for efficiently solving nonlinear constrained optimization problems. Each step approximately solves a constrained quadratic programming sub-problem. The elimination strategy for new energy subject power terminal devices is reflected in the fact that if the privacy budget consumed by a certain new energy subject power terminal device increases sharply but its utility loss improves little, the present invention will eliminate it from training. When the gain of a single new energy subject power terminal device is less than the threshold, the supervisory node can refuse the new energy subject power terminal device from continuing to participate in the training.
[0112] The supervisory node constructs a multi-objective optimization problem weighted by utility loss and privacy loss to ensure the overall privacy budget ε tot On the premise of not being broken through, SQP optimization and the mechanism of gradually eliminating the main power terminal equipment of new energy entities are adopted to dynamically balance model performance and privacy compliance, and realize robust, efficient, and regulatory-level federated learning scheduling.
[0113] Step 3-2: Build a "multi-scenario scheduling and orchestration engine" to enable supervisory nodes to perceive task scenarios, dynamically combine scheduling processes based on task types, and introduce a cross-scenario data sharing mechanism to achieve data sharing.
[0114] Step 3-2-1: Build a "multi-scenario scheduling and orchestration engine" to monitor the task scenario and dynamically combine the scheduling process according to the task type. Dynamically select appropriate strategies based on different scenarios and achieve rapid adaptation by changing the strategy layer configuration. The strategy layer configuration includes participating nodes, scheduling frequency, asynchronous or not, etc., which are dynamically set according to the complexity and confidentiality of the scenario. The formula for evaluating the comprehensive score of the power terminal equipment of a single new energy entity is:
[0115]
[0116] Among them, S i reflects the comprehensive score, μ is the weight parameter, f b It is the benchmark dispatch frequency of the power terminal equipment of the single new energy subject, which is set by the supervision node. sum is the historical dispatch frequency of the power terminal equipment of the single new energy entity, Q i Reflects the data quality score, ∈ i is the remaining privacy budget of the single new energy entity's power terminal equipment, ∈ max is the maximum privacy budget allowed by the supervisory node, and γ is a preset parameter. Reflects the scheduling frequency, Reflects the privacy level. The supervisory node dynamically sets the data quality of the power terminal equipment of a single new energy entity based on the scenario, evaluates the privacy level, dispatch frequency, calculates the comprehensive score, and dynamically selects participating nodes.
[0117] Step 3-2-2: Introduce a cross-scenario data sharing mechanism to achieve data sharing. The supervisory node maintains a shared security zone responsible for storing shared data, categorizing it according to different scenarios. Cross-scenario data is uploaded to the shared zone using homomorphic encryption technology to achieve data sharing.
[0118] The power distribution management system server decryption and update phase mainly includes the following steps:
[0119] Step 4-1: The distribution management system server performs weighted aggregation on the encrypted cluster-level updates in the ciphertext domain to obtain the global update, decrypts it, and then updates the federated learning basic model. The specific steps are as follows:
[0120] Step 4-1-1: The distribution management system server calculates the weight. The weight is calculated by the distribution management system server for each cluster uploaded anchor soft label Calculate the degree of conflict:
[0121]
[0122] in is the anchor soft label vector of the K-th cluster, representing the output of the cluster model for some standard samples, is the average soft label vector of all clusters, representing the common direction. k is the calculated conflict degree.
[0123] Then the aggregation weight is assigned according to the following formula:
[0124]
[0125] where η k is the degree of conflict calculated in the previous step, and β is the coefficient that controls the steepness of the weight.
[0126] Step 4-1-2: After receiving the encrypted cluster-level update from the regional aggregation node, the distribution management system server performs weighted aggregation in the ciphertext domain. The specific formula is as follows:
[0127]
[0128] where ω k is the weight calculated in the previous step, It is the global update after aggregation.
[0129] The present invention calculates and The cosine similarity of each cluster is used to reflect the similarity between each cluster and the whole. If the cluster has a higher similarity with the whole, the conflict degree of the cluster is lower and the cluster task is more consistent with the global one. In this case, the present invention assigns a larger weight to it, and vice versa. This method can effectively reduce the impact of noise clusters on the performance of the global model.
[0130] Step 4-1-3: Finally, the power distribution management system server After decryption, the federated learning basic model (W) is updated.
[0131] This paper proposes a continuous learning method for a privacy-preserving federated basic model in a multi-source renewable energy collaborative scheduling scenario. This method can ultimately train a model that balances privacy and utility in a multi-party environment. Specifically, regional aggregation nodes cluster renewable energy power terminal devices based on their task signatures, grouping them into clusters with similar tasks for local training and aggregation. The cluster centers are then continuously adjusted, thereby improving model training efficiency. Regarding security, the system incorporates multiple security technologies, including differential privacy perturbation, one-time masked noisy, and homomorphic encryption, to ensure data privacy during transmission. Furthermore, the system incorporates a supervisory node that monitors the contribution of renewable energy power terminal devices to global indicators and implements different strategies, effectively improving privacy protection while significantly reducing communication overhead. Finally, the model trained by the distribution management system server demonstrates excellent performance and can make reasonable predictions based on electricity consumption and generation data uploaded by renewable energy power terminal devices, thereby adaptively optimizing the power structure and distribution level. This system can be widely applied in intelligent scenarios involving renewable energy power systems, such as distributed renewable energy forecasting, load forecasting, demand response, and distribution automation.
Claims
1. A privacy-preserving federated basic model continuous learning method in a multi-source new energy subject collaborative scheduling scenario, the method is implemented by a new energy subject client and a distribution management server, and is characterized by: It includes the phase of data collection and privacy protection of new energy main power terminal equipment, the phase of regional aggregation and parameter upload, the phase of dynamic monitoring of privacy utility, and the phase of decryption and update of the distribution management system server. The above phases are executed cyclically until the final training is completed. The data collection and privacy protection stage of the new energy subject power terminal device includes the following steps: the new energy subject power terminal device collects the power consumption data of devices such as smart meters and the power generation data of devices such as distributed photovoltaic inverters within a certain time interval, and executes the embedded function to obtain task signature data; Perform differential privacy perturbation on the task signature data and then upload it to the regional aggregation node; The power terminal equipment of the new energy subject uses local data to fine-tune the shared global model, and uses a one-time mask to add noise to the fine-tuning parameters, and uploads the noised results to the regional aggregation node; The regional aggregation and parameter uploading phase specifically includes the following steps: Perform streaming hierarchical clustering on the collected task signature data and keep the clusters stable during the cooling time; The regional aggregation node collects the masks uploaded by the clustered new energy main power terminal devices, sums them, and then cancels the masks to obtain cluster-level updates. Finally, it is homomorphically encrypted and uploaded to the distribution management system server; The privacy utility dynamic monitoring stage specifically includes the following steps: The supervisory node checks the global indicators for each round and calculates the gain of the power terminal equipment of a single new energy subject. It then determines the gain of the power terminal equipment of a single new energy subject to determine whether it can participate in subsequent training. It also scores the power terminal equipment of a single new energy subject that can participate in subsequent training and determines the priority of the equipment's participation. The decryption and update phase of the power distribution management system server specifically includes the following steps: The distribution management system server calculates the weights and then distributes them. Based on the weights, the distribution management system server performs weighted aggregation on the cluster-level updates encrypted from the regional aggregation nodes in the ciphertext domain to obtain the aggregated global update. Finally, the distribution management system server decrypts the aggregated global update and updates the federated learning basic model.
2. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The task signature data is specifically: s i represents the encoded feature vector used to represent the task signature after the embedded function is executed, f enc is an encoding function that maps the original input features into a vector of fixed dimension, FFT(p i ) represents the original power time series characteristic data p i Perform fast Fourier transform to extract the frequency domain features of the time series data.
3. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The specific formula for performing differential privacy perturbation on task signature data is as follows: in is the result after adding noise, which satisfies ((ε,δ))-DP, s i It is the signature data of the produced task. The mean is 0, and the covariance matrix is Gaussian noise, I d is the d-dimensional identity matrix, σ 2 is the variance of the Gaussian distribution, indicating the intensity of the noise, δ indicates the allowed failure probability, and ε is the privacy budget, which measures the degree of privacy leakage.
4. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The calculation formula for the fine-tuning increment of the global shared model is as follows: Where W is the global shared model, It is the time window data of the new energy main power terminal device i. LoRA is a lightweight parameter structure used for local tuning. ΔA i It is the model fine-tuning increment learned locally by the new energy entity power terminal device i; The specific formula for adding noise using a one-time mask is as follows Among them (r i,out ,r i,in ) is the coordinated access mask pair between the new energy main power terminal equipment, satisfying the mutual offset relationship, is the noise increment result; A disposable mask satisfies: Among them br i,out ,br i,in Indicates that the mask vector is bitwise added or subtracted according to each dimension component, (r i,out ,r i,in )The mask follows a uniform distribution and the range is (-λ,λ).
5. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The cluster merging criterion for performing streaming hierarchical clustering on the task signature data is: in is the task signature vector of the two new energy main power terminal devices, τ is the preset similarity threshold, d cos Used to represent the similarity of the task signature vectors of the power terminal equipment of two new energy entities; The intra-cluster similarity is determined by Euclidean distance. When the variance is greater than a specified threshold, the cluster is split. The specific formula is as follows: in For a cluster, is the task signature vector of each new energy main power terminal device in the cluster, is the cluster center, i.e., the mean vector, and η is the maximum tolerable threshold of the variance.
6. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The cluster-level update is the model fine-tuning increment after noise addition uploaded by each new energy main power terminal device in the cluster. The sum of The homomorphic encryption is CKKS encryption.
7. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The formula for the comprehensive score of power terminal equipment of a single new energy entity is: Among them, S i reflects the comprehensive score, μ is the weight parameter, f b It is the benchmark dispatch frequency of the power terminal equipment of the single new energy subject, which is set by the supervision node. sum is the historical dispatch frequency of the power terminal equipment of the single new energy entity, Q i Reflects the data quality score, ∈ i is the remaining privacy budget of the single new energy entity's power terminal equipment, ∈ max is the maximum privacy budget allowed by the supervisory node, γ is a preset parameter, Reflects the scheduling frequency, Reflects the privacy level.
8. The privacy-preserving federated basic model continuous learning method in the scenario of multi-source new energy subject collaborative scheduling according to claim 1 is characterized in that: The specific formula for weighted aggregation is as follows: where ω k is the weight, is the global update after aggregation, where η k is the degree of conflict, β is the coefficient that controls the steepness of the weight, in is the anchor soft label vector of the K-th cluster, representing the output of the cluster model for some standard samples, is the average soft label vector of all clusters, representing the common direction.
Citation Information
Cited By
Big data privacy computing method and system based on cloud computing, electronic equipment and storage medium
CN120880800A
A cloud computing-based big data privacy computing method and system, electronic equipment and storage medium
CN120880800B
Federal learning and differential privacy protection-based carbon emission prediction method and system
CN120910914A
Intelligent analysis and identification method for resident travel characteristics based on multi-source data fusion
CN121167328A
Edge side load prediction method and device based on federated learning
CN121706932A