A Privacy-Enhanced Aggregation Method and System for Multi-Source Trusted Federated Learning
This paper proposes a privacy-enhancing aggregation method for multi-source trusted federated learning, which solves the privacy protection and Byzantine attack problems of multi-source data aggregation in smart grids. It achieves in-depth protection of data privacy and improves model stability, and is applicable to fields such as smart grids.
Patent Information
- Application Number
- CN202511082906.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-04
AI Technical Summary
In smart grids, multi-source data aggregation technology faces challenges such as the urgent need for user data privacy protection, significant risks of Byzantine attacks, and low efficiency of cross-regional multi-party collaboration. Traditional methods struggle to balance privacy protection with modeling efficiency.
We adopt a privacy-enhancing aggregation method for multi-source trusted federated learning. By performing privacy protection processing on the original gradient, cluster identity obfuscation, abnormal feature extraction and robust aggregation, combined with the Byzantine robust aggregation algorithm, we achieve data privacy protection and model stability.
It effectively reduces the risk of privacy leaks in electricity consumption data, enhances the ability to resist Byzantine attacks, ensures the stability and credibility of the aggregation process, strengthens the decentralized characteristics, and improves the fitting ability of the global model.
Smart Images

Figure CN120579220B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of federated learning and privacy computing, and in particular to a privacy-enhancing aggregation method and system for multi-source trusted federated learning. Background Technology
[0002] In the development of smart grids, data needs to be collected from multiple independent data sources, such as residential electrical equipment, industrial systems, and local substations, to support key decisions such as grid load forecasting and energy dispatch. These distributed data sources achieve collaborative modeling through a federated learning framework, improving modeling efficiency while avoiding the sharing of raw data.
[0003] However, current multi-source data aggregation technologies face several challenges when dealing with complex application scenarios such as power systems: First, there is an urgent need for user data privacy protection: household electricity consumption curves, industrial load data, etc., contain sensitive behavioral characteristics, and traditional direct aggregation methods may lead to the leakage of privacy information such as personal work and rest habits and enterprise production processes; Second, the risk of Byzantine attacks is significant: under the federated learning framework with multi-party participation, there is a risk of Byzantine attacks. Malicious third-party institutions may disguise themselves as legitimate data sources, submit forged model updates or interfere with data, mislead global model training, and affect the accuracy of power grid dispatching decisions; Third, cross-regional multi-party collaboration is inefficient: there is a contradiction between the characteristics of local data storage and the needs of collaborative modeling, and traditional centralized aggregation models struggle to balance data privacy and modeling efficiency, often falling into a dilemma where privacy protection and modeling efficiency cannot be simultaneously achieved.
[0004] Therefore, there is an urgent need to build an innovative and efficient trusted federated learning architecture to achieve secure, reliable, and efficient multi-source data aggregation. Summary of the Invention
[0005] To address the challenges faced by existing multi-source data aggregation technologies in complex application scenarios such as power systems, this invention provides a privacy-enhanced aggregation method and system for multi-source trusted federated learning.
[0006] In a first aspect, embodiments of the present invention provide a privacy-enhancing aggregation method for multi-source trusted federated learning, comprising:
[0007] The original gradient of each client is obtained by training based on the local power consumption data of each client, and privacy protection processing is performed on each original gradient to obtain the reported gradient of the corresponding client.
[0008] Cluster identity obfuscation is performed on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and the true and false update ratio is adjusted for each multi-cluster perturbation update set to obtain the anonymized upload data of the interest cluster corresponding to the client.
[0009] Anomaly features are extracted from the anonymized uploaded data of the interest clusters of each client to obtain high deviation update seeds, and the high deviation update seeds are subjected to derivation perturbation synthesis to obtain a synthetic client update set;
[0010] Robust aggregation is performed on the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain the global model update;
[0011] Based on the parsing of the global model update by each client, the multi-cluster association data corresponding to the client is obtained, and the model reconstruction solution is performed on each of the multi-cluster association data to obtain the real cluster model parameters of the corresponding client. The real cluster model parameters are used to support the corresponding client to perform the next round of local training.
[0012] Preferably, the step of training the original gradient corresponding to each client based on the local power consumption data of each client, and performing privacy protection processing on each original gradient to obtain the reported gradient corresponding to the client, includes:
[0013] The backpropagation algorithm is used to train the local power consumption data of each client to obtain the original gradient of the corresponding client.
[0014] For each original gradient, several layer gradient vectors corresponding to the original gradient are obtained by layer decomposition, and orthogonal complement subspaces are constructed for each layer gradient vector to obtain a perturbation candidate set for the gradient vector of the corresponding layer.
[0015] For each of the perturbation candidate sets, a local loss evaluation is performed to obtain the loss value corresponding to the perturbation candidate set. Then, the loss value of each of the perturbation candidate sets is subjected to cold Bayesian posterior filtering to obtain the reporting gradient corresponding to the client.
[0016] Preferably, the step of performing cluster identity obfuscation processing on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and adjusting the true / false update ratio of each multi-cluster perturbation update set to obtain anonymized upload data of the interest cluster corresponding to the client, includes:
[0017] For each client's reported gradient, a real cluster identity matching is performed to obtain the real cluster perturbation update corresponding to the client. Then, a fake cluster spoofing processing is performed on each real cluster perturbation update to obtain a multi-cluster perturbation update set corresponding to the client. The multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates.
[0018] Based on a preset true / false update ratio, each multi-cluster perturbation update set is filtered and combined to obtain a corresponding ratio-adapted update combination. Then, each ratio-adapted update combination is associated with anonymized cluster identities to obtain the anonymized upload data of the interest clusters of the corresponding client.
[0019] Preferably, the step of extracting abnormal features from the anonymized uploaded data of the interest clusters of each client to obtain a high-deviation update seed, and then performing derivative perturbation synthesis on the high-deviation update seed to obtain a synthetic client update set, includes:
[0020] For each client, the anonymized uploaded data of the interest cluster is evaluated for deviation to obtain an updated deviation distribution, and the updated deviation distribution is screened for extreme values to obtain high deviation update seeds;
[0021] The high-deviation update seed is structurally replicated and slightly perturbed to obtain a synthetic client update set.
[0022] Preferably, the robust aggregation of the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain the global model update includes:
[0023] The Byzantine-Lupin aggregation algorithm is used to aggregate the synthetic client update set and the anonymized uploaded data of the interest cluster of each client to obtain a global model update resistant to Byzantine attacks.
[0024] Preferably, the step of obtaining multi-cluster association data corresponding to each client based on the parsing of the global model update by each client, and performing model reconstruction and solving on each of the multi-cluster association data to obtain the real cluster model parameters corresponding to the client, includes:
[0025] Each client performs cluster identity matching on the received global model update to obtain an associated cluster model set, and performs model decomposition on the associated cluster model set to obtain multi-cluster associated data corresponding to the client.
[0026] For each of the multi-cluster associated data, a mapping relationship model is performed to obtain the mapping matrix corresponding to the client. The least squares method is then used to solve the parameters of each mapping matrix to obtain the actual cluster model parameters corresponding to the client.
[0027] Secondly, embodiments of the present invention provide a privacy-enhancing aggregation system for multi-source trusted federated learning, comprising:
[0028] The gradient reporting module is used to train the original gradient corresponding to each client based on the local power consumption data of each client, and to perform privacy protection processing on each original gradient to obtain the reported gradient corresponding to the client.
[0029] The data upload determination module is used to perform cluster identity obfuscation processing on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and to adjust the true and false update ratio of each multi-cluster perturbation update set to obtain anonymized upload data of the interest cluster corresponding to the client.
[0030] The synthesis update determination module is used to extract abnormal features from the anonymized uploaded data of the interest cluster of each client to obtain a high deviation update seed, and to perform derivative perturbation synthesis on the high deviation update seed to obtain a synthesis client update set.
[0031] The robust aggregation module is used to robustly aggregate the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain a global model update;
[0032] The model parameter determination module is used to obtain the multi-cluster association data corresponding to each client based on the parsing of the global model update of each client, and to perform model reconstruction and solving on each multi-cluster association data to obtain the real cluster model parameters corresponding to the client. Each real cluster model parameter is used to support the next round of local training for the corresponding client.
[0033] Preferably, the gradient reporting module includes:
[0034] The original gradient determination unit is used to train the local power consumption data of each client using the backpropagation algorithm to obtain the original gradient of the corresponding client.
[0035] The perturbation candidate determination unit is used to decompose each original gradient by layer to obtain several layer gradient vectors corresponding to the original gradient, and to construct an orthogonal complement subspace for each layer gradient vector to obtain a perturbation candidate set for the corresponding layer gradient vector.
[0036] The posterior filtering unit is used to perform local loss evaluation on each of the perturbation candidate sets to obtain the loss value corresponding to the perturbation candidate set, and to perform cold Bayesian posterior filtering on the loss value of each of the perturbation candidate sets to obtain the reporting gradient corresponding to the client.
[0037] Preferably, the uploaded data determination module includes:
[0038] The perturbation update determination unit is used to perform real cluster identity matching on the reported gradient of each client to obtain the real cluster perturbation update corresponding to the client, and to perform fake cluster spoofing processing on each real cluster perturbation update to obtain a multi-cluster perturbation update set corresponding to the client, wherein the multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates;
[0039] The matching anonymization association unit is used to filter and combine each of the multi-cluster perturbation update sets based on a preset true and false update ratio to obtain the corresponding matching adaptation update combination, and to perform anonymous cluster identity association on each of the matching adaptation update combinations to obtain the anonymized upload data of the interest cluster of the corresponding client.
[0040] Preferably, the synthesis update determination module includes:
[0041] The seed determination unit is used to evaluate the deviation of the anonymized uploaded data of the interest cluster of each client to obtain the updated deviation distribution, and to perform extreme value screening on the updated deviation distribution to obtain high deviation update seeds.
[0042] The seed expansion unit is used to perform structural replication and slight perturbation on the high-deviation update seed to obtain a synthetic client update set.
[0043] Compared with existing technologies, the privacy-enhancing aggregation method and system for multi-source trusted federated learning, as described in this invention, offers the following advantages: By protecting the privacy of the original gradients and obfuscating cluster identities, it masks the characteristics of local electricity consumption data from the client while significantly reducing the risk of reverse inference of the client's true electricity consumption preferences through multi-cluster perturbation updates and true / false ratio control, thus achieving deep protection of electricity consumption data privacy. Through abnormal feature extraction and derived perturbation synthesis, combined with a robust aggregation mechanism, it effectively filters abnormal update interference from malicious clients, significantly improving the global model update's resistance to Byzantine attacks and ensuring the stability and credibility of the aggregation process. The client autonomously reconstructs the true cluster model based on the global model update, without relying on the server for local model reconstruction, thus strengthening the decentralized nature of the federated intelligent system and avoiding the privacy leakage risk caused by secondary transmission of model parameters. This invention is particularly suitable for multi-source heterogeneous electricity consumption data scenarios. While ensuring the privacy and security of each client's data, it enhances the global model's ability to fit complex electricity consumption patterns through federated collaborative learning, balancing privacy protection, model performance, and training efficiency, providing reliable technical support for the application of multi-source trusted federated learning in fields such as smart grids. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating a privacy-enhancing aggregation method for multi-source trusted federated learning according to an embodiment of the present invention.
[0045] Figure 2 This is a schematic diagram of the process for determining the reporting gradient in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the process for determining the uploaded data according to an embodiment of the present invention;
[0047] Figure 4 This is a schematic diagram of the process for determining the synthetic client update set in an embodiment of the present invention;
[0048] Figure 5 This is a flowchart illustrating the process of determining the parameters of the real cluster model in an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the structure of a privacy-enhancing aggregation system for multi-source trusted federated learning according to an embodiment of the present invention;
[0050] Figure label:
[0051] 1. Gradient reporting module; 2. Data upload module; 3. Synthesis and update module; 4. Lubang aggregation module; 5. Model parameter determination module. Detailed Implementation
[0052] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0053] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by those skilled in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0054] like Figure 1 The diagram shown is a flowchart illustrating a privacy-enhancing aggregation method for multi-source trusted federated learning according to an embodiment of the present invention. (Refer to...) Figure 1 This invention provides a privacy-enhancing aggregation method for multi-source trusted federated learning, comprising the following steps:
[0055] S1. Train the original gradient of the corresponding client based on the local power consumption data of each client, and perform privacy protection processing on each original gradient to obtain the reported gradient of the corresponding client.
[0056] like Figure 2 The diagram shown is a flowchart illustrating step S1 of an embodiment of the present invention. (Refer to...) Figure 2 Step S1 includes:
[0057] S101. The backpropagation algorithm is used to train the local power consumption data of each client to obtain the original gradient of the corresponding client.
[0058] The client can be understood as a power data node, encompassing residential electricity consumption terminals, industrial load monitoring points, and substation data units. Local electricity consumption data refers to the electricity consumption data collected locally by the client, including but not limited to residential electricity consumption curves and industrial load data.
[0059] Specifically, each client constructs a neural network model based on locally collected electricity consumption data. It calculates the loss function between the predicted and actual values using forward propagation, and then calculates the gradient of the loss function with respect to the model parameters—the original gradient—using the backpropagation algorithm. This process is completed locally on the client, ensuring that the original electricity consumption data does not leave the local device, thus complying with the data privacy protection principles of federated learning.
[0060] S102. For each original gradient, decompose it into several layer gradient vectors corresponding to the original gradient, and construct an orthogonal complement subspace for each layer gradient vector to obtain a perturbation candidate set for the corresponding layer gradient vector.
[0061] Specifically, the original gradient of each client is decomposed into several gradient vectors according to the neural network hierarchy, and the gradient vector of the i-th layer is denoted as . For each layer gradient vector The perturbation candidate set is generated by constructing orthogonal complement subspaces, that is, by sampling random vectors of the same dimension from the standard normal distribution. After orthogonal projection, eliminate and The directional correlation is used to obtain orthogonal perturbation vectors, which are then scaled using norm scaling to make the perturbation vectors align with... By maintaining the same scale, multiple perturbation candidates corresponding to the gradient of this layer are eventually formed, which together constitute the perturbation candidate set.
[0062] The above process disrupts the feature correlation of the original gradient through hierarchical orthogonal perturbation, while ensuring the scale consistency between the perturbation vector and the original gradient, laying the foundation for subsequent privacy protection and model stability.
[0063] S103. Perform local loss evaluation on each perturbation candidate set to obtain the loss value of the corresponding perturbation candidate set, and perform cold Bayesian posterior screening on the loss value of each perturbation candidate set to obtain the reporting gradient of the corresponding client.
[0064] Specifically, for each perturbation candidate in the perturbation candidate set for each client, it is applied to the local model, and the loss function value on the local electricity consumption data is calculated to obtain the corresponding candidate's loss value. Further, based on a cold Bayesian posterior probability model, the candidate set is screened using the loss value as the key indicator. That is, the matching degree between each candidate and the original gradient optimization direction is quantified through the posterior probability distribution, and the perturbation candidate with the highest probability is selected as the reported gradient for that client. This step ensures the effectiveness of the model optimization of the perturbation gradient through loss evaluation and enhances the robustness of candidate selection through Bayesian screening, avoiding bias caused by a single indicator and achieving a balance between privacy protection and model performance.
[0065] S2. Perform cluster identity obfuscation processing on the reported gradient of each client to obtain the multi-cluster perturbation update set of the corresponding client, and adjust the real and fake update ratio of each multi-cluster perturbation update set to obtain the anonymized upload data of the interest cluster of the corresponding client.
[0066] Each client uploads the perturbed gradient, i.e. the reported gradient, to multiple cluster identities simultaneously to achieve anonymization of cluster identities.
[0067] like Figure 3 The diagram shown is a flowchart illustrating step S2 of an embodiment of the present invention. (Refer to...) Figure 3 Step S2 includes:
[0068] S201. Perform real cluster identity matching on the reported gradient of each client to obtain the real cluster perturbation update of the corresponding client, and perform fake cluster masquerading on each real cluster perturbation update to obtain the multi-cluster perturbation update set of the corresponding client.
[0069] The multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates.
[0070] Specifically, for each client's reported gradient, it is first matched based on the identifier of its real cluster to obtain the real cluster perturbation update corresponding to the cluster features. Then, by forging the identity information of non-cluster members, the real cluster perturbation update is disguised to generate several fake cluster updates. Finally, the real cluster perturbation update and the forged fake cluster updates are integrated to form a set of multi-cluster perturbation updates containing real and fake information, thereby masking the client's real cluster affiliation and enhancing the privacy protection of cluster identity.
[0071] S202. Based on the preset true and false update ratio, each multi-cluster perturbation update set is filtered and combined to obtain the corresponding ratio-adapted update combination, and the anonymous cluster identity is associated with each ratio-adapted update combination to obtain the anonymized upload data of the interest cluster of the corresponding client.
[0072] This step generates uploaded data through ratio adjustment and anonymization. Specifically, for each client's multi-cluster perturbation update set, it is filtered and combined according to a preset ratio of genuine to fake updates to form a matching update combination that meets the ratio requirements. Then, an anonymized cluster identity identifier is attached to each update information in this matching update combination, completing the association and binding between cluster identity and update content, and finally obtaining the client's anonymized upload data of interest clusters. Among them, the anonymized upload data of interest clusters includes genuine cluster perturbation updates and fake cluster spoof updates bound to the anonymous cluster identity.
[0073] This process controls the exposure of real information by setting a preset ratio and eliminates the direct mapping between cluster identifiers and clients by combining anonymous identity association. This not only ensures the effectiveness of updates required for federated collaboration, but also further strengthens the protection of cluster identity privacy and can resist attacks based on identity inference.
[0074] S3. Extract abnormal features from the anonymized uploaded data of each client's interest cluster to obtain high deviation update seeds, and perform derivative perturbation synthesis on the high deviation update seeds to obtain a synthetic client update set.
[0075] After receiving all real and fake updates uploaded by clients, the server generates a synthetic update based on the simulated client behavior, simulating normal but deviating client upload behavior.
[0076] like Figure 4 The diagram shown is a flowchart illustrating step S3 of an embodiment of the present invention. (Refer to...) Figure 4 Step S3 includes:
[0077] S301. For each client's anonymized uploaded data of interest clusters, the deviation is evaluated to obtain the updated deviation distribution, and the extreme value of the updated deviation distribution is screened to obtain high deviation update seeds.
[0078] For client-uploaded data, anomaly features are extracted. The server anonymizes the uploaded data for each client's interest cluster and calculates the degree of deviation from the global update distribution (the deviation of each dimension's update value from the mean) to generate an update deviation distribution. Based on this distribution, extreme value screening is performed, and the update data with the highest deviation is selected as the high deviation update seed.
[0079] By quantifying the anomalous characteristics of updated data, we can accurately locate update samples that may contain malicious perturbations or anomalous patterns, providing representative anomalous seeds for subsequent derivative perturbation synthesis and supporting the robust handling of anomalous updates by the federated intelligence system.
[0080] S302. Perform structural replication and slight perturbation on the high deviation update seed to obtain the synthetic client update set.
[0081] Specifically, a synthetic update set is generated through structural replication and perturbation. The server performs parameter structure analysis on the selected high-deviation update seeds, extracts their gradient direction, dimensional distribution and other feature structures, replicates them based on the structure, and generates multiple pseudo-updates with similar parameter distributions to the seeds. Then, a slight Gaussian perturbation (such as adding random noise that follows a normal distribution) is applied to each pseudo-update, so that it produces subtle differences while maintaining the original structural features. Finally, all perturbed pseudo-updates are integrated with the original high-deviation seeds to form a synthetic client update set.
[0082] By employing a dual approach of structural replication and mild perturbation, key features of anomalous updates are preserved for robust training, while diversity is introduced to reduce the risk of attacks from single samples.
[0083] S4. Robustly aggregate the updated set of the synthetic clients and the anonymized uploaded data of each client's interest clusters to obtain the global model update;
[0084] The Byzantine-Lupin aggregation algorithm is used to aggregate the synthetic client update set and the anonymized uploaded data of each client's interest clusters to obtain a global model update resistant to Byzantine attacks.
[0085] Specifically, the server inputs the synthesized client update set and the anonymized uploaded data of all clients' interest clusters into the Byzantine robust aggregation algorithm. The algorithm first calculates the geometric distance between each update and other updates, quantifies its deviation from the overall distribution, and selects the K samples with the smallest distance that are most similar to the majority of updates as consistent updates based on the distance metric, and removes outlier Byzantine updates. Then, the selected consistent updates are weighted and averaged to obtain a global model update resistant to Byzantine attacks.
[0086] S5. Based on the parsing of the global model update by each client, obtain the multi-cluster association data of the corresponding client, and perform model reconstruction and solution for each multi-cluster association data to obtain the real cluster model parameters of the corresponding client.
[0087] The parameters of each real cluster model are used to support the corresponding client in the next round of local training.
[0088] like Figure 5 The diagram shown is a flowchart illustrating step S5 of an embodiment of the present invention. (Refer to...) Figure 5 Step S5 includes:
[0089] S501. Each client performs cluster identity matching on the received global model update to obtain the associated cluster model set, and performs model decomposition on the associated cluster model set to obtain the multi-cluster associated data of the corresponding client.
[0090] Specifically, after receiving the global model update, each client, based on its locally stored historical cluster identity information, performs similarity matching with the parameters of each cluster in the global model, selecting cluster models with a correlation higher than a threshold to form a set of associated cluster models. Each cluster model in this set is then parameter-decomposed, its key parameter matrices (such as weight matrices and bias vectors) are extracted, and cross-mapped with its own historical update records to generate structured data containing multi-cluster feature relationships, i.e., multi-cluster associated data. The threshold is dynamically set based on the statistical characteristics of the client's historical cluster correlation distribution combined with the convergence requirements of federated learning, ensuring that the selected associated cluster models are both representative and meet the accuracy requirements of collaborative training.
[0091] S502. For each multi-cluster associated data, a mapping relationship model is performed to obtain the mapping matrix of the corresponding client. The least squares method is used to solve the parameters of each mapping matrix to obtain the real cluster model parameters of the corresponding client.
[0092] Specifically, for the multi-cluster associated data of each client, a mapping relationship between the hybrid model and the local real model is constructed to generate a mapping matrix. Then, the least squares method is used to optimize the parameters of the mapping matrix. The parameters are adjusted with the goal of minimizing the matrix error to obtain the real cluster model parameters that can accurately reflect the characteristics of the client's local data.
[0093] By connecting global aggregated information with local model features through a mapping matrix, and leveraging the numerical optimization properties of the least squares method, it is possible to efficiently restore the true model parameters from multi-cluster associated data, ensuring the accuracy of the next round of local training on the client side.
[0094] This invention presents a privacy-enhancing aggregation method for multi-source trusted federated learning. By protecting the privacy of the original gradients and obfuscating cluster identities, it masks the characteristics of local electricity consumption data from the client. Through multi-cluster perturbation updates and true / false data ratio adjustments, it significantly reduces the risk of reverse inference of the client's true electricity consumption preferences, achieving deep protection of electricity consumption data privacy. Through abnormal feature extraction and derived perturbation synthesis, combined with a robust aggregation mechanism, it effectively filters abnormal update interference from malicious clients, significantly improving the global model update's resistance to Byzantine attacks and ensuring the stability and credibility of the aggregation process. The client autonomously reconstructs the true cluster model based on the global model update, without relying on the server for local model reconstruction. This strengthens the decentralized nature of the federated intelligent system and avoids the privacy leakage risk caused by secondary transmission of model parameters. This invention is particularly suitable for multi-source heterogeneous electricity consumption data scenarios. While ensuring the privacy and security of data from each client, it improves the global model's ability to fit complex electricity consumption patterns through federated collaborative learning, balancing privacy protection, model performance, and training efficiency. This provides reliable technical support for the application of multi-source trusted federated learning in fields such as smart grids.
[0095] like Figure 6 The diagram shown is a schematic representation of a privacy-enhancing aggregation system for multi-source trusted federated learning according to an embodiment of the present invention. (Refer to...) Figure 6 This invention provides a privacy-enhancing aggregation system for multi-source trusted federated learning, comprising:
[0096] The gradient reporting module 1 is used to train the original gradient of the corresponding client based on the local power consumption data of each client, and to perform privacy protection processing on each original gradient to obtain the reported gradient of the corresponding client.
[0097] Specifically, the gradient determination module is reported, including:
[0098] The original gradient determination unit is used to train the local power consumption data of each client using the backpropagation algorithm to obtain the original gradient of the corresponding client.
[0099] The perturbation candidate determination unit is used to decompose each original gradient into several layer gradient vectors corresponding to the original gradient by layer, and to construct an orthogonal complement subspace for each layer gradient vector to obtain a perturbation candidate set for the corresponding layer gradient vector.
[0100] The posterior filtering unit is used to perform local loss evaluation on each perturbation candidate set to obtain the loss value of the corresponding perturbation candidate set, and to perform cold Bayesian posterior filtering on the loss value of each perturbation candidate set to obtain the reported gradient of the corresponding client.
[0101] Upload data determination module 2 is used to perform cluster identity obfuscation processing on the reported gradient of each client to obtain the multi-cluster perturbation update set of the corresponding client, and to adjust the true and false update ratio of each multi-cluster perturbation update set to obtain the anonymized upload data of the interest cluster of the corresponding client.
[0102] Specifically, the data upload determination module includes:
[0103] The perturbation update determination unit is used to perform real cluster identity matching on the reported gradient of each client to obtain the real cluster perturbation update of the corresponding client, and to perform fake cluster spoofing processing on each real cluster perturbation update to obtain the multi-cluster perturbation update set of the corresponding client, wherein the multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates;
[0104] The matching anonymization association unit is used to filter and combine each multi-cluster perturbation update set based on a preset true and false update ratio to obtain the corresponding matching adaptation update combination, and to perform anonymous cluster identity association on each matching adaptation update combination to obtain the anonymized upload data of the interest cluster of the corresponding client.
[0105] The synthesis update determination module 3 is used to extract abnormal features from the anonymized uploaded data of each client's interest cluster to obtain high deviation update seeds, and to synthesize the high deviation update seeds by performing derivative perturbation to obtain a synthetic client update set.
[0106] Specifically, the synthesis update determination module includes:
[0107] The seed determination unit is used to evaluate the deviation of the anonymized uploaded data of each client's interest cluster to obtain the updated deviation distribution, and to perform extreme value screening on the updated deviation distribution to obtain the high deviation update seed;
[0108] The seed expansion unit is used to perform structural replication and slight perturbation on the high-deviation update seed to obtain the synthetic client update set.
[0109] The robust aggregation module 4 is used to robustly aggregate the synthetic client update set and the anonymized upload data of each client's interest cluster to obtain the global model update;
[0110] The model parameter determination module 5 is used to obtain the multi-cluster association data of the corresponding client based on the parsing of the global model update of each client, and to perform model reconstruction and solve for each multi-cluster association data to obtain the real cluster model parameters of the corresponding client. The real cluster model parameters are used to support the corresponding client to perform the next round of local training.
[0111] It should be noted that the modules in the aforementioned privacy-enhancing aggregation system for multi-source trusted federated learning can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module. For specific limitations regarding the privacy-enhancing aggregation system for multi-source trusted federated learning, please refer to the limitations of the privacy-enhancing aggregation method for multi-source trusted federated learning described above; both have the same function and role, and will not be repeated here.
[0112] In summary, this invention provides a privacy-enhancing aggregation method and system for multi-source trusted federated learning. By protecting the privacy of the original gradients and obfuscating cluster identities, it masks the characteristics of local electricity consumption data from clients while significantly reducing the risk of reverse inference of clients' true electricity consumption preferences through multi-cluster perturbation updates and true / false ratio control, thus achieving deep protection of electricity consumption data privacy. Through abnormal feature extraction and derived perturbation synthesis, combined with a robust aggregation mechanism, it effectively filters abnormal update interference from malicious clients, significantly improving the global model update's resistance to Byzantine attacks and ensuring the stability and credibility of the aggregation process. Clients autonomously reconstruct the true cluster model based on global model updates, without relying on servers for local model reconstruction. This strengthens the decentralized nature of the federated intelligent system and avoids the privacy leakage risk caused by secondary transmission of model parameters. This invention is particularly suitable for multi-source heterogeneous electricity consumption data scenarios. While ensuring the privacy and security of data from each client, it improves the global model's ability to fit complex electricity consumption patterns through federated collaborative learning, balancing privacy protection, model performance, and training efficiency. This provides reliable technical support for the application of multi-source trusted federated learning in fields such as smart grids.
[0113] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0114] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention.
Claims
1. A privacy-enhancing aggregation method for multi-source trusted federated learning, characterized in that, include: The original gradient of each client is obtained by training based on the local power consumption data of each client, and privacy protection processing is performed on each original gradient to obtain the reported gradient of the corresponding client. Cluster identity obfuscation is performed on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and the true and false update ratio is adjusted for each multi-cluster perturbation update set to obtain the anonymized upload data of the interest cluster corresponding to the client. Anomaly features are extracted from the anonymized uploaded data of the interest clusters of each client to obtain high deviation update seeds, and the high deviation update seeds are subjected to derivation perturbation synthesis to obtain a synthetic client update set; Robust aggregation is performed on the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain the global model update; Based on the parsing of the global model update by each client, the multi-cluster association data corresponding to the client is obtained, and the model reconstruction solution is performed on each multi-cluster association data to obtain the real cluster model parameters corresponding to the client. The real cluster model parameters are used to support the client to perform the next round of local training. The process of performing cluster identity obfuscation on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and adjusting the true / false update ratio of each multi-cluster perturbation update set to obtain anonymized upload data of the interest cluster corresponding to the client, includes: For each client's reported gradient, a real cluster identity matching is performed to obtain the real cluster perturbation update corresponding to the client. Then, a fake cluster spoofing processing is performed on each real cluster perturbation update to obtain a multi-cluster perturbation update set corresponding to the client. The multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates. Based on a preset ratio of true to false updates, each set of multi-cluster perturbation updates is filtered and combined to obtain a corresponding ratio-adapted update combination. Then, each ratio-adapted update combination is associated with anonymized cluster identities to obtain the anonymized upload data of the interest clusters of the corresponding client. The anonymized upload data of the interest clusters includes real cluster perturbation updates and fake cluster false updates bound to the anonymous cluster identities.
2. The privacy-enhancing aggregation method for multi-source trusted federated learning according to claim 1, characterized in that, The process of training the original gradient corresponding to each client based on the local power consumption data of each client, and performing privacy protection processing on each original gradient to obtain the reported gradient corresponding to the client, includes: The backpropagation algorithm is used to train the local power consumption data of each client to obtain the original gradient of the corresponding client. For each original gradient, several layer gradient vectors corresponding to the original gradient are obtained by layer decomposition, and orthogonal complement subspaces are constructed for each layer gradient vector to obtain a perturbation candidate set for the gradient vector of the corresponding layer. For each of the perturbation candidate sets, a local loss evaluation is performed to obtain the loss value corresponding to the perturbation candidate set. Then, the loss value of each of the perturbation candidate sets is subjected to cold Bayesian posterior filtering to obtain the reporting gradient corresponding to the client.
3. The privacy-enhancing aggregation method for multi-source trusted federated learning according to claim 1, characterized in that, The step of extracting abnormal features from the anonymized uploaded data of each client's interest cluster to obtain a high-deviation update seed, and then performing derivative perturbation synthesis on the high-deviation update seed to obtain a synthetic client update set, includes: For each client, the anonymized uploaded data of the interest cluster is evaluated for deviation to obtain an updated deviation distribution, and the updated deviation distribution is screened for extreme values to obtain high deviation update seeds; The high-deviation update seed is structurally replicated and slightly perturbed to obtain a synthetic client update set.
4. The privacy-enhancing aggregation method for multi-source trusted federated learning according to claim 1, characterized in that, The robust aggregation of the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain the global model update includes: The Byzantine-Lupin aggregation algorithm is used to aggregate the synthetic client update set and the anonymized uploaded data of the interest cluster of each client to obtain a global model update resistant to Byzantine attacks.
5. The privacy-enhancing aggregation method for multi-source trusted federated learning according to claim 1, characterized in that, The step of parsing the global model update for each client to obtain the multi-cluster association data corresponding to that client, and then performing model reconstruction and solving on each of the multi-cluster association data to obtain the real cluster model parameters corresponding to that client, includes: Each client performs cluster identity matching on the received global model update to obtain an associated cluster model set, and performs model decomposition on the associated cluster model set to obtain multi-cluster associated data corresponding to the client. For each of the multi-cluster associated data, a mapping relationship model is performed to obtain the mapping matrix corresponding to the client. The least squares method is then used to solve the parameters of each mapping matrix to obtain the actual cluster model parameters corresponding to the client.
6. A privacy-enhancing aggregation system for multi-source trusted federated learning, characterized in that, include: The gradient reporting module is used to train the original gradient corresponding to each client based on the local power consumption data of each client, and to perform privacy protection processing on each original gradient to obtain the reported gradient corresponding to the client. The data upload determination module is used to perform cluster identity obfuscation processing on the reported gradient of each client to obtain a multi-cluster perturbation update set corresponding to the client, and to adjust the true and false update ratio of each multi-cluster perturbation update set to obtain anonymized upload data of the interest cluster corresponding to the client. The synthesis update determination module is used to extract abnormal features from the anonymized uploaded data of the interest cluster of each client to obtain a high deviation update seed, and to perform derivative perturbation synthesis on the high deviation update seed to obtain a synthesis client update set. The robust aggregation module is used to robustly aggregate the synthetic client update set and the anonymized upload data of the interest cluster of each client to obtain a global model update; The model parameter determination module is used to obtain the multi-cluster association data corresponding to each client based on the parsing of the global model update of each client, and to perform model reconstruction and solution on each multi-cluster association data to obtain the real cluster model parameters corresponding to the client. Each real cluster model parameter is used to support the next round of local training for the corresponding client. The uploaded data determination module includes: The perturbation update determination unit is used to perform real cluster identity matching on the reported gradient of each client to obtain the real cluster perturbation update corresponding to the client, and to perform fake cluster spoofing processing on each real cluster perturbation update to obtain a multi-cluster perturbation update set corresponding to the client, wherein the multi-cluster perturbation update set includes real cluster perturbation updates and fake cluster spoof updates; The matching anonymization association unit is used to filter and combine each of the multi-cluster perturbation update sets based on a preset real and fake update ratio to obtain a corresponding matching adaptation update combination, and to perform anonymous cluster identity association on each of the matching adaptation update combinations to obtain the anonymized upload data of the interest clusters corresponding to the client. The anonymized upload data of the interest clusters includes real cluster perturbation updates and fake cluster false updates bound to the anonymous cluster identity.
7. The privacy-enhancing aggregation system for multi-source trusted federated learning according to claim 6, characterized in that, The gradient reporting module includes: The original gradient determination unit is used to train the local power consumption data of each client using the backpropagation algorithm to obtain the original gradient of the corresponding client. The perturbation candidate determination unit is used to decompose each original gradient by layer to obtain several layer gradient vectors corresponding to the original gradient, and to construct an orthogonal complement subspace for each layer gradient vector to obtain a perturbation candidate set for the corresponding layer gradient vector. The posterior filtering unit is used to perform local loss evaluation on each of the perturbation candidate sets to obtain the loss value corresponding to the perturbation candidate set, and to perform cold Bayesian posterior filtering on the loss value of each of the perturbation candidate sets to obtain the reporting gradient corresponding to the client.
8. The privacy-enhancing aggregation system for multi-source trusted federated learning according to claim 6, characterized in that, The synthesis update determination module includes: The seed determination unit is used to evaluate the deviation of the anonymized uploaded data of the interest cluster of each client to obtain the updated deviation distribution, and to perform extreme value screening on the updated deviation distribution to obtain high deviation update seeds. The seed expansion unit is used to perform structural replication and slight perturbation on the high-deviation update seed to obtain a synthetic client update set.
Citation Information
Patent Citations
Federal machine learning method and system considering robustness and privacy protection
CN117828627A
Federal learning-based privacy improvement method, system and device and medium
CN118350051A