A cloud-edge collaborative industrial security model evolution method based on federated learning
Patent Information
- Application Number
- CN202610896285.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-22
AI Technical Summary
根据行业监管要求,工厂严禁将原始数据上传至云端,直接汇聚数据训练全局模型的方式面临严重的合规风险
[0018]本申请提出的一种基于联邦学习的云边协同工业安全模型进化方法,与传统的联邦学习方法相比,至少具有以下提升。
Smart Images

Figure CN122420001B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the fields of industrial artificial intelligence and edge computing technology, and in particular to a cloud-edge collaborative industrial security model evolution method based on federated learning. Background Technology
[0002] With the deep integration of the Industrial Internet and Artificial Intelligence (AI) technologies, AI-based industrial safety systems have been widely applied in high-risk industries such as energy, chemical, and power. A typical system architecture adopts a collaborative model of "edge perception + cloud decision-making." Edge nodes (such as AI gateways and edge servers in the factory) deploy lightweight models to perform real-time analysis of data from cameras and sensors. The cloud is responsible for aggregating data from each edge node, training a global model using large-scale samples, and then distributing the global model to each edge node for updates.
[0003] However, this traditional architecture contains two fundamental contradictions.
[0004] The first issue is the conflict between data privacy and compliance. Industrial data may involve information such as production processes, equipment parameters, and trade secrets. According to industry regulations, factories are strictly prohibited from uploading raw data to the cloud, and directly aggregating data to train global models faces serious compliance risks.
[0005] The second challenge is the heterogeneity of data. Different factories have highly heterogeneous production environments, tasks, equipment models, and accident types. For example, chemical plant A primarily faces the risk of pipeline leaks, while chemical plant B primarily faces the risk of reactor overheating. If a traditional federated learning approach is used to uniformly aggregate the model across all nodes, it will lead to a decline in model performance and even "negative transfer," meaning that for some nodes, the global model's performance is worse than their locally trained, independently developed model.
[0006] In summary, although some solutions have attempted to address data privacy issues using federated learning frameworks, most of these solutions assume that the data of each node is independent and identically distributed, failing to fully consider the strong heterogeneity of industrial scenarios. They also lack targeted model evolution strategies for newly connected edge nodes (with sparse data) or extremely skewed nodes. Summary of the Invention
[0007] To address the aforementioned technical issues, embodiments of this application propose a cloud-edge collaborative industrial security model evolution method based on federated learning. This method, while protecting the data privacy of each edge node, aggregates edge nodes with similar data distributions into federated groups for collaborative training through a task similarity-aware federated grouping mechanism. Furthermore, it combines differential privacy noise addition, secure multi-party computation aggregation, and knowledge distillation-assisted cold-start strategies to achieve a high-precision, adaptive multi-factory collaborative model evolution that meets industrial data compliance requirements.
[0008] To achieve the above objectives, embodiments of this application propose a cloud-edge collaborative industrial safety model evolution method based on federated learning. This method is implemented using an aggregation server deployed in the cloud and edge nodes deployed in different factories. The method includes the following steps: a basic model is pre-set in the aggregation server; multiple edge nodes each register with the cloud and hold local datasets; the aggregation server distributes the basic model to the edge nodes registered in the cloud; the aggregation server collects metadata from each edge node and calculates the task similarity between any two edge nodes based on the metadata, grouping edge nodes with task similarity greater than a preset similarity threshold into the same federated family; wherein, the metadata includes equipment type distribution, historical accident type distribution, and initial local update gradient; the aggregation server traverses each federated family, performs initial aggregation based on the local update gradients of each edge node in the current federated family, obtains the aggregated model of the current federated family, and distributes the aggregated model of the current federated family to the current federated family. Each edge node in a federation uses its local dataset to train the received aggregated model locally and calculates a noisy gradient. The encrypted noisy gradient is then uploaded to the aggregation server. The aggregation server uses secure multi-party computation and secret sharing techniques to aggregate the noisy gradients uploaded by edge nodes within the same federation, resulting in an updated aggregated model. This updated model is then distributed to the corresponding edge nodes within the federation. Upon receiving the updated aggregated model, each edge node uses its local dataset for personalized fine-tuning, thus achieving model evolution. The aggregation server designates newly registered edge nodes and edge nodes with local dataset sample counts below a preset threshold as sparse nodes. Mature aggregated models are used as teacher models for these sparse nodes to perform knowledge distillation training, enabling the first local update. The aggregation server also periodically recalculates the task similarity between edge nodes, dynamically adjusting the federation division.
[0009] To achieve the above objectives, embodiments of this application also propose an electronic device, including: a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement the above-described cloud-edge collaborative industrial security model evolution method based on federated learning.
[0010] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of a cloud-edge collaborative industrial security model evolution method based on federated learning as described above.
[0011] Optionally, the local dataset held by the edge nodes centrally includes local basic data and local personalized data. Local basic data refers to the common data collected by each edge node, including general image data collected by high-definition cameras, general sensor data collected by sensors, and fault record data. Local personalized data refers to the proprietary data collected by the edge nodes in their respective factories and production processes, including the operating status parameter sequence of specific production equipment, proprietary environmental monitoring data for specific process environments, and proprietary image data with specific industrial accident characteristics. For chemical plants, localized personalized data includes temperature and pressure data of reactors, concentration data of special gases, and infrared thermal imaging images of pipeline leaks. For machinery manufacturing plants, local personalized data includes high-frequency vibration signals of machine tool spindles, current and voltage data and spatter data of the welding process, and visual images of surface defects of specific materials; For power energy plants, local personalized data includes ultrasonic signals of partial discharge in transformers and infrared heating sequences of power generation equipment contacts.
[0012] Optionally, the aggregation server collects metadata from each edge node and calculates the task similarity between any two edge nodes based on the metadata. Edge nodes with task similarity greater than a preset similarity threshold are grouped into the same federation family, including: After receiving the base model, the edge node performs the first local training on the base model based on the local base data in the local dataset, according to the set first local training round, and obtains the first local update gradient. Edge nodes extract device type distribution probability vectors and historical accident type distribution probability vectors from local datasets, and package them together with the first local update gradient as metadata, which is then encrypted and uploaded to the aggregation server. After collecting the metadata uploaded by each edge node, the aggregation server calculates the task similarity between any two edge nodes based on the metadata, and assigns edge nodes with task similarity greater than a preset similarity threshold to the same federation family, ultimately resulting in several federation families. Denote any two edge nodes and The task similarity between them is The calculation formula is as follows: ; in, This represents the total number of layers in the base model. and They represent and In the The first local gradient update of the layer, Indicates the first The cosine weighting parameter of the layer, and They represent and The device type distribution probability vector, and They represent and The probability vector of historical accident types. This represents finding the Jason-Shannon divergence between two probability vectors. , , All are preset similarity weight coefficients.
[0013] Optionally, the aggregation server traverses each federation, performs the first aggregation based on the local update gradients of each edge node in the current federation, obtains the aggregation model of the current federation, and distributes the aggregation model of the current federation to each edge node of the current federation, including: The aggregation server traverses each federation family, calculates the mean of the local update gradient of each edge node in the current federation family as the reference gradient direction, and orthogonally projects the first local update gradient of each edge node in the current federation family onto the reference gradient direction to extract the common gradient components that are consistent with the global optimization direction of the current federation family, and then separates the orthogonal residual components that reflect the specificity of the edge nodes. The aggregation server calculates the aggregation weight of each edge node in the current federation family based on the projection magnitude of the common gradient components of each edge node. The larger the projection magnitude of the common gradient components of the edge node, the more consistent the local update of the edge node is with the global optimization direction of its federation family, and the higher the aggregation weight it is assigned. The aggregation server performs a weighted average of the common gradient components of each edge node of the current federation family based on the aggregation weight, generates the aggregation gradient of the current federation family, updates the base model using the aggregation gradient, obtains the aggregation model of the current federation family, and distributes it to each edge node of the current federation family. Assume there are currently a total of The edge node, the first edge nodes The first local update gradient is Then the baseline gradient direction of the current federation family The calculation formula is as follows: ; Common gradient components The following projection formula is used for calculation: ; in, This is a preset smoothing term used to prevent the denominator from being zero; Aggregate weights based on The projection amplitude is calculated using the following formula: ; in, for The projection amplitude; The aggregation server generates the aggregated gradient of the current federation family by weighted averaging of the common gradient components of each edge node based on the aggregation weights. This is achieved through the following formula: ; in, This represents the aggregation gradient of the current federation family.
[0014] Optionally, the edge node uses its local dataset to train the received aggregated model locally, computes the noisy gradient, and uploads the encrypted noisy gradient to the aggregation server, including: Edge nodes use their local datasets to train the received aggregated model locally, and divide the trained gradients into shallow feature gradients and deep decision gradients. Edge nodes use an adaptive sensitivity function to calculate the clipping thresholds corresponding to the shallow feature gradient and the deep decision gradient, respectively, and inject Laplacian noise of different intensities into the shallow feature gradient and the deep decision gradient according to the clipping threshold, thereby obtaining the noisy gradient; Edge nodes generate pseudo-random streams based on the parameters of the current federated training round and the federated family to which they belong, and use the mask matrix to perform symbolic XOR obfuscation encryption on the noisy gradients, and then upload the encrypted noisy gradients to the aggregation server. remember The encrypted, noisy gradient uploaded to the aggregation server is The calculation formula is as follows: ; ; in, for The noisy gradient, This indicates the Hadamard product. For symbolic functions, For the mask matrix, For adaptive sensitivity function, It is Laplace noise. for Shallow feature gradient, for The deep decision gradient, and These are the clipping thresholds corresponding to the shallow feature gradient and the deep decision gradient, respectively. , For shallow sensitivity, For deep sensitivity, This is the shallow privacy budget coefficient. For deep privacy budgeting coefficients, .
[0015] Optionally, the aggregation server uses secure multi-party computation and secret sharing techniques to aggregate the noisy gradients uploaded by edge nodes of the same federation family to obtain an updated aggregation model, and then distributes the updated aggregation model to each edge node of the corresponding federation family, including: After receiving the encrypted and noisy gradients uploaded by each edge node of the same federation family, the aggregation server uses secret sharing technology to split the encrypted and noisy gradients uploaded by each edge node into multiple random fragments. The number of random fragments is the same as the number of edge nodes in the federation family, and the sum of all random fragments is equal to the corresponding encrypted and noisy gradient. The aggregation server distributes a random shard of each edge node to the other edge nodes in the federation, so that each edge node holds a random shard of each edge node in the federation. Each edge node sums up all the random shards it holds locally to obtain the local shard aggregation result, and then sends the local shard aggregation result back to the aggregation server; The aggregation server receives the local shard aggregation results returned by each edge node, uses the secret sharing reconstruction algorithm to accumulate all the local shard aggregation results, restores the global aggregation gradient of the federation family, updates the current federation family's aggregation model based on the global aggregation gradient, obtains the updated aggregation model, and sends the updated aggregation model to each edge node of the corresponding federation family. Let the first There are a total of [number] federal families. The edge node, the first edge nodes The uploaded encrypted noisy gradient is The aggregation server splits it into Random partitions , ; The aggregation server will randomly shard the data. Distributed to the The first federal family edge nodes , Locally held The local sharding aggregation result is obtained by summing the random shards of all edge nodes in each federation family. , ; The aggregation server receives the first The local fragment aggregation results returned by each edge node of the federation family are summed using a secret-sharing reconstruction algorithm to reconstruct the fragment aggregation result. Global aggregation gradient of each federal family , .
[0016] Optionally, after receiving the updated aggregated model, the edge nodes use their held local datasets to perform personalized fine-tuning, thereby achieving model evolution, including: After receiving the updated aggregation model, the edge nodes split it into a frozen backbone network for extracting general industrial features and a personalized classification head for the device's local environment. The parameters of the edge nodes remain unchanged while the parameters of the frozen backbone network are fine-tuned using local personalized data in the local dataset. The personalized fine-tuning loss function used in the personalized fine-tuning process Represented as: ; in, Parameters representing personalized classification headers, The task loss function for the production tasks of the factory and production process where the edge node is located. This represents the parameters of the updated aggregation model. This represents local personalized data. For proximal constraint terms, To adjust the proximal balance coefficient for the specific intensity of a plant-specific policy, .
[0017] Optionally, the aggregation server calls a mature aggregation model as the teacher model for sparse nodes to perform knowledge distillation training to achieve the first local update, including: The aggregation server selects nodes from all federations with the highest task similarity to sparse nodes based on the device type distribution and historical incident type distribution of sparse nodes. A federal tribe, based on this The latest aggregation model of each federal family constitutes the teacher model set. And distribute it to the edge nodes; in, , The task similarity between the node and the sparse node is represented by the first... The latest aggregation model of high-level federation families, namely the first A teacher model; Sparse nodes use the base model as the student model It utilizes its local dataset to perform knowledge distillation training on the student model to achieve the first local update; Knowledge distillation training loss function Represented as: ; in, Let cross-entropy be the loss function. For the output of the student model, The soft labels for the local dataset of the edge nodes. This is a preset balance coefficient used to balance annotation supervision and knowledge distillation. Relative entropy loss function, For the Softmax function, The preset distillation temperature parameter is used to smooth the probability distribution. For the first The output of the teacher model.
[0018] This application proposes a cloud-edge collaborative industrial security model evolution method based on federated learning, which has at least the following improvements compared with traditional federated learning methods.
[0019] First, it achieves closed-loop protection for data privacy compliance across the entire chain. This application ensures that sensitive industrial data exists in ciphertext or perturbed gradient form throughout its entire lifecycle of training, transmission, and aggregation by combining local privacy noise addition, symbol masking obfuscation encryption, and secure multi-party computation in the cloud. This completely blocks the possibility of leakage of the original data and effectively meets the compliance requirements for the classification and hierarchical management of industrial data.
[0020] Second, it achieves anti-interference performance optimization under heterogeneous working conditions. This application introduces a task similarity-aware federated clustering mechanism, abandoning the traditional one-size-fits-all aggregation mode. It accurately clusters from multiple dimensions such as equipment type distribution, historical accidents and gradient geometry direction, effectively suppressing the "negative migration" phenomenon caused by cross-heterogeneous working conditions, and significantly improving the model evolution accuracy of each edge node in complex industrial scenarios.
[0021] Third, it achieves a deep balance between group generalization and "one policy per factory". After the edge node receives the updated aggregation model, this application adds a local fine-tuning strategy based on near-end regularization constraints. This strategy not only fully incorporates the group security knowledge shared within the federation family, but also achieves personalized evolution by adapting to local personalized data, perfectly balancing the model's generalization ability with the specific needs of the local production environment.
[0022] Fourth, it enables rapid cold start for sparse nodes with low barriers to entry. For newly connected nodes or edge nodes with extremely scarce data, this application introduces a knowledge distillation mechanism, which uses mature teacher models in the federation family to quickly train student models in a targeted manner. This significantly shortens the cycle from random initialization to industrial usability, and significantly reduces the threshold and annotation cost of industrial intelligent deployment.
[0023] Fifth, it provides adaptive dynamic evolution capabilities for industrial scenarios. This application sets up a periodic task similarity recalculation mechanism and a federation dynamic adjustment mechanism within the aggregation server, which can perceive the industrial data distribution drift caused by production process upgrades and equipment changes in real time, ensuring that the model always adapts to the current production environment and ensuring the robustness of the industrial safety system in long-term operation. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0025] Figure 1 This is a flowchart of a cloud-edge collaborative industrial security model evolution method based on federated learning provided in one embodiment of this application; Figure 2 This is a schematic diagram of metadata collection and federal family division provided in one embodiment of this application; Figure 3 This is a schematic diagram of the uploading of encrypted noisy gradients, the aggregation of noisy gradients, and the distribution of the updated aggregated model provided in one embodiment of this application; Figure 4 This is a schematic diagram of knowledge distillation training provided in one embodiment of this application; Figure 5 This is an edge node in the federation family Group_A provided in one embodiment of this application. A diagram illustrating the performance improvement of the models; Figure 6This is a sparse node provided in one embodiment of this application. A diagram illustrating the comparison of cold start performance; Figure 7 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0027] One embodiment of this application proposes a cloud-edge collaborative industrial security model evolution method based on federated learning, which is implemented based on an aggregation server deployed in the cloud and edge nodes deployed in different factories. The implementation details of the cloud-edge collaborative industrial security model evolution method based on federated learning proposed in this embodiment are described below. The following details are provided for ease of understanding and are not necessary for implementing this solution.
[0028] The specific process of the cloud-edge collaborative industrial security model evolution method based on federated learning proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 101: Pre-configure the basic model in the aggregation server. Multiple edge nodes register with the cloud and hold local datasets. The aggregation server then distributes the basic model to the edge nodes registered in the cloud.
[0029] Specifically, the aggregation server is deployed in the cloud, and multiple edge nodes corresponding to factories, production workshops, and production lines are deployed in different factories. Each edge node holds its own private, non-transferable local dataset and registers itself with the cloud. The aggregation server has a pre-built base model, which is distributed to the edge nodes registered in the cloud when needed.
[0030] In one example, the base model can be a lightweight convolutional neural network model such as YOLOv8, ResNet-18, or EfficientNet-B0, or a temporal model such as LSTM or GRU. The number of parameters of the base model needs to be strictly controlled to adapt to the computing power limitations of the edge nodes.
[0031] In one example, the local dataset held by the edge node contains local basic data and local personalized data. Local basic data refers to the common data collected by each edge node, including general image data collected by high-definition cameras, general sensor data collected by sensors, and fault record data. Local personalized data refers to the proprietary data collected by the edge node in its factory and production process, including the operating status parameter sequence of specific production equipment, proprietary environmental monitoring data for specific process environments, and proprietary image data with specific industrial accident characteristics.
[0032] In one example, for a chemical plant, localized personalized data includes temperature and pressure data of the reactor, concentration data of specialty gases, and infrared thermal imaging images of pipeline leaks.
[0033] The core safety hazards in chemical plants are reactor overheating and material leakage. Because the exothermic reaction mechanisms differ across chemical production lines, their peripheral nodes must collect not only basic temperature sensor data but also core mechanism data (e.g., for Plant A, which involves hydrogenation, the specific data includes the hydrogen flow rate and pressure interlock ratio; for Plant B, which involves polymerization, the specific data includes the monomer feed rate and agitator current). These data determine the risk assessment benchmark for the model and are unique to chemical plants.
[0034] In one example, for a machinery manufacturing plant, local personalized data includes high-frequency vibration signals from machine tool spindles, current and voltage data and spatter data from the welding process, and visual images of surface defects in specific materials.
[0035] The core safety hazards in machinery manufacturing plants are mechanical entanglement, tool breakage, and equipment overload. The personalized data collected by edge nodes mainly comes from physical mechanical signals (such as acceleration vibration sensor data of spindle drive bearings under specific cutting processes, and high-frequency acoustic signals when tools cut metal). This data can be linked with general image data to capture the moment when operators illegally enter dangerous areas or when machinery suddenly fails.
[0036] In one example, for a power plant, local personalized data includes ultrasonic signals of partial discharge from transformers and infrared heating sequences of power generation equipment contacts.
[0037] The core safety hazards of power energy plants are high-temperature melting splashes and leaks. Edge nodes must collect data on specific physical quantities (such as: high-temperature radiation wavelengths collected by infrared thermal imagers, instantaneous current changes in the high-voltage electrodes of electric furnaces, and negative pressure sequences in dust removal pipelines).
[0038] Step 102: The aggregation server collects the metadata of each edge node, calculates the task similarity between any two edge nodes based on the metadata, and classifies edge nodes with task similarity greater than a preset similarity threshold into the same federation family.
[0039] Specifically, after federated learning begins, the aggregation server collects metadata from each edge node and calculates the task similarity between any two edge nodes based on this metadata. Edge nodes with task similarity greater than a preset similarity threshold are grouped into the same federated family. The metadata includes device type distribution, historical accident type distribution, and the initial local update gradient.
[0040] In one example, let any two edge nodes be... and The task similarity between them is The calculation process can be expressed by the formula: ; in, This represents the total number of layers in the base model. and They represent and In the The first local gradient update of the layer, Indicates the first The cosine weighting parameter of the layer, and They represent and The device type distribution probability vector, and They represent and The probability vector of historical accident types. This represents finding the Jason-Shannon divergence between two probability vectors. , , All are preset similarity weight coefficients.
[0041] The accident distribution across different factories is extremely skewed (e.g., 85% of accidents in chemical plant A are leaks, while 90% of accidents in chemical plant B are overheating). The KL divergence is asymmetric and unsuitable for calculating symmetric task similarity. The Jason Shannon divergence (JS divergence), on the other hand, is symmetric and its value range is strictly limited. It can be perfectly converted into a metric for task similarity.
[0042] In one example, a chemical industrial park has 5 chemical plants, and each chemical plant has an edge node, denoted as . , , , , They are engaged in the production of different chemical products. and The main risk is pipeline leakage (the dataset mainly consists of leak images and pressure sensor data). and The main risk is overheating of the reactor (the dataset mainly consists of temperature sequences and videos inside the reactor). For newly commissioned chemical plants, the amount of data is extremely limited.
[0043] Each edge node packages the device type distribution probability vector, the historical accident type distribution probability vector, and the first local update gradient into metadata and uploads it to the aggregation server in the cloud. (Choose not to upload). Leakage accidents accounted for 85% of all incidents. Leakage accidents accounted for 78% of all incidents. Overheating accidents accounted for 90% of all incidents. Overheating accidents accounted for 82% of all incidents. Without sufficient statistics, it does not participate in the federal family division. The aggregation server calculates task similarity. , , The preset similarity threshold is 0.75, therefore and Belonging to the same federal family (leaking sensitive family), denoted as Group_A, and It belongs to another federal group (the hyperthermophilic group), denoted as Group_B.
[0044] Step 103: The aggregation server traverses each federation family, performs the first aggregation based on the local update gradient of each edge node of the current federation family, obtains the aggregation model of the current federation family, and distributes the aggregation model of the current federation family to each edge node of the current federation family.
[0045] Specifically, after completing the division of federations, the aggregation server can traverse each federation, perform the first aggregation based on the local update gradient of each edge node of the current federation, obtain the aggregation model of the current federation, and distribute the aggregation model of the current federation to each edge node of the current federation.
[0046] In one example, the aggregation server traverses each federation, calculates the mean of the local update gradients of each edge node in the current federation as the baseline gradient direction, and orthogonally projects the initial local update gradients of each edge node in the current federation onto the baseline gradient direction. This extracts common gradient components consistent with the global optimization direction of the current federation, thus separating orthogonal residual components reflecting the specificity of the edge nodes. Subsequently, based on the projection magnitude of the common gradient components of each edge node in the current federation, the aggregation weight of each edge node within the current federation is calculated (the larger the projection magnitude of the common gradient components of an edge node, the more consistent its local update is with the global optimization direction of its federation, and the higher its assigned aggregation weight). Finally, based on the aggregation weights, a weighted average of the common gradient components of each edge node in the current federation is calculated to generate the aggregated gradient of the current federation. This aggregated gradient is then used to update the base model, resulting in the aggregated model of the current federation, which is then distributed to each edge node in the current federation.
[0047] Assume there are currently a total of The edge node, the first edge nodes The first local update gradient is Then the baseline gradient direction of the current federation family The calculation formula is as follows: .
[0048] Common gradient components The following projection formula is used for calculation: ; in, This is a preset smoothing term used to prevent the denominator from being zero.
[0049] Aggregate weights based on The projection amplitude is calculated using the following formula: ; in, for The projection amplitude.
[0050] The aggregation server generates the aggregated gradient of the current federation family by weighted averaging of the common gradient components of each edge node based on the aggregation weights. This is achieved through the following formula: ; in, This represents the aggregation gradient of the current federation family.
[0051] The traditional FedAvg method directly performs a weighted average of the gradients of all edge nodes. However, in industrial scenarios, even if nodes are grouped within the same federation, some edge nodes may have gradient directions that deviate from the mainstream direction due to extremely limited local data or severe noise (such as sensor malfunctions). Forcing an average can lead to "negative migration," skewing the originally good model. Therefore, this implementation introduces a "gradient orthogonal projection" mechanism. Through projection, we retain only the portion of the gradients of each edge node that is parallel to the "reference direction" (common gradient components), directly filtering out "orthogonal residual components" (i.e., unique, inconsistent noise or extreme specificity of each node) that may cause model oscillations. Instead of simply assigning weights based on data volume, we assign weights based on contribution; the node whose gradient direction is most consistent with the majority (larger projection amplitude) has greater influence during aggregation.
[0052] Step 104: The edge node uses its local dataset to train the received aggregated model locally, calculates the noisy gradient, and uploads the encrypted noisy gradient to the aggregated server.
[0053] Specifically, after receiving the aggregated model from the aggregation server, the edge node can use its local dataset to train the received aggregated model locally, calculate the noisy gradient, and upload the encrypted noisy gradient to the aggregation server.
[0054] In one example, such as Figure 3 As shown, edge nodes use their local datasets to train the received aggregation model locally, dividing the trained gradients into shallow feature gradients and deep decision gradients. Then, an adaptive sensitivity function is used to calculate the pruning thresholds for the shallow feature gradients and deep decision gradients, respectively. Based on these pruning thresholds, Laplacian noise of varying intensities is injected into the shallow feature gradients and deep decision gradients to obtain noisy gradients. Finally, a pseudo-random stream is generated based on the parameters of the current federated training round and the federated family, serving as a mask matrix. This mask matrix is used to perform sign-based XOR obfuscation encryption on the noisy gradients, and the encrypted noisy gradients are uploaded to the aggregation server. remember The encrypted, noisy gradient uploaded to the aggregation server is The calculation formula is as follows: ; ; in, for The noisy gradient, This indicates the Hadamard product. For symbolic functions, For the mask matrix, For adaptive sensitivity function, It is Laplace noise. for Shallow feature gradient, for The deep decision gradient, and These are the clipping thresholds corresponding to the shallow feature gradient and the deep decision gradient, respectively. , For shallow sensitivity, For deep sensitivity, This is the shallow privacy budget coefficient. For deep privacy budgeting coefficients, .
[0055] Conventional differential privacy typically uses only general Laplacian noise. This embodiment proposes an asymmetric noise addition strategy of "emphasizing privacy in shallow layers and emphasizing utility in deep layers" based on the physical characteristics of deep learning models, which further enhances privacy protection and model training accuracy.
[0056] In one example, such as Figure 3 As shown, after receiving the aggregation model, the edge nodes undergo three rounds of local training with a learning rate of 0.001.
[0057] Step 105: The aggregation server uses secure multi-party computation and secret sharing technology to aggregate the noisy gradients uploaded by each edge node of the same federation family to obtain an updated aggregation model, and then distributes the updated aggregation model to each edge node of the corresponding federation family.
[0058] Specifically, the aggregation server uses secure multi-party computation and secret sharing technology to aggregate the noisy gradients uploaded by each edge node of the same federation family to obtain an updated aggregation model, and then distributes the updated aggregation model to each edge node of the corresponding federation family.
[0059] In one example, after receiving encrypted and noisy gradients uploaded by edge nodes within the same federation family, the aggregation server uses a secret-sharing technique to split each encrypted and noisy gradient into multiple random fragments. The number of random fragments matches the number of edge nodes in the federation family, and the sum of all random fragments equals the corresponding encrypted and noisy gradient. Then, one random fragment from each edge node is distributed to the other edge nodes in the federation family, ensuring that each edge node holds one random fragment from each of the edge nodes within the federation family. Next, each edge node locally sums all its random fragments to obtain a local fragment aggregation result, which it then sends back to the aggregation server. The aggregation server receives the local fragment aggregation results from each edge node, uses a secret-sharing reconstruction algorithm to accumulate all the local fragment aggregation results, reconstructs the global aggregation gradient of the federation family, updates the current federation family's aggregation model based on the global aggregation gradient, obtains the updated aggregation model, and distributes the updated aggregation model to the corresponding edge nodes of the federation family.
[0060] Let the first There are a total of [number] federal families. The edge node, the first edge nodes The uploaded encrypted noisy gradient is The aggregation server splits it into Random partitions , .
[0061] The aggregation server will randomly shard the data. Distributed to the The first federal family edge nodes , Locally held The local sharding aggregation result is obtained by summing the random shards of all edge nodes in each federation family. , .
[0062] The aggregation server receives the first The local fragment aggregation results returned by each edge node of the federation family are summed using a secret-sharing reconstruction algorithm to reconstruct the fragment aggregation result. Global aggregation gradient of each federal family , .
[0063] This embodiment introduces a secret sharing mechanism. The aggregation server no longer directly holds the complete ciphertext gradient, but instead breaks it down into fragments. Even if any single edge node or external attacker intercepts a fragment, they cannot reconstruct the original ciphertext gradient. Only after the aggregation server has collected the local fragment aggregation results from all nodes can the final global gradient be reconstructed, thus achieving "double insurance."
[0064] Step 106: After receiving the updated aggregated model, the edge nodes use the local dataset they hold to perform personalized fine-tuning, thereby achieving model evolution.
[0065] Specifically, after receiving the updated aggregated model, edge nodes can use their local datasets to fine-tune it, thereby enabling model evolution.
[0066] In one example, after receiving the updated aggregation model, the edge node splits it into a frozen backbone network for extracting general industrial features and a personalized classification head for the local environment of the device. The parameters of the frozen backbone network remain unchanged, and the parameters of the personalized classification head are fine-tuned using local personalized data in the local dataset.
[0067] The personalized fine-tuning loss function used in the personalized fine-tuning process Represented as: ; in, Parameters representing personalized classification headers, The task loss function for the production tasks of the factory and production process where the edge node is located. This represents the parameters of the updated aggregation model. This represents local personalized data. For proximal constraint terms, To adjust the proximal balance coefficient for the specific intensity of a plant-specific policy, .
[0068] In one example, the aggregation server treats newly registered edge nodes and edge nodes in the local dataset with a sample count below a preset holding threshold as sparse nodes, and calls a mature aggregation model as a teacher model for the sparse nodes to perform knowledge distillation training in order to achieve the first local update.
[0069] like Figure 4 As shown, the aggregation server selects nodes from all federations with the highest task similarity to sparse nodes based on the device type distribution and historical incident type distribution of sparse nodes. A federal tribe, based on this The latest aggregation model of each federal family constitutes the teacher model set. And distribute it to the edge nodes.
[0070] in, , The first task similarity with sparse nodes represents the task similarity. The latest aggregation model of high-level federation families, namely the first A teacher model.
[0071] Sparse nodes use the base model as the student model It utilizes its local dataset to perform knowledge distillation training on the student model to achieve the first local update.
[0072] Knowledge distillation training loss function Represented as: ; in, Let cross-entropy be the loss function. For the output of the student model, The soft labels for the local dataset of the edge nodes. This is a preset balance coefficient used to balance annotation supervision and knowledge distillation. Relative entropy loss function, For the Softmax function, The preset distillation temperature parameter is used to smooth the probability distribution. For the first The output of the teacher model.
[0073] Traditional knowledge distillation often uses only a single teacher model. In multi-process industrial scenarios, sparse nodes may have requirements for both leak detection and over-temperature detection. By introducing an integrated set of teacher models and dynamic weights, sparse nodes can absorb safety knowledge from multiple process scenarios in the initial stage, significantly improving the success rate of cold starts.
[0074] In one example, the aggregation server also needs to periodically (e.g., monthly) recalculate the task similarity between edge nodes and dynamically adjust the division of federation families.
[0075] In one example, we compared the performance of the cloud-edge collaborative industrial security model evolution method based on federated learning proposed in this embodiment (this method) with the traditional FedAvg method. Figure 5 Displaying edge nodes in the federation group Group_A The comparison of model performance improvements Figure 6 Showing sparse nodes A comparison of cold start performance. From Figure 5 , Figure 6As can be seen, this method achieves a better performance improvement, reducing the cold start time of edge nodes from 5 weeks to 2 weeks.
[0076] This embodiment proposes a cloud-edge collaborative industrial security model evolution method based on federated learning, which has at least the following improvements compared with traditional federated learning methods.
[0077] First, it achieves end-to-end protection for data privacy compliance. This embodiment combines local privacy noise addition, symbol masking obfuscation encryption, and cloud-based secure multi-party computation to ensure that sensitive industrial data exists in ciphertext or perturbed gradient form throughout its entire lifecycle of training, transmission, and aggregation. This completely blocks the possibility of original data leakage and effectively meets the compliance requirements for industrial data classification and hierarchical management.
[0078] Second, it achieves anti-interference performance optimization under heterogeneous working conditions. This embodiment introduces a task similarity-aware federated clustering mechanism, abandoning the traditional one-size-fits-all aggregation mode. It accurately clusters from multiple dimensions such as equipment type distribution, historical accidents and gradient geometry direction, effectively suppressing the "negative migration" phenomenon caused by cross-heterogeneous working conditions, and significantly improving the model evolution accuracy of each edge node in complex industrial scenarios.
[0079] Third, a deep balance between group generalization and "one policy per factory" is achieved. This embodiment adds a local fine-tuning strategy based on near-end regularization constraints after the edge node receives the updated aggregation model. This strategy not only fully incorporates the group security knowledge shared within the federation family, but also achieves personalized evolution by adapting to local personalized data, perfectly balancing the model's generalization ability with the specific needs of the local production environment.
[0080] Fourth, it enables rapid cold start for sparse nodes with low barriers to entry. For newly connected nodes or edge nodes with extremely scarce data, this embodiment introduces a knowledge distillation mechanism, which uses mature teacher models in the federation family to quickly train student models in a targeted manner. This significantly shortens the cycle from random initialization to industrial usability, and significantly reduces the threshold and annotation cost of industrial intelligent deployment.
[0081] Fifth, it provides adaptive dynamic evolution capabilities for industrial scenarios. This embodiment sets up a periodic task similarity recalculation mechanism and a federation dynamic adjustment mechanism within the aggregation server, which can perceive the industrial data distribution drift caused by production process upgrades and equipment changes in real time, ensuring that the model always adapts to the current production environment and ensuring the robustness of the industrial safety system in long-term operation.
[0082] The steps described above are merely for clarity in describing the technical solution. In actual implementation, they can be combined into one step, or certain steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Any insignificant modifications or designs added to the algorithm or process, as long as they do not change the core of the algorithm or process, are also within the scope of protection of this application.
[0083] Another embodiment of this application provides an electronic device, such as Figure 7 As shown, it includes a processor 201 and a memory 202. The memory 202 stores instructions that the processor 201 can execute. When the processor 201 is configured to execute the instructions, the electronic device can implement a cloud-edge collaborative industrial security model evolution method based on federated learning as described in the above method embodiment.
[0084] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges. The bus can connect various circuits of one or more processors and memories, as well as other circuits such as peripherals, voltage regulators, and power management circuits—all well-known in the art and therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which also receives and transmits data to the processor.
[0085] The processor manages the bus and handles general processing, providing various functions, including but not limited to timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory, on the other hand, is used to store data used by the processor during operation.
[0086] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a cloud-edge collaborative industrial security model evolution method based on federated learning as described in the above method embodiments.
[0087] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0088] It will be understood by those skilled in the art that the above embodiments are specific implementations of this application, and various changes in form and detail can be made in practical applications without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A cloud-edge collaborative industrial security model evolution method based on federated learning, implemented using an aggregation server deployed in the cloud and edge nodes deployed in different factories, characterized in that... The method includes: A basic model is pre-configured in the aggregation server, and multiple edge nodes register themselves to the cloud and hold local datasets. The aggregation server then distributes the basic model to the edge nodes registered in the cloud. The aggregation server collects metadata from each edge node and calculates the task similarity between any two edge nodes based on the metadata. Edge nodes with task similarity greater than a preset similarity threshold are grouped into the same federation family. The metadata includes device type distribution, historical accident type distribution, and first local update gradient. The aggregation server traverses each federation, performs the first aggregation based on the local update gradient of each edge node of the current federation, obtains the aggregation model of the current federation, and distributes the aggregation model of the current federation to each edge node of the current federation. Edge nodes use their local datasets to train the received aggregated model locally, calculate the noisy gradient, and upload the encrypted noisy gradient to the aggregation server. The aggregation server uses secure multi-party computation and secret sharing technology to aggregate the noisy gradients uploaded by each edge node of the same federation family to obtain an updated aggregation model, and then distributes the updated aggregation model to each edge node of the corresponding federation family. After receiving the updated aggregate model, the edge nodes use their local datasets to fine-tune it individually, thereby enabling model evolution. The aggregation server uses newly registered edge nodes and edge nodes with fewer than a preset holding threshold in the local dataset as sparse nodes. It calls a mature aggregation model as a teacher model for the sparse nodes to perform knowledge distillation training to achieve the first local update. The aggregation server also periodically recalculates the task similarity between edge nodes and dynamically adjusts the division of federation families.
2. The cloud-edge collaborative industrial security model evolution method based on federated learning according to claim 1, characterized in that, The local dataset held by the edge nodes centrally includes local basic data and local personalized data. Local basic data refers to the common data collected by each edge node, including general image data collected by high-definition cameras, general sensor data collected by sensors, and fault record data. Local personalized data refers to the proprietary data collected by the edge nodes in their respective factories and production processes, including the operating status parameter sequence of specific production equipment, proprietary environmental monitoring data for specific process environments, and proprietary image data with specific industrial accident characteristics. For chemical plants, localized personalized data includes temperature and pressure data of reactors, concentration data of special gases, and infrared thermal imaging images of pipeline leaks. For machinery manufacturing plants, local personalized data includes high-frequency vibration signals of machine tool spindles, current and voltage data and spatter data of the welding process, and visual images of surface defects of specific materials; For power energy plants, local personalized data includes ultrasonic signals of partial discharge in transformers and infrared heating sequences of power generation equipment contacts.
3. The cloud-edge collaborative industrial security model evolution method based on federated learning according to claim 2, characterized in that, The aggregation server collects metadata from each edge node and calculates the task similarity between any two edge nodes based on the metadata. Edge nodes with task similarity greater than a preset similarity threshold are grouped into the same federation family, including: After receiving the base model, the edge node performs the first local training on the base model based on the local base data in the local dataset, according to the set first local training round, and obtains the first local update gradient. Edge nodes extract device type distribution probability vectors and historical accident type distribution probability vectors from local datasets, and package them together with the first local update gradient as metadata, which is then encrypted and uploaded to the aggregation server. After collecting the metadata uploaded by each edge node, the aggregation server calculates the task similarity between any two edge nodes based on the metadata, and assigns edge nodes with task similarity greater than a preset similarity threshold to the same federation family, ultimately resulting in several federation families. Denote any two edge nodes and The task similarity between them is The calculation formula is as follows: ; in, This represents the total number of layers in the base model. and They represent and In the The first local gradient update of the layer, Indicates the first The cosine weighting parameter of the layer, and They represent and The device type distribution probability vector, and They represent and The probability vector of historical accident types. This represents finding the Jason-Shannon divergence between two probability vectors. , , All are preset similarity weight coefficients.
4. The cloud-edge collaborative industrial security model evolution method based on federated learning according to claim 3, characterized in that, The aggregation server traverses each federation, performs the first aggregation based on the local update gradients of each edge node in the current federation, obtains the aggregation model of the current federation, and distributes the aggregation model of the current federation to each edge node of the current federation, including: The aggregation server traverses each federation family, calculates the mean of the local update gradient of each edge node in the current federation family as the reference gradient direction, and orthogonally projects the first local update gradient of each edge node in the current federation family onto the reference gradient direction to extract the common gradient components that are consistent with the global optimization direction of the current federation family, and then separates the orthogonal residual components that reflect the specificity of the edge nodes. The aggregation server calculates the aggregation weight of each edge node in the current federation family based on the projection magnitude of the common gradient components of each edge node. The larger the projection magnitude of the common gradient components of the edge node, the more consistent the local update of the edge node is with the global optimization direction of its federation family, and the higher the aggregation weight it is assigned. The aggregation server performs a weighted average of the common gradient components of each edge node of the current federation family based on the aggregation weight, generates the aggregation gradient of the current federation family, updates the base model using the aggregation gradient, obtains the aggregation model of the current federation family, and distributes it to each edge node of the current federation family. Suppose that there are currently a total of The edge node, the first edge nodes The first local update gradient is Then the baseline gradient direction of the current federation family The calculation formula is as follows: ; Common gradient components The following projection formula is used for calculation: ; in, This is a preset smoothing term used to prevent the denominator from being zero; Aggregate weights based on The projection amplitude is calculated using the following formula: ; in, for The projection amplitude; The aggregation server generates the aggregated gradient of the current federation family by weighted averaging of the common gradient components of each edge node based on the aggregation weights. This is achieved through the following formula: ; in, This represents the aggregation gradient of the current federation family.
5. The cloud-edge collaborative industrial security model evolution method based on federated learning according to claim 4, characterized in that, Edge nodes use their local datasets to train the received aggregated model locally, compute noisy gradients, and upload the encrypted noisy gradients to the aggregation server, including: Edge nodes use their local datasets to train the received aggregated model locally, and divide the trained gradients into shallow feature gradients and deep decision gradients. Edge nodes use an adaptive sensitivity function to calculate the clipping thresholds corresponding to the shallow feature gradient and the deep decision gradient, respectively, and inject Laplacian noise of different intensities into the shallow feature gradient and the deep decision gradient according to the clipping threshold, thereby obtaining the noisy gradient; Edge nodes generate pseudo-random streams based on the parameters of the current federated training round and the federated family to which they belong, and use the mask matrix to perform symbolic XOR obfuscation encryption on the noisy gradients, and then upload the encrypted noisy gradients to the aggregation server. remember The encrypted, noisy gradient uploaded to the aggregation server is The calculation formula is as follows: ; ; in, for The noisy gradient, This indicates the Hadamard product. For symbolic functions, For the mask matrix, For adaptive sensitivity function, It is Laplace noise. for Shallow feature gradient, for The deep decision gradient, and These are the clipping thresholds corresponding to the shallow feature gradient and the deep decision gradient, respectively. , For shallow sensitivity, For deep sensitivity, This is the shallow privacy budget coefficient. For deep privacy budgeting coefficients, .
6. The cloud-edge collaborative industrial security model evolution method based on federated learning according to claim 5, characterized in that, After receiving the updated aggregated model, edge nodes use their local datasets to fine-tune it, thereby enabling model evolution, including: After receiving the updated aggregation model, the edge nodes split it into a frozen backbone network for extracting general industrial features and a personalized classification head for the device's local environment. The parameters of the edge nodes remain unchanged while the parameters of the frozen backbone network are kept constant. The parameters of the personalized classification head are then fine-tuned using local personalized data from the local dataset. The personalized fine-tuning loss function used in the personalized fine-tuning process Represented as: ; in, Parameters representing personalized classification headers, The task loss function for the production tasks of the factory and production process where the edge node is located. This represents the parameters of the updated aggregation model. This represents local personalized data. For proximal constraint terms, To adjust the proximal balance coefficient for the specific intensity of a plant-specific policy, .
7. A cloud-edge collaborative industrial security model evolution method based on federated learning according to any one of claims 1 to 6, characterized in that, The aggregation server calls a mature aggregation model as the teacher model, which is then used by sparse nodes for knowledge distillation training to achieve the first local update, including: The aggregation server selects nodes from all federations with the highest task similarity to sparse nodes based on the device type distribution and historical incident type distribution of sparse nodes. A federal tribe, based on this The latest aggregation model of each federal family constitutes the teacher model set. And distribute it to sparse nodes; in, , The first task similarity with sparse nodes represents the task similarity. The latest aggregation model of high-level federation families, namely the first A teacher model; Sparse nodes use the base model as the student model It utilizes its local dataset to perform knowledge distillation training on the student model to achieve the first local update; Knowledge distillation training loss function Represented as: ; in, Let cross-entropy be the loss function. For the output of the student model, The soft labels for the local dataset of the edge nodes. This is a preset balance coefficient used to balance annotation supervision and knowledge distillation. Relative entropy loss function, For the Softmax function, The preset distillation temperature parameter is used to smooth the probability distribution. For the first The output of the teacher model.
8. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores instructions that the processor can execute, and the processor is configured to, when executing the instructions, enable the electronic device to implement a cloud-edge collaborative industrial security model evolution method based on federated learning as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a cloud-edge collaborative industrial security model evolution method based on federated learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Layered personalized federal learning method and device in edge computing network and medium
CN116579417A
Data and resource heterogeneous sensing cluster federated learning equipment selection method
CN118760513A