Multi-source data privacy calculation method and system based on federal learning
By performing multi-source data collection and privacy calculation on devices in the industrial Internet of Things environment, hierarchical division and incremental updates of device hardware resources are achieved, solving the problems of low training efficiency and communication bottlenecks caused by device heterogeneity, and improving the system's operating efficiency and data processing capabilities.
Patent Information
- Application Number
- CN202511157628.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the industrial Internet of Things environment, the heterogeneity of hardware resources of devices participating in federated learning leads to low training efficiency, high communication overhead, limited and unstable network bandwidth, and affects the progress and accuracy of model training.
By evaluating the hardware resources of the device, assigning tasks hierarchically, and adopting distributed key processing, privacy-enhancing perturbation processing, and homomorphic encryption technology, fuzzy data is generated for secure transmission and aggregation, and incremental updates are performed using edge computing nodes and cloud servers.
It improves the operational efficiency and security of the federated learning system, reduces communication delays and resource waste, ensures data privacy and integrity, and enhances data processing capabilities.
Smart Images

Figure CN120785642A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a multi-source data privacy computing method and system based on federated learning. BACKGROUND
[0002] In the industrial Internet of Things environment, the devices participating in federated learning (such as sensors, edge servers, data centers, etc.) have significant differences in hardware resources (CPU, GPU, memory, storage, etc.). For example, edge devices usually have limited computing power, while data centers have powerful computing resources. This heterogeneity leads to low training efficiency of federated learning, especially the devices with weak computing power significantly slow down the overall training progress, affecting the overall performance of the system. Federated learning requires frequent communication to synchronize model parameters, especially in the multi-source data scenario, the number of participants is large, and the communication overhead will increase significantly. In the industrial Internet of Things environment, the network bandwidth between devices and servers is limited, and the communication environment is unstable, which further aggravates the communication delay. Communication overhead becomes the main bottleneck of federated learning, especially in the case of large-scale device participation, frequent model updating and transmission may cause network congestion, prolong the aggregation waiting time. The data distribution on different devices is usually non-independent and identically distributed (Non-IID), which leads to large differences in the training quality of local models, and further affects the accuracy and generalization ability of the global model. SUMMARY
[0003] Therefore, it is necessary to provide a multi-source data privacy computing method and system based on federated learning to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a multi-source data privacy computing method based on federated learning, the method comprising the following steps: Step S1: In the industrial Internet of Things device environment, collecting device running state information, device production data and device environment parameters from a plurality of heterogeneous devices to obtain device multi-source privacy data; performing distributed key processing on the device multi-source privacy data to generate multi-source privacy key fragments, and each device holds part of the key fragments; Step S2: Obtain the devices participating in federated learning and the federated learning task; evaluate the device computing power, storage capacity and network condition resources of the federated learning devices to obtain device hardware resource evaluation data; divide the federated learning devices into different resource levels according to the device hardware resource evaluation data to generate device hardware resource levels; and distribute the federated learning task to devices at different levels according to the device hardware resource levels, and assign each device a unique random identifier; Step S3: Each device participating in federated learning performs privacy computation on the device multi-source privacy data locally, and performs information blurring processing on the local data through privacy-enhanced perturbation processing to generate device blurred data; the device blurred data is homomorphically encrypted through multi-source privacy key fragments to obtain device homomorphic encryption data; the device homomorphic encryption data is transmitted to the edge computing node, and the device homomorphic encryption data from multiple devices is aggregated to generate device aggregated data; Step S4: The edge computing node transmits the device aggregated data to the cloud server, and the cloud server performs incremental update on the device aggregated data to obtain industrial Internet of Things device update data; the industrial Internet of Things device update data is verified and signed and broadcast to each device for use in the next round of privacy computation.
[0005] Preferably, step S1 comprises the following steps: Step S11: In the industrial Internet of Things device environment, the running state information of each heterogeneous device is periodically collected to obtain device running state information; each heterogeneous device generates corresponding device production data according to the device running state information; Step S12: The device production data is segmented and processed, and the data is divided into multiple data blocks, each of which is processed independently, wherein the division of the data blocks is based on the time nodes of the production process, and each time node of the production process corresponds to a data block; Step S13: Obtain the device location information of the device, and collect the device environment parameters layer by layer, the first layer collects temperature and humidity parameters, the second layer collects air pressure and illumination parameters, and the third layer collects electromagnetic interference and vibration parameters, and each layer of parameters generates an independent feature vector; Step S14: Integrate and label the device running state information, device production data and device environment parameters as device multi-source privacy data; Step S15: The device multi-source privacy data is encrypted to generate multi-source privacy key fragments, wherein the encryption process includes local encryption of the data of each device to generate device-level key fragments, then grouping the device-level key fragments, each group generating a group-level key fragment, and finally summarizing into a global key fragment.
[0006] Preferably, the device computing capability, storage capacity and network condition resource assessment of the federated learning device in step S2 comprises: The computing capability of the federated learning device is evaluated by running a preset benchmark test task, recording the time required by the device to complete the task, and generating a device computing capability index; The storage capacity of the federated learning device is evaluated by scanning the storage space of the device to count the total storage capacity, used storage capacity and available storage capacity of the device; The storage capacity utilization rate of the device for federated learning is calculated, and the ratio of the used storage capacity to the total storage capacity is taken as the device storage capacity indicator; The network condition of the device for federated learning is evaluated, the network delay, data transmission rate and packet loss rate of the device within a preset time period are measured by sending data packets to a preset network test server, and the device network condition indicator is determined according to the network delay, data transmission rate and packet loss rate; The device computing capacity indicator, the device storage capacity indicator and the device network condition indicator are resource-weighted mapped to obtain device hardware resource evaluation data.
[0007] Preferably, the devices for federated learning are divided into different resource levels according to the device hardware resource evaluation data in step S2, including: The computing capacity of the device is divided into three levels according to the device computing capacity indicator, if the number of CPU cores of the device is greater than or equal to 4 and the frequency is not less than 2.5 GHz, it is divided into high computing capacity level; if the number of CPU cores of the device is 2 or 3 and the frequency is between 1.5 GHz and 2.5 GHz, it is divided into medium computing capacity level; if the number of CPU cores of the device is 1 and the frequency is less than 1.5 GHz, it is divided into low computing capacity level; The storage capacity of the device is divided into three levels according to the device storage capacity indicator, if the total storage capacity of the device is greater than or equal to 128 GB, it is divided into high storage capacity level; if the total storage capacity of the device is between 32 GB and 128 GB, it is divided into medium storage capacity level; if the total storage capacity of the device is less than 32 GB, it is divided into low storage capacity level; The network condition of the device is divided into three levels according to the device network condition indicator, if the device supports 5G network or Wi-Fi6 standard and the network bandwidth is not less than 100 Mbps, it is divided into high network condition level; if the device supports 4G network or Wi-Fi5 standard and the network bandwidth is between 50 Mbps and 100 Mbps, it is divided into medium network condition level; if the device only supports 3G network or Wi-Fi4 standard and the network bandwidth is less than 50 Mbps, it is divided into low network condition level.
[0008] Preferably, the federated learning task is assigned to devices at different levels according to the device hardware resource level in step S2, and a unique random identifier is assigned to each device, including: The resource load of the federated learning task is evaluated according to the device hardware resource level to obtain learning task load data; The learning task load data is divided into high-load learning task, medium-load learning task and low-load learning task; For high-load learning tasks, high-resource-level devices are adopted, and multiple sub-tasks are simultaneously allocated to the devices through a parallel processing strategy; For medium-load learning tasks, medium-resource-level devices are adopted, and sub-tasks are allocated to the devices in time sequence through a time-sharing processing strategy; For low-load learning tasks, low-resource-level devices are adopted, and sub-tasks are allocated to the devices in time sequence through a time-sharing processing strategy.
[0009] Preferably, in step S3, each device participating in federated learning performs privacy calculation on device multi-source privacy data locally, and performs information blurring processing on local data through a privacy-enhanced perturbation process, including: Classifying device multi-source privacy data, dividing the data into static data and dynamic data; wherein the static data includes basic configuration information and historical running data of the device, and the dynamic data includes real-time running state and real-time data in the production process; Encrypting the static data, using a symmetric encryption algorithm to generate an encrypted static data segment; Hash processing the dynamic data to generate a hash value, and binding the hash value with the dynamic data to generate a bound dynamic data segment; Adding random noise to the encrypted static data segment, and replacing the bound dynamic data segment with a hash value to generate device fuzzed data.
[0010] Preferably, in step S3, the homomorphic encryption of the device fuzzed data by the multi-source privacy key fragment includes: Format and integrity check the multi-source privacy key fragment held by each device to generate a check value; if the check value does not meet the preset standard check value, regenerate the key fragment and perform the check until the check passes; Divide the device's fuzzed data into multiple fuzzed data segments, each corresponding to a specific key fragment; Label each fuzzed data segment, wherein the label content includes data type, data purpose, and corresponding key fragment number; Homomorphic encryption processing of each fuzzed data segment using the corresponding key fragment; In the encryption process, the fuzzed data segment is divided into multiple fuzzed small blocks, each of which is independently encrypted to obtain device homomorphic encryption data.
[0011] Preferably, in step S3, the device homomorphic encryption data is transmitted to the edge computing node, and the device homomorphic encryption data from multiple devices is aggregated, including: The device homomorphic encryption data is encapsulated, and a device identifier, a timestamp and a data check code are added in the encapsulation process to generate encapsulated data; The encapsulated data is divided into a plurality of encapsulated data packets, and each encapsulated data packet contains part of the homomorphic encryption data; Before transmitting each encapsulated data packet, a check sum of the encapsulated data packet is calculated, the check sum is attached to the encapsulated data packet, and a redundant transmission mechanism is used to transmit each encapsulated data packet multiple times; The edge computing node receives the encapsulated data packet from the device, and checks each received encapsulated data packet; If a certain encapsulated data packet fails the check, a retransmission request is sent to the device; the device retransmits the encapsulated data packet until the check passes; The edge computing node reassembles all the received encapsulated data packets, and recombines the encapsulated data packets in order into complete device homomorphic encryption data according to the device identifier and the timestamp in the encapsulated data packets to obtain encapsulated reassembled data; The encapsulated reassembled data is decapsulated, and the device identifier and the timestamp of the device are extracted; The homomorphic encryption data from a plurality of devices is aggregated according to the device identifier and the timestamp, and the homomorphic encryption data of each device is added in the aggregation process to generate device aggregated data.
[0012] Preferably, step S4 comprises the following steps: Step S41: The edge computing node uses an asymmetric encryption algorithm to encrypt the device aggregated data using a public key of a cloud server to obtain device aggregated encrypted data blocks; Step S42: The device aggregated encrypted data blocks are transmitted to the cloud server one by one, and a unique sequence number is generated before each device aggregated encrypted data block is transmitted and attached to the data block; Step S43: When the cloud server receives the device aggregated encrypted data blocks, the device aggregated encrypted data blocks are sorted and integrity checked according to the sequence numbers, and the check content includes the integrity of the data blocks and the correctness of the encryption format; Step S44: The cloud server decrypts the received device aggregated data using a private key to restore the device aggregated data; Step S45: The device aggregated data is incrementally updated, the newly received aggregated data is compared with the historical data stored in the cloud server, only the changed part is updated, and industrial Internet of Things device update data is generated; Step S46: The cloud server digitally signs the industrial Internet of Things device update data using a private key to generate a signature value for the update data, and the signature value is attached to the update data. Step S47: Broadcast the industrial internet of things device update data with signature to each device, and segment the data during the broadcast process.
[0013] The present application collects device running state information, device production data and device environment parameters from multiple heterogeneous devices to obtain comprehensive device multi-source privacy data. The device multi-source privacy data is processed by distributed key, and multi-source privacy key fragments are generated, and each device holds part of the key fragments, ensuring that the privacy is effectively protected during data collection and storage, and the integrity and confidentiality of the data will not be damaged due to the leakage of a single device, providing a secure foundation for subsequent data processing and utilization. The multi-source data such as device running state information, device production data and device environment parameters are collected, covering various key information of the device during operation, and the data sources are extensive and diverse, which can fully reflect the actual operation of the device in the industrial Internet of Things environment, providing rich and complete data basis for subsequent analysis and learning, and helping to more accurately grasp the device operation rules and characteristics. After obtaining the devices participating in federated learning and the federated learning task, the device computing capability, storage capacity and network condition resource of the federated learning device are evaluated to obtain device hardware resource evaluation data. According to the device hardware resource evaluation data, the federated learning devices are divided into different resource levels, and the device hardware resource level is generated, and the federated learning task is allocated to the devices of different levels. This fine evaluation and level division based on device hardware resources realizes the reasonable allocation of federated learning tasks, so that each device can efficiently complete the task within its capability, fully utilizes the resource value of each device, avoids the waste and uneven distribution of resources, and improves the operation efficiency and resource utilization efficiency of the entire federated learning system. A unique randomized identifier is assigned to each device to effectively avoid the risk of device identity exposure during federated learning. During the process of device participating in federated learning, only the randomized identifier is used for interaction and identification, further enhancing the security and anonymity of the device, preventing the device identity information from being maliciously used, and ensuring the safe and stable operation of the entire federated learning system. Each device participating in federated learning performs privacy calculation on the device multi-source privacy data locally, and performs information fuzzing processing on the local data through privacy-enhanced perturbation processing to generate device fuzzed data. This process further enhances the privacy of the data, so that the key information of the original data is effectively hidden during transmission and sharing, and even if the data is intercepted or leaked during transmission, it is difficult for attackers to restore the true content of the original data, thereby effectively preventing the risk of data privacy leakage. The device fuzzed data is homomorphically encrypted by the multi-source privacy key fragments to obtain device homomorphic encryption data, ensuring the security and integrity of the data during transmission. The device homomorphic encryption data is transmitted to the edge computing node, and the device homomorphic encryption data from multiple devices is aggregated to generate device aggregated data. During the aggregation process, since the data has been processed by homomorphic encryption, the data of each device does not need to be decrypted during aggregation, avoiding the risk of privacy leakage of data during aggregation, while ensuring the accuracy and effectiveness of the aggregation result.The edge computing node is used to transmit the device aggregated data to the cloud server, and the device aggregated data is incrementally updated through the cloud server to obtain the industrial Internet of Things device update data. This update mechanism based on edge computing and cloud server cooperation can quickly and efficiently complete the data update processing, reduce the data transmission amount and processing time, and improve the response speed and data update efficiency of the entire system. At the same time, the incremental update method avoids repeated processing and transmission of all data, further saving system resources and communication costs. The industrial Internet of Things device update data is verified and signed and broadcast to each device for use in the next round of privacy calculation. Through the verification and signature method, the integrity and authenticity of the update data are ensured, preventing data from being tampered with or forged during transmission, and ensuring the quality and reliability of the data. After receiving the verified and signed update data, each device can use it for the next round of privacy calculation, thereby ensuring the stable operation and continuous optimization of the federated learning system, and improving the data processing and analysis level in the entire industrial Internet of Things device environment.
[0014] The present specification also provides a multi-source data privacy computing system based on federated learning, which is used to execute the multi-source data privacy computing method based on federated learning as described above, and the multi-source data privacy computing system based on federated learning comprises: A multi-source data acquisition module is used to acquire device running state information, device production data and device environment parameters from a plurality of heterogeneous devices in an industrial Internet of Things device environment to obtain device multi-source privacy data, and to perform distributed key processing on the device multi-source privacy data to generate multi-source privacy key fragments, each device holding part of the key fragments. A federated learning resource allocation module is used to obtain devices participating in federated learning and federated learning tasks, to perform resource evaluation on device computing capability, storage capacity and network conditions of the devices participating in federated learning to obtain device hardware resource evaluation data, to divide the devices participating in federated learning into different resource levels according to the device hardware resource evaluation data to generate device hardware resource levels, and to allocate the federated learning tasks to the devices at different levels according to the device hardware resource levels and to allocate a unique random identifier to each device. A privacy computing module is used to perform privacy computation on the device multi-source privacy data locally by each device participating in federated learning, to perform information blurring processing on the local data through privacy-enhanced perturbation processing to generate device blurred data, to perform homomorphic encryption on the device blurred data through the multi-source privacy key fragments to obtain device homomorphic encryption data, to transmit the device homomorphic encryption data to the edge computing node, and to aggregate the device homomorphic encryption data from a plurality of devices to generate device aggregated data. The data updating module is configured to transmit the device aggregated data to a cloud server by using the edge computing node, and perform incremental updating on the device aggregated data by the cloud server to obtain industrial Internet of Things device updating data; and the industrial Internet of Things device updating data is verified and signed and broadcast to each device for use in the next round of privacy calculation.
[0015] The multi-source data privacy calculation system based on federated learning realizes comprehensive collection of device running state information, device production data and device environment parameters from multiple heterogeneous devices through the multi-source data collection module, generates device multi-source privacy data, and performs distributed key processing to generate multi-source privacy key fragments, thereby ensuring data privacy and providing a secure basis for subsequent processing. The federated learning resource allocation module evaluates the resources of the devices participating in federated learning according to the device computing capability, storage capacity and network condition, divides the resource hierarchy and allocates tasks and random identifiers, thereby optimizing the task allocation efficiency and improving the overall operation performance of the system. The privacy calculation module performs privacy calculation and perturbation processing on the data locally to generate device fuzzification data, and performs homomorphic encryption through the multi-source privacy key fragments, thereby further strengthening the data privacy protection, and at the same time, the device homomorphic encryption data is aggregated to generate device aggregated data, thereby ensuring the security and accuracy of the data transmission and aggregation process. The data updating module transmits the device aggregated data to the cloud server by using the edge computing node for incremental updating, generates industrial Internet of Things device updating data, and broadcasts the verified and signed industrial Internet of Things device updating data to each device, thereby realizing efficient data updating and synchronization, ensuring the integrity and authenticity of the updating data, and providing reliable support for the next round of privacy calculation, thereby realizing effective utilization of multi-source data in the industrial Internet of Things device environment and efficient implementation of federated learning under the premise of ensuring data privacy, and improving the intelligent level and data processing capability of the industrial Internet of Things system. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 It is a step flowchart of a multi-source data privacy calculation method based on federated learning. Figure 2 It is Figure 1 It is a detailed implementation step flowchart of step S1. The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0017] The technical method of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0018] Furthermore, the accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application. In the drawings:
[0019] It should be understood that, although terms such as "first", "second", and so on can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be referred to as a second element, and similarly a second element can be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0020] To achieve the above object, please refer to Figures 1 to 2 A multi-source data privacy computing method based on federated learning, the method comprising the following steps: Step S1: In an industrial Internet of Things device environment, collecting device running state information, device production data and device environment parameters from a plurality of heterogeneous devices to obtain device multi-source privacy data; performing distributed key processing on the device multi-source privacy data to generate multi-source privacy key fragments, each device holding part of the key fragments; Step S2: Obtain devices participating in federated learning and federated learning tasks; evaluate the device computing capability, storage capacity and network condition resources of the devices participating in federated learning to obtain device hardware resource evaluation data; divide the devices participating in federated learning into different resource levels according to the device hardware resource evaluation data to generate device hardware resource levels; allocate the federated learning tasks to the devices in different levels according to the device hardware resource levels, and assign each device a unique random identifier; Step S3: Each device participating in federated learning performs privacy computation on the device multi-source privacy data locally, and performs information blurring processing on the local data through privacy-enhanced perturbation processing to generate device blurred data; performs homomorphic encryption on the device blurred data through the multi-source privacy key fragments to obtain device homomorphic encryption data; transmits the device homomorphic encryption data to an edge computing node, and aggregates the device homomorphic encryption data from a plurality of devices to generate device aggregated data; Step S4: transmitting the device aggregated data to the cloud server by using the edge computing node, and updating the device aggregated data by the cloud server to obtain industrial Internet of Things device update data; verifying and signing the industrial Internet of Things device update data and broadcasting to each device for use in the next round of privacy calculation.
[0021] In the embodiment of the application, as shown in the reference Figure 1 The method comprises the following steps: Step S1: in the industrial Internet of Things device environment, collecting device running state information, device production data and device environment parameters from a plurality of heterogeneous devices to obtain device multi-source privacy data; and performing distributed key processing on the device multi-source privacy data to generate multi-source privacy key fragments, each device holding part of the key fragments; In the embodiment of the application, in the industrial Internet of Things device environment, the sensor module deployed on each industrial device is used to collect device running state information, device production data and device environment parameters. The sensor module includes temperature sensors, pressure sensors, vibration sensors, current and voltage sensors, and environmental temperature and humidity sensors, etc. These sensors acquire data in real time at a preset sampling frequency (for example, 10 times per second). The collected device running state information covers parameters such as the speed and load rate of the device; the device production data includes indicators such as the yield and defective rate in the production process; and the device environment parameters involve environmental factors such as the temperature, humidity and dust concentration in the area where the device is located. These data are integrated into device multi-source privacy data and stored in the data storage unit of each device. Then, distributed key processing is performed on the device multi-source privacy data. The Shamir distributed key generation algorithm is used to split the key into multiple key fragments. Specifically, the total length of the key is determined to be 256 bits, and the allocation parameters of the key fragments are set according to the number of participating devices and security requirements. Assuming that there are 10 heterogeneous devices participating, the threshold is set to 6, that is, the key fragments held by any 6 devices can be combined to restore the complete key. In the key generation process, a polynomial on a finite field is used to construct the key fragments, a 256-bit prime number is selected as the modulus of the finite field, a polynomial of degree 5 (threshold minus 1) is generated based on the prime number, and the coefficients of the polynomial are generated by a secure random number generator. The values of the polynomial at 10 different points on the finite field are taken as the key fragments, which are distributed to the 10 devices. Each device only holds one key fragment, which is stored in the local secure storage area, and the storage area is encrypted and protected by a hardware encryption module to prevent the key fragments from being illegally read or tampered with.
[0022] Step S2: Obtain the devices participating in the federated learning and the federated learning task; perform resource evaluation on the devices participating in the federated learning, including device computing capability, storage capacity and network condition, to obtain device hardware resource evaluation data; divide the devices participating in the federated learning into different resource levels according to the device hardware resource evaluation data, and generate a device hardware resource level; and distribute the federated learning task to the devices in different levels according to the device hardware resource level, and assign a unique random identifier to each device; In the embodiment of the present application, the device management system sends a device information query instruction to each industrial Internet of Things device. After receiving the instruction, the device returns the basic information of the device through its built-in device management module, including the device model, hardware configuration, network interface type, etc. These information is transmitted back to the device management system in JSON format through a predefined communication protocol. At the same time, specific federated learning tasks are defined in the federated learning task management system, including task type, task target, task data requirement, and task execution time window, etc. The task information is stored in the database of the task management system in a structured data form. Then, the device management system sends a computing capacity test instruction to each device. The device starts the pre-installed computing performance test program, executes the preset computing task such as matrix multiplication or AES encryption algorithm operation, and records the task execution time before returning to the device management system. The device management system calculates the computing performance indicators of the device according to the task complexity and execution time, for example, the floating point operations per second (FLOPS) of the device is calculated by dividing the total operation amount of the task by the task execution time. The device management system sends a storage capacity query instruction to the device. The device calculates the total capacity and available capacity of the storage medium through the file system management module and returns the data in bytes. The device management system records the storage capacity data of each device. The device management system sends a network bandwidth test instruction to the device. The device transmits a fixed-size data file between the network interface and the device management system and records the transmission time. The device management system calculates the network bandwidth of the device according to the transmission time and the size of the data file, with the unit of bits per second (bps). At the same time, the device management system sends a network delay test instruction to the device. The device sends a network request with a timestamp to the device management system through the network interface and records the round-trip time (RTT) of the request, with the unit of milliseconds (ms). The device management system records the network bandwidth and network delay data of each device. Subsequently, the device management system divides the devices into three hardware resource levels according to the preset grading standard, i.e. high, medium and low. The high resource level device has a computing performance indicator greater than 1000 GFLOPS, an available storage capacity greater than 1TB, a network bandwidth greater than 100Mbps and a network delay less than 10ms. The medium resource level device has a computing performance indicator between 500 and 1000 GFLOPS, an available storage capacity between 500GB and 1TB, a network bandwidth between 50 and 100Mbps and a network delay between 10 and 50ms. The low resource level device has a computing performance indicator less than 500 GFLOPS, an available storage capacity less than 500GB, a network bandwidth less than 50Mbps and a network delay greater than 50ms. The device management system traverses the device list and checks the computing performance indicator, storage capacity and network condition of each device one by one. According to the grading rule, the device is assigned to the high, medium or low resource level, and the device level information is stored in the device management system database.Finally, the device management system allocates tasks to devices at different levels according to the requirements of the federated learning task and the hardware resource levels of the devices, high-resource-level devices are allocated computationally intensive tasks, medium-resource-level devices are allocated medium-complexity tasks, and low-resource-level devices are allocated lightweight tasks. The device management system generates a task allocation table according to the task allocation results, recording information such as the type of task each device is assigned, task parameters, and task execution time window; the device management system uses a secure random number generator to generate a 128-bit random number for each device and converts it to a hexadecimal string form as a unique random identifier. The identifier is bound to the hardware resource level of the device and the allocated task information and stored in the device management system database. The random identifier is sent to the corresponding device through the network communication module, and the device stores it in the local secure storage area for subsequent identity authentication and data interaction in the federated learning process.
[0023] Step S3: Each device participating in federated learning performs privacy computation on the device multi-source privacy data locally, and performs information blurring processing on the local data through privacy-enhanced perturbation processing to generate device blurred data; the device blurred data is homomorphically encrypted through multi-source privacy key fragments to obtain device homomorphic encryption data; the device homomorphic encryption data is transmitted to the edge computing node, and the device homomorphic encryption data from multiple devices is aggregated to generate device aggregated data; In the embodiment of the present application, each device participating in federated learning performs privacy calculation on the device multi-source privacy data locally. By calling a local privacy protection module, differential privacy technology is used to perform information fuzzing on the data. The privacy protection module adds noise conforming to Laplace distribution to each data item according to a preset privacy parameter ε value of 1.0. The generation formula of the Laplace noise is Laplace(0, 1 / ε), where 0 is the mean value and 1 / ε is the scale parameter. The device adds each data item to the corresponding Laplace noise to obtain device fuzzed data and stores the device fuzzed data in a local encrypted storage area. Then the device calls a local encryption module to perform homomorphic encryption on the device fuzzed data using the multi-source privacy key fragments held by the device. The Paillier homomorphic encryption algorithm is used, which supports the homomorphic properties of addition and multiplication operations. The encryption module generates the public key and private key fragments required for homomorphic encryption according to the key fragments held by the device. For each device fuzzed data item m, the encryption process is to generate a random number r in a finite field, calculate the homomorphic encryption data c equal to m times the nth power of g multiplied by n squared, where g and n are public key parameters of the Paillier algorithm, n is the product of two large prime numbers, m is the device fuzzed data item, and c is the encrypted device homomorphic encryption data. The device stores the generated device homomorphic encryption data in a local encrypted cache area. The device transmits the device homomorphic encryption data to the edge computing node using the TLS encryption protocol through a secure communication module. After receiving the device homomorphic encryption data from multiple devices, the edge computing node calls an aggregation module to perform aggregation processing on the data. The aggregation module performs addition operation on the device homomorphic encryption data of multiple devices based on the homomorphic properties of the Paillier homomorphic encryption algorithm. The specific operation is that for the homomorphic encryption data ci from device i, the aggregation module calculates C equal to the product of all ci modulo n squared, where N is the number of devices participating in aggregation. The generated C is the device aggregated data, which is still in an encrypted state and can be aggregated without decryption, thereby completing the fusion processing of multi-source data under the premise of protecting data privacy.
[0024] Step S4: transmitting the device aggregated data to the cloud server using the edge computing node, and performing incremental updating of the device aggregated data through the cloud server to obtain industrial Internet of Things device update data; verifying and signing the industrial Internet of Things device update data and broadcasting the same to each device for use in the next round of privacy calculation.
[0025] In the embodiment of the present application, the edge computing node transmits the device aggregation data to the cloud server through a secure network communication protocol, for example, using the TLS encryption protocol. After receiving the device aggregation data, the cloud server calls the data updating module and processes the device aggregation data using the incremental updating algorithm. Based on the preset updating parameters, for example, the learning rate a is set to 0.01, the incremental updating algorithm performs weighted average calculation on the device aggregation data, fuses the newly aggregated data with the historical data stored in the cloud server, and generates the industrial Internet of Things device updating data. The update process is as follows: for the historical data vector H stored in the cloud server and the device aggregation data vector A, the update data vector U is calculated, wherein each element of U is equal to the weighted sum of the corresponding elements in H and A, that is, U[i] = (1 - a) x H[i] + a x A[i], i represents the index position of the vector. Subsequently, the cloud server calls the verification signature module, uses the preset private key to digitally sign the industrial Internet of Things device updating data, the signature algorithm uses the RSA algorithm, the private key length is 2048 bits, and the signature data S is generated. The cloud server encapsulates the industrial Internet of Things device updating data and the corresponding signature data S as a broadcast data packet, and sends the broadcast data packet to each device participating in federated learning through the network broadcast module using the UDP broadcast protocol for use by the device in the next round of privacy calculation.
[0026] As an example of the present application, reference is made to Figure 2 In this example, the step S1 includes: Step S11: In the industrial Internet of Things device environment, the running state information of each heterogeneous device is periodically collected to obtain device running state information; each heterogeneous device generates corresponding device production data according to the device running state information; Step S12: The device production data is processed by segmentation, and the data is divided into multiple data blocks, each data block is processed independently, wherein the division of the data block is based on the time node of the production process, and each production process time node corresponds to a data block; Step S13: Obtain the device location information of the device, and collect the device environment parameters layer by layer, the first layer collects temperature and humidity parameters, the second layer collects air pressure and illumination parameters, and the third layer collects electromagnetic interference and vibration parameters, and each layer of parameters generates an independent feature vector; Step S14: Integrate and label the device running state information, device production data and device environment parameters as device multi-source privacy data; Step S15: encrypting the device multi-source privacy data to generate multi-source privacy key fragments, wherein the encryption process includes locally encrypting the data of each device to generate device-level key fragments, grouping the device-level key fragments, generating group-level key fragments for each group, and finally aggregating the global key fragments.
[0027] In the embodiment of the present application, in the industrial Internet of Things device environment, first, step S11 is performed, and the running state information of the device is collected by the sensor module deployed on each heterogeneous device at a preset period (for example, every 5 minutes). The collection content includes the key parameters of the device such as the rotating speed, the load rate, the current voltage and the like, which are stored as the device running state information. Subsequently, each heterogeneous device generates corresponding device production data according to the collected running state information and through the built-in data processing unit thereof according to a preset production data generation rule, for example, the production efficiency data calculated according to the device load rate and the rotating speed. In step S12, the generated device production data is processed in segments. According to the time nodes of the production process, the device production data is divided into multiple data blocks, each data block corresponding to a time node of the production process, for example, taking the start and end time of the production batch as the division basis, the data of each batch is processed as an independent data block, and the size of the data block is dynamically adjusted according to the complexity of the production process and the data generation rate. In step S13, the device location information of the device is obtained, and the device environment parameters are collected in layers through a multi-layer sensor network. The first layer of sensors collects the temperature and humidity parameters around the device, the second layer of sensors collects the air pressure and illumination parameters, and the third layer of sensors collects the electromagnetic interference and vibration parameters. The parameter data collected by each layer is respectively processed by a data preprocessing module to generate independent feature vectors, the dimension of the feature vector is determined according to the number and accuracy of the collected parameters, for example, the feature vector generated by the temperature and humidity parameters has a dimension of 2, the feature vector generated by the air pressure and illumination parameters also has a dimension of 2, the feature vector generated by the electromagnetic interference and vibration parameters has a dimension of 2, and the element values of each feature vector are normalized to ensure the consistency of the data. In step S14, the device running state information, the device production data and the feature vectors of the device environment parameters are integrated to form complete device multi-source privacy data. The integration process is completed by a data fusion module, which associates the data from different sources according to the time sequence and the device identifier according to a preset data structure, generates structured device multi-source privacy data, and adds a unique data identifier to it for subsequent processing. Finally, in step S15, the device multi-source privacy data is encrypted. First, each device calls a local encryption module to encrypt the locally generated device multi-source privacy data using a symmetric encryption algorithm (such as AES-256) to generate a device-level key fragment with a key length of 256 bits, and a preset initial vector (IV) is used in the encryption process to ensure the randomness of the encryption. Then, the device-level key fragments are grouped according to the groups to which the devices belong, each group containing key fragments of multiple devices, and a group-level key fragment is generated within the group through a key agreement protocol (such as Diffie-Hellman). Finally, all group-level key fragments are aggregated to a central node to generate a global key fragment through a key aggregation algorithm, which is used for data decryption and verification operations in subsequent federated learning processes.
[0028] Preferably, the device computing capability, storage capacity and network condition resource evaluation of the device for federated learning in step S2 comprises: evaluating the computing capability of the device for federated learning by running a preset benchmark test task, recording the time required for the device to complete the task, and generating a device computing capability index; evaluating the storage capacity of the device for federated learning by scanning the storage space of the device, and counting the total storage capacity, used storage capacity and available storage capacity of the device; calculating the storage capacity utilization of the device for federated learning, and taking the ratio of the used storage capacity to the total storage capacity as the device storage capacity index; evaluating the network condition of the device for federated learning by sending data packets to a preset network test server, measuring the network delay, data transmission rate and packet loss rate of the device within a preset time period, and determining the device network condition index according to the network delay, data transmission rate and packet loss rate; mapping the device computing capability index, device storage capacity index and device network condition index with resource weighting to obtain device hardware resource evaluation data.
[0029] In the embodiment of the present application, when evaluating the resources of the devices participating in federated learning, first, the computing power is evaluated, the device management system sends instructions to each participating device to start a preset benchmarking task, which includes a series of complex computing operations such as matrix operations and encryption algorithm processing. The device records the total time from the start to the completion of the task in seconds. The device management system collects these time data and calculates the computing power index of the device, which is expressed as the number of operations per second. The formula is the total number of operations of the benchmark task divided by the completion time. Then the storage capacity of the device is evaluated. The device management system instructs the device to scan its storage space and counts the total storage capacity, used storage capacity and available storage capacity in bytes. After the device returns these data, the device management system calculates the storage capacity utilization rate, which is the ratio of used storage capacity to total storage capacity, resulting in a value between 0 and 1, as the device storage capacity index. Then the network condition is evaluated. The device management system sends instructions to the device to send fixed-size data packets to the network test server within a preset time period, for example 60 seconds. The device records the network delay, data transfer rate and packet loss rate during this period. The network delay is in milliseconds, the data transfer rate is in bits per second, and the packet loss rate is in percentage. The device management system determines the device network condition index according to these parameters. The specific calculation method is to weight the inverse of the network delay, the data transfer rate and the negative value of the packet loss rate and sum them up. The weights are 0.4, 0.5 and 0.1 respectively. Finally, the device management system maps the device computing power index, device storage capacity index and device network condition index to the resource weight, calculates the device hardware resource evaluation data, and allocates weights as follows: computing power index 0.5, storage capacity index 0.3 and network condition index 0.2. The final device hardware resource evaluation data is obtained by weighted summation. This data is used for subsequent device resource allocation and task scheduling.
[0030] Preferably, the step S2 of classifying the devices participating in federated learning into different resource levels according to the device hardware resource evaluation data comprises: According to the device computing power index, the computing power of the device is divided into three levels. If the number of CPU cores of the device is greater than or equal to 4 and the frequency is not less than 2.5 GHz, it is classified as high computing power level; if the number of CPU cores of the device is 2 or 3 and the frequency is between 1.5 GHz and 2.5 GHz, it is classified as medium computing power level; if the number of CPU cores of the device is 1 and the frequency is less than 1.5 GHz, it is classified as low computing power level. According to the device storage capacity index, the storage capacity of the device is divided into three levels. If the total storage capacity of the device is greater than or equal to 128 GB, it is classified as a high storage capacity level. If the total storage capacity of the device is between 32 GB and 128 GB, it is classified as a medium storage capacity level. If the total storage capacity of the device is less than 32 GB, it is classified as a low storage capacity level. According to the device network condition index, the network condition of the device is divided into three levels. If the device supports 5G network or Wi-Fi6 standard, and the network bandwidth is not less than 100 Mbps, it is classified as a high network condition level. If the device supports 4G network or Wi-Fi5 standard, and the network bandwidth is between 50 Mbps and 100 Mbps, it is classified as a medium network condition level. If the device only supports 3G network or Wi-Fi4 standard, and the network bandwidth is less than 50 Mbps, it is classified as a low network condition level.
[0031] In the embodiment of the present application, when the devices are resource classified in the federated learning environment, the device management system first sends instructions to each participating device through a preset communication protocol (such as MQTT or CoAP), requiring the device to return the core number and main frequency information of its CPU. After receiving the instruction, the device reads the hardware parameters of the CPU through its hardware management module and returns these information to the device management system in a structured data format (such as JSON). After receiving the data, the device management system classifies the computing power of the device: if the CPU core number of the device is greater than or equal to 4 and the main frequency is not less than 2.5 GHz, the device is classified as high computing power level; if the CPU core number of the device is 2 or 3 and the main frequency is between 1.5 GHz and 2.5 GHz, the device is classified as medium computing power level; if the CPU core number of the device is 1 and the main frequency is less than 1.5 GHz, the device is classified as low computing power level. Then, the device management system sends instructions to the device, requiring it to scan the local storage space and return the total storage capacity information. The device calculates the total capacity of the storage medium through its storage management module and returns the data to the device management system in bytes. The device management system classifies the storage capacity of the device according to the total storage capacity returned: if the total storage capacity of the device is greater than or equal to 128 GB, it is classified as high storage capacity level; if the total storage capacity of the device is between 32 GB and 128 GB, it is classified as medium storage capacity level; if the total storage capacity of the device is less than 32 GB, it is classified as low storage capacity level. Finally, the device management system sends instructions to the device, requiring it to communicate with the preset network test server to detect the network standard supported by the device and the network bandwidth. The device establishes a connection with the network test server through its network interface module, the test server sends a fixed size of data packet to the device, the device measures the time of receiving the data packet and calculates the network bandwidth, and detects the network standard (such as 3G, 4G, 5G or Wi-Fi version) supported by the device. The device returns the network standard and network bandwidth information to the device management system. The device management system classifies the network conditions of the device according to these information: if the device supports 5G network or Wi-Fi6 standard and the network bandwidth is not less than 100 Mbps, it is classified as high network condition level; if the device supports 4G network or Wi-Fi5 standard and the network bandwidth is between 50 Mbps and 100 Mbps, it is classified as medium network condition level; if the device only supports 3G network or Wi-Fi4 standard and the network bandwidth is less than 50 Mbps, it is classified as low network condition level. The device management system stores the classification results of the computing power, storage capacity and network conditions of the device in the database, providing a basis for subsequent resource allocation and task scheduling.
[0032] Preferably, the step S2 of allocating the federated learning task to the devices at different levels according to the hardware resource levels of the devices and assigning a unique randomized identifier to each device comprises: According to the device hardware resource level, the resource load of the federated learning task is evaluated to obtain learning task load data; The learning task load data is divided into high-load learning tasks, medium-load learning tasks, and low-load learning tasks; For high-load learning tasks, high-resource-level devices are used, and multiple sub-tasks are simultaneously allocated to the devices through a parallel processing strategy; For medium-load learning tasks, medium-resource-level devices are used, and sub-tasks are allocated to the devices in time sequence through a time-sharing processing strategy; For low-load learning tasks, low-resource-level devices are used, and sub-tasks are allocated to the devices in time sequence through a time-sharing processing strategy.
[0033] In the embodiment of the present application, in the federated learning environment, first, the resource load of the federated learning task is evaluated according to the device hardware resource level. The device management system obtains the detailed information of the federated learning task from the task management module, including task type, data volume, calculation complexity, and expected completion time and other parameters. According to these parameters, the resource load value of each task is calculated, which is calculated by a preset load evaluation formula, for example, the load value is equal to the task data volume multiplied by the calculation complexity coefficient and then divided by the expected completion time. The load evaluation module compares the calculated load value with the preset threshold value, thereby dividing the learning task load data into high-load learning tasks, medium-load learning tasks, and low-load learning tasks. Specifically, if the load value is greater than 100, it is classified as a high-load learning task; if the load value is between 50 and 100, it is classified as a medium-load learning task; and if the load value is less than 50, it is classified as a low-load learning task. For high-load learning tasks, the task allocation module selects high-resource level devices from the device resource pool, which have the characteristics of CPU core number greater than or equal to 4 and main frequency not less than 2.5GHz, total storage capacity greater than or equal to 128GB, and support 5G network or Wi-Fi6 standard and network bandwidth not less than 100Mbps. The task allocation module adopts a parallel processing strategy, splits the high-load learning task into multiple subtasks, and assigns one or more subtasks to each high-resource level device according to the computing power and storage capacity of the device, ensuring that the subtasks can be executed in parallel on multiple devices. The task allocation module dynamically adjusts the number of subtasks according to the real-time load of the device through the task scheduling algorithm, so as to fully utilize the computing power of the high-resource level device and improve the task execution efficiency. For medium-load learning tasks, the task allocation module selects medium-resource level devices from the device resource pool, which have the characteristics of CPU core number of 2 or 3 and main frequency between 1.5GHz and 2.5GHz, total storage capacity between 32GB and 128GB, and support 4G network or Wi-Fi5 standard and network bandwidth between 50Mbps and 100Mbps. The task allocation module adopts a time-sharing processing strategy, splits the medium-load learning task into multiple subtasks, and assigns the subtasks to the medium-resource level devices in time sequence. The task allocation module assigns the execution time window of the subtasks to each medium-resource level device according to the available time of the device and the priority of the task through the time scheduling algorithm, ensuring that the subtasks can be executed in sequence within the available time of the device, avoiding device overload, and ensuring the timely completion of the task. For low-load learning tasks, the task allocation module selects low-resource level devices from the device resource pool, which have the characteristics of CPU core number of 1 and main frequency less than 1.5GHz, total storage capacity less than 32GB, and only support 3G network or Wi-Fi4 standard and network bandwidth less than 50Mbps.The task allocation module also adopts a time-sharing processing strategy, splits the low-load learning task into multiple sub-tasks, and sequentially allocates the sub-tasks to the low-resource level devices in time sequence. The task allocation module allocates the execution time window of the sub-tasks for each low-resource level device through a time scheduling algorithm according to the available time of the device and the priority of the task, ensures that the sub-tasks can be executed in sequence within the available time of the device, fully utilizes the idle time of the low-resource level device, and improves the utilization rate of the device.
[0034] Preferably, in step S3, each device participating in federated learning performs privacy calculation on device multi-source privacy data locally, and performs information blurring processing on local data through privacy-enhanced perturbation processing, including: Classifying the device multi-source privacy data, the data is divided into static data and dynamic data; wherein the static data includes basic configuration information and historical running data of the device, and the dynamic data includes real-time running state and real-time data in the production process; Encrypting the static data, using a symmetric encryption algorithm to generate an encrypted static data segment; Hash processing the dynamic data to generate a hash value, and binding the hash value with the dynamic data to generate a bound dynamic data segment; Adding random noise to the encrypted static data segment, and replacing the bound dynamic data segment with a hash value to generate device blurred data.
[0035] In the embodiment of the present application, when processing device multi-source privacy data, the data is first classified, and the data is divided into static data and dynamic data. The static data includes the basic configuration information of the device, such as device model, hardware parameters, software version, etc., and historical operation data, such as past production efficiency records, fault logs, etc.; the dynamic data includes real-time running state, such as current device temperature, pressure, speed, etc., and real-time data in the production process, such as real-time yield, quality detection results, etc. For static data, symmetric encryption algorithm is used for encryption processing. The specific operation is: using AES-256 encryption algorithm, a 256-bit random key is generated. The key is generated through a secure key management module and stored in the secure storage area of the device. The static data is processed in segments, and the size of each data segment is set according to the data type and storage requirements, for example, each data segment is 1KB. Then, each static data segment is encrypted using the AES-256 algorithm to generate an encrypted static data segment. In the encryption process, a preset initial vector (IV) is used to ensure the randomness of the encryption result each time, preventing security risks caused by repeated encryption. For dynamic data, hash processing is performed. Each dynamic data item is processed using the SHA-256 hash algorithm to generate a corresponding hash value. The length of the hash value is 256 bits, represented in hexadecimal string form. The generated hash value is bound with the dynamic data item to form a bound dynamic data segment. The binding process is completed through a data encapsulation module, and the hash value is used as an additional attribute of the dynamic data item to ensure data integrity and consistency. After completing the encryption of static data and the hash binding of dynamic data, random noise is added to the encrypted static data segment. Random noise is generated by a Gaussian noise generator, and the standard deviation of the noise is set according to the sensitivity of the data, for example, set to 0.1. Random noise is added bit by bit to the encrypted static data segment to increase the fuzziness of the data and prevent data leakage. At the same time, the hash value of the bound dynamic data segment is replaced. The specific operation is: using the SHA-256 algorithm to recalculate the hash value of the bound dynamic data segment to generate a new hash value, and replacing the original hash value with the new hash value to further enhance data privacy protection. The final device fuzzification data not only retains the characteristics of the original data, but also effectively protects data privacy through encryption, hashing and noise processing.
[0036] Especially important is that each device participating in federated learning in step S3 performs privacy calculation on device multi-source privacy data locally, and the information fuzzing processing on local data through privacy-enhanced perturbation processing further includes: performing hierarchical processing on local multi-source privacy data, dividing the data into high-privacy-sensitivity data and low-privacy-sensitivity data according to privacy sensitivity; For data with high privacy sensitivity, multi-round encrypted privacy calculation is used for processing, and the key of each round of encryption algorithm is dynamically generated based on the output of the previous round; For data with low privacy sensitivity, lightweight encrypted privacy calculation is used for processing, while the original features of part of the data are retained; The local multi-source privacy data after privacy calculation is divided into multiple dynamic-size multi-source privacy data blocks, wherein the size of the multi-source privacy data block is adaptively adjusted according to the complexity of the data and the result of the privacy calculation; Each multi-source privacy data block is independently subjected to privacy enhancement processing, and is subjected to privacy-enhanced perturbation processing. In the perturbation processing, the multi-source privacy data block is subjected to multiple iterations of perturbation, and the perturbation strength is adjusted according to the perturbation result of the previous iteration each time to generate device fuzzification data.
[0037] In the embodiment of the present application, when processing local multi-source privacy data, the data is first processed in layers. The data is divided into high privacy sensitivity data and low privacy sensitivity data according to privacy sensitivity. The division of privacy sensitivity is based on the type and use of the data. For example, data related to user identity information and financial data are classified as high privacy sensitivity data, while data such as device model and environmental temperature are classified as low privacy sensitivity data. For high privacy sensitivity data, multi-round encrypted privacy calculation is used for processing. The first round of encryption uses the AES-256 algorithm to generate a 256-bit initial key for encrypting the data. Subsequently, the key for each round of encryption is dynamically generated based on the output of the previous round of encryption. The previous round of ciphertext is processed through a hash function (such as SHA-256) to generate a new key for the next round of encryption. The entire process is repeated for 3 rounds to ensure high security of the data. The data after each round of encryption is stored in a local secure storage area to prevent data leakage. For low privacy sensitivity data, lightweight encrypted privacy calculation is used for processing. The AES-128 algorithm is used to generate a 128-bit key for encrypting the data. During the encryption process, the original features of part of the data are preserved. For example, through selective encryption, only the key fields in the data are encrypted, while other non-sensitive fields remain in plaintext state. In this way, the privacy of the data is guaranteed, and the usability of the data is also preserved. After completing the privacy calculation, the processed local multi-source privacy data is divided into multiple dynamically sized multi-source privacy data blocks. The size of the data block is adaptively adjusted according to the complexity of the data and the results of the privacy calculation. Specifically, for data with high complexity (such as large data volume and multiple fields), larger data blocks are generated; for data with low complexity, smaller data blocks are generated. The size of the data block is set to a range of 1KB to 10KB, and is dynamically allocated according to the actual data situation. Each multi-source privacy data block is independently processed for privacy enhancement, and perturbation processing for privacy enhancement is performed. The perturbation processing uses differential privacy technology to achieve this by adding noise that conforms to the Laplace distribution to the data block. The initial perturbation intensity is set according to the size of the data block and the privacy requirements. For example, for a 1KB data block, the initial perturbation intensity is 0.1. During the perturbation processing, the multi-source privacy data block is iteratively perturbed multiple times. The perturbation intensity is adjusted according to the perturbation result of the previous iteration. The specific adjustment method is as follows: if the privacy of the data after the previous perturbation does not meet the preset threshold, the perturbation intensity is increased; if the privacy meets the requirements, the perturbation intensity remains unchanged or is appropriately reduced. The entire process is repeated for 5 iterations to finally generate device fuzzification data, ensuring that the data is protected while still having some usability.
[0038] Preferably, the homomorphic encryption of the device fuzzification data by the multi-source privacy key fragment in step S3 comprises: The multi-source privacy key fragments held by each device are formatted and integrity checked to generate a check value; if the check value does not conform to a preset standard check value, the key fragments are regenerated and checked until the check passes; The obfuscated data of the device is divided into multiple obfuscated data segments, each of which corresponds to a specific key fragment; Each obfuscated data segment is labeled, and the label content includes data type, data purpose, and corresponding key fragment number; For each obfuscated data segment, the corresponding key fragment is used for homomorphic encryption processing; In the encryption process, the obfuscated data segment is divided into multiple obfuscated small blocks, each of which is independently encrypted to obtain device homomorphic encryption data.
[0039] In the embodiment of the present application, in the federated learning environment, first, the multi-source privacy key fragments held by each device are formatted and integrity checked. The specific operation is: the device management system sends a key verification instruction to each device, and after the device receives the instruction, the key management module of the device performs format checking on the key fragment to ensure that the length and format of the key fragment meet the preset standard, for example, the length of the key fragment is 128 bits and the format is a hexadecimal string. Then, the device calculates the check value of the key fragment, and uses the SHA-256 hash algorithm to hash process the key fragment to generate a 256-bit check value. The device returns the check value to the device management system, and the device management system compares the returned check value with the preset standard check value. If the check value does not meet the preset standard check value, the device management system instructs the device to regenerate the key fragment, and repeats the above verification process until the verification is passed. After completing the key verification, the device divides the local blurred data into multiple blurred data segments, and the size of each blurred data segment is set according to the type of data and the privacy requirement, for example, each data segment is 512 bytes. Each blurred data segment corresponds to a specific key fragment, ensuring a one-to-one correspondence between the data segment and the key fragment. The device labels each blurred data segment, and the label content includes the data type (such as device running state data, production data, etc.), the data use (such as for model training, data sharing, etc.), and the corresponding key fragment number. The label information is attached in the form of metadata at the head of the blurred data segment. Then, the device uses the corresponding key fragment to perform homomorphic encryption processing on each blurred data segment. The Paillier homomorphic encryption algorithm is used, which supports the homomorphic properties of addition and multiplication operations. In the encryption process, the device divides the blurred data segment into multiple blurred small blocks, and the size of each blurred small block is set according to the encryption efficiency and data security, for example, each small block is 64 bytes. The device independently encrypts each blurred small block, and the specific operation is: a random number r is generated (randomly selected in a finite field), and for each blurred small block m, the homomorphic encryption data c = g^m × r^n mod n^2 is calculated, where g and n are public key parameters of the Paillier algorithm, and n is the product of two large prime numbers. Each encrypted blurred small block generates corresponding homomorphic encryption data, and the device concatenates all the homomorphic encryption data in order to form complete device homomorphic encryption data for subsequent federated learning process.
[0040] Especially important is that the homomorphic encryption of the device blurred data by the multi-source privacy key fragment in step S3 also includes: For each device participating in federated learning, a set of unique multi-source privacy key fragments is generated, each key fragment is associated with the unique identifier of the device, and is distributed to the corresponding device through a secure channel; After the device receives the key fragments, the locally generated obfuscated data is format adapted and converted into a segmented encryption structure to obtain obfuscated encryption structure data; The obfuscated encryption structure data is sequentially encrypted according to the order of the key fragments, each encryption step uses only one key fragment, and the result of each encryption operation is used as the input of the next encryption, to obtain the encryption result of each block of data; During the encryption process, the encryption result of each block of data is hashed to generate an intermediate hash value; The encryption result of each block of data is checked by comparing the intermediate hash value with a preset check value to verify the correctness of the encryption process, and the data blocks that pass the verification are recorded; All data blocks that pass the verification are recombined in the original order to form homomorphic encryption data.
[0041] In the embodiment of the present application, in a federated learning environment, a set of unique multi-source privacy key fragments is first generated for each participating device. The key management system generates key fragments based on the unique identifier of the device (such as the device ID), each key fragment is 128 bits long, and is generated using the AES encryption algorithm. The generated key fragments are distributed to the corresponding devices through a secure TLS channel to ensure the confidentiality and integrity of the key fragments during transmission. After receiving the key fragments, the device performs format adaptation processing on the locally generated obfuscated data. The obfuscated data is first divided into multiple data blocks, each data block is 256 bytes in size. The device converts each data block into a segmented encryption structure, i.e., each data block is divided into multiple small segments, each segment is 64 bytes in size. The purpose of this segmented encryption structure is to adapt to subsequent encryption operations, ensuring that each small segment can be independently encrypted, thereby obtaining obfuscated encrypted structure data. For the obfuscated encrypted structure data, the device performs encryption operations in sequence according to the order of the key fragments. The encryption process uses the AES-128 algorithm, and each encryption step uses only one key fragment. The specific operation is as follows: the first key fragment is used to encrypt the first data segment to obtain the encryption result, which is used as the input for the next encryption, and the second key fragment is used to encrypt the next data segment. This process is repeated until all data segments are encrypted, and the encryption result of each block of data is finally obtained. During the encryption process, the device performs a hash calculation on the encryption result of each block of data. The SHA-256 hash algorithm is used to process the encryption result of each block of data to generate a 256-bit intermediate hash value. The device management system pre-sets a check value, and the device compares the calculated intermediate hash value with the pre-set check value to verify the correctness of the encryption process. If the intermediate hash value is consistent with the pre-set check value, the data block passes the verification, and the device records the hash value and encryption result of the data block. Finally, the device recombines all the verified data blocks in the original order to form complete homomorphic encryption data. This process ensures the integrity of the data and the correctness of the encryption, and at the same time, through segmented encryption and hash verification, enhances the security and privacy protection of the data. The homomorphic encryption data can then be used for secure data sharing and calculation in the federated learning process.
[0042] Preferably, the step S3 of transmitting the device homomorphic encryption data to the edge computing node and performing aggregation processing on the device homomorphic encryption data from multiple devices comprises: performing encapsulation processing on the device homomorphic encryption data, adding a device identifier, a timestamp, and a data check code during the encapsulation process to generate encapsulated data; segmenting the encapsulated data into multiple encapsulated data packets, each encapsulated data packet containing part of the homomorphic encryption data; Before transmitting each encapsulated data packet, a checksum of the encapsulated data packet is calculated, the checksum is appended to the encapsulated data packet, and a redundant transmission mechanism is adopted to transmit each encapsulated data packet multiple times; At the edge computing node, encapsulated data packets from the devices are received, and each received encapsulated data packet is checked; If a certain encapsulated data packet fails the check, a retransmission request is sent to the device; the device retransmits the encapsulated data packet until the check passes; The edge computing node reassembles all received encapsulated data packets, and according to the device identifier and timestamp in the encapsulated data packet, the encapsulated data packets are recombined in order to form complete device homomorphic encryption data, obtaining encapsulated reassembled data; The encapsulated reassembled data is unpacked and the device identifier and timestamp of the device are extracted; According to the device identifier and timestamp, homomorphic encryption data from multiple devices is aggregated, and in the aggregation process, the homomorphic encryption data of each device is subjected to an additive operation to generate device aggregated data.
[0043] In the embodiment of the present application, in the federated learning environment, the homomorphic encryption data of the device is first encapsulated. The encapsulation module adds the device identifier (such as device ID), timestamp (accurate to milliseconds) and data check code to the homomorphic encryption data. The data check code is calculated by the CRC-32 algorithm, which is used to detect whether the data has errors during transmission. The encapsulated data is called encapsulated data, and its structure includes device identifier, timestamp, data check code and homomorphic encryption data. Then, the encapsulated data is divided into multiple encapsulated data packets, each of which contains part of the homomorphic encryption data. The size of the encapsulated data packet is dynamically adjusted according to the network transmission conditions and device performance, for example, the size of each encapsulated data packet is set to 1KB. During the segmentation process, it is ensured that each encapsulated data packet contains a complete data segment to avoid information loss caused by data segmentation. Before transmitting each encapsulated data packet, the checksum of the encapsulated data packet is calculated. The checksum is calculated by a simple XOR operation and is attached to the tail of the encapsulated data packet. A redundant transmission mechanism is used to transmit each encapsulated data packet multiple times, for example, each encapsulated data packet is transmitted 3 times, to improve the reliability of data transmission. After receiving the encapsulated data packet from the device, the edge computing node checks each received encapsulated data packet. The checking process is performed by calculating the checksum of the encapsulated data packet and comparing it with the attached checksum. If a certain encapsulated data packet fails the check, the edge computing node sends a retransmission request to the device. After receiving the retransmission request, the device retransmits the encapsulated data packet until the check passes. The encapsulated data packet that passes the check is stored in the temporary storage area of the edge computing node. The edge computing node recombines all received encapsulated data packets, recombposes the encapsulated data packets into complete device homomorphic encryption data in the original order according to the device identifier and timestamp in the encapsulated data packet, and obtains encapsulated recombination data. During the recombination process, the device identifier and timestamp are used to ensure the integrity and order of the data. The encapsulated recombination data is de-encapsulated to extract the device identifier and timestamp of the device. The de-encapsulation module parses the encapsulated recombination data to separate the device identifier, timestamp and homomorphic encryption data, ensuring the integrity and availability of the data. Finally, the homomorphic encryption data from multiple devices is aggregated according to the device identifier and timestamp. In the aggregation process, the addition operation of the homomorphic encryption algorithm is used to add each device's homomorphic encryption data bit by bit. For example, the homomorphic encryption data from device A and the homomorphic encryption data from device B are added bit by bit, and the result is taken modulo to ensure that the result meets the requirements of the homomorphic encryption algorithm. In this way, device aggregation data is generated for subsequent federated learning calculation.
[0044] Preferably, step S4 comprises the following steps: Step S41: Through the edge computing node, using asymmetric encryption algorithm, and with the public key of the cloud server, the device aggregation data is encrypted to obtain the device aggregation encrypted data block; Step S42: The device aggregation encrypted data block is transmitted to the cloud server one by one, and a unique serial number is generated before each device aggregation encrypted data block is transmitted and attached to the data block; Step S43: When the cloud server receives the device aggregation encrypted data block, the device aggregation encrypted data block is sorted and integrity checked according to the serial number, and the checking content includes the integrity of the data block and the correctness of the encryption format; Step S44: The cloud server decrypts the received device aggregation data, uses the private key to decrypt the data block, and restores the device aggregation data; Step S45: The device aggregation data is incrementally updated, the newly received aggregation data is compared with the historical data stored in the cloud server, only the changed part is updated, and the industrial Internet of Things device update data is generated; Step S46: The cloud server digitally signs the industrial Internet of Things device update data, uses the private key to generate a signature value for the update data, and attaches the signature value to the update data; Step S47: The industrial Internet of Things device update data with signature is broadcasted to each device, and the data is segmented and transmitted during the broadcast process.
[0045] In the embodiment of the present application, in the federated learning environment, first, step S41 is performed, and the device aggregated data is encrypted by the edge computing node. The edge computing node uses an asymmetric encryption algorithm to encrypt the device aggregated data with the public key of the cloud server. The specific operation is as follows: using the RSA encryption algorithm, the public key length is 2048 bits, the device aggregated data is divided into multiple data blocks, and the size of each data block is set to 245 bytes according to the requirement of the encryption algorithm (considering the padding mechanism of RSA encryption). Each data block is encrypted separately to generate device aggregated encrypted data blocks. In step S42, the device aggregated encrypted data blocks are transmitted to the cloud server one by one. Before transmitting each device aggregated encrypted data block, a unique serial number is generated. The serial number is generated by the UUID (universal unique identifier) generation algorithm, ensuring that each serial number is unique in the global scope. The generated serial number is attached to the head of the device aggregated encrypted data block to form a complete transmission data block. In step S43, when the cloud server receives the device aggregated encrypted data block, the data block is sorted and integrity checked according to the serial number. The checking content includes the integrity of the data block and the correctness of the encryption format. The specific operation is as follows: the cloud server first parses the serial number in the data block, and sorts the data block according to the serial number. Then, the SHA-256 hash algorithm is used to calculate the hash value of each data block. The calculated hash value is compared with the preset check value, and the format of the data block is checked to ensure that it meets the requirements of the RSA encryption algorithm, ensuring the integrity of the data block and the correctness of the encryption format. In step S44, the cloud server decrypts the received device aggregated data. The private key (length of 2048 bits) corresponding to the encryption is used to decrypt the device aggregated encrypted data block. The decryption process is performed block by block to recover the device aggregated data. In step S45, the device aggregated data is updated incrementally. The cloud server compares the newly received aggregated data with the historical data stored in the cloud server, and only updates the changed part. The specific operation is as follows: the new and old data are compared field by field, and if the field value is found to have changed, the new field value is updated to the historical data to generate industrial internet of things device update data. In step S46, the industrial internet of things device update data is digitally signed by the cloud server. The private key (length of 2048 bits) is used to generate a signature value for the update data. The signature process uses the SHA-256withRSA algorithm, which first performs SHA-256 hash calculation on the update data, and then encrypts the hash value with the private key to generate the signature value. The signature value is attached to the tail of the update data to form the industrial internet of things device update data with signature. Finally, in step S47, the industrial internet of things device update data with signature is broadcast to each device. The data is transmitted in segments during the broadcast process. The specific operation is as follows: the update data is divided into multiple data segments, and the size of each data segment is set to 1KB according to the network transmission conditions.The data segments are sent to each device one by one through the UDP broadcast protocol, ensuring efficient transmission and distribution of data.
[0046] The specification also provides a multi-source data privacy computing system based on federated learning for performing the multi-source data privacy computing method based on federated learning as described above, the multi-source data privacy computing system based on federated learning comprising: A multi-source data acquisition module is configured to acquire device running state information, device production data and device environment parameters from a plurality of heterogeneous devices in an industrial Internet of Things device environment to obtain device multi-source privacy data, and to perform distributed key processing on the device multi-source privacy data to generate multi-source privacy key fragments, each device holding part of the key fragments. A federated learning resource allocation module is configured to obtain devices participating in federated learning and federated learning tasks, to perform resource evaluation on device computing capability, storage capacity and network conditions of the devices participating in federated learning to obtain device hardware resource evaluation data, to divide the devices participating in federated learning into different resource levels according to the device hardware resource evaluation data to generate device hardware resource levels, and to allocate the federated learning tasks to the devices in different levels according to the device hardware resource levels and to allocate a unique random identifier to each device. A privacy computing module is configured to perform privacy computation on the device multi-source privacy data locally by each device participating in federated learning, to perform information blurring processing on the local data through privacy-enhanced perturbation processing to generate device blurred data, to perform homomorphic encryption on the device blurred data through the multi-source privacy key fragments to obtain device homomorphic encrypted data, to transmit the device homomorphic encrypted data to an edge computing node, and to perform aggregation processing on the device homomorphic encrypted data from a plurality of devices to generate device aggregated data. A data updating module is configured to transmit the device aggregated data to a cloud server through the edge computing node, to perform incremental updating on the device aggregated data through the cloud server to obtain industrial Internet of Things device updating data, to verify and sign the industrial Internet of Things device updating data and broadcast the same to each device for use in the next round of privacy computation.
[0047] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the application being defined by the appended claims rather than the above description, and it is intended to encompass all variations falling within the meaning and scope of the equivalent elements of the application file.
[0048] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and it is intended to embrace all such modifications and changes that fall within the scope of the application. Accordingly, the application is not to be restricted in scope to the specific embodiments disclosed herein but is to be accorded the full scope that the principles and novel features request appropriately granted.
Claims
1. A multi-source data privacy computing method based on federated learning, characterized in that: The following steps are involved: Step S1: In an industrial IoT device environment, device operating status information, device production data, and device environmental parameters are collected from multiple heterogeneous devices to obtain device multi-source private data. Distributed key processing is performed on the device multi-source private data to generate multi-source private key fragments, with each device holding a portion of the key fragments. Step S2: Obtain the devices participating in federated learning and the federated learning tasks; evaluate the computing power, storage capacity, and network condition resources of the federated learning devices to obtain device hardware resource evaluation data; Based on the device hardware resource evaluation data, the federated learning devices are divided into different resource levels to generate the device hardware resource level. Distribute federated learning tasks to devices at different levels based on their hardware resource levels, and assign a unique randomized identifier to each device. Step S3: Each device participating in federated learning performs privacy computation on its multi-source private data locally, and performs information obfuscation on the local data through privacy-enhancing perturbation processing to generate device obfuscated data. The obfuscated data is homomorphically encrypted using multi-source private key fragments to obtain device homomorphically encrypted data. The device homomorphically encrypted data is transmitted to the edge computing node, and the device homomorphically encrypted data from multiple devices is aggregated to generate device aggregated data. Step S4: Use the edge computing node to transmit the device aggregated data to the cloud server, and incrementally update the device aggregated data through the cloud server to obtain the industrial Internet of Things device update data; verify and sign the industrial Internet of Things device update data and broadcast it to each device for use in the next round of privacy computing.
2. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: In the industrial Internet of Things device environment, periodically collect the operating status information of each heterogeneous device to obtain the device operating status information; each heterogeneous device generates corresponding device production data based on the device operating status information; Step S12: Segment the equipment production data into multiple data blocks, and process each data block independently. The data blocks are divided based on the time nodes of the production process, and each time node of the production process corresponds to a data block. Step S13: Acquire the device location information of the device, and perform layered collection of device environment parameters based on the device location information. The first layer collects temperature and humidity parameters, the second layer collects air pressure and light parameters, and the third layer collects electromagnetic interference and vibration parameters. Independent feature vectors are generated for each layer of parameters. Step S14: Integrate the device operating status information, device production data, and device environmental parameters and mark them as device multi-source privacy data; Step S15: Encrypt the multi-source private data of the device to generate multi-source private key fragments. The encryption process includes locally encrypting the data of each device to generate device-level key fragments, then grouping the device-level key fragments, generating group-level key fragments for each group, and finally aggregating them into global key fragments.
3. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: The evaluation of computing power, storage capacity, and network resource conditions of the federated learning device in step S2 includes: Evaluate the computing power of federated learning devices by running pre-set benchmark tasks, recording the time required for the devices to complete the tasks, and generating device computing power indicators. Evaluate the storage capacity of federated learning devices by scanning the device's storage space and counting the device's total storage capacity, used storage capacity, and available storage capacity. Calculate the storage capacity utilization of the federated learning device, and use the ratio of used storage capacity to total storage capacity as the device storage capacity indicator; Evaluate the network conditions of federated learning devices by sending data packets to a preset network test server to measure the device's network latency, data transmission rate, and packet loss rate within a preset time period; determine the device's network condition indicators based on network latency, data transmission rate, and packet loss rate; The device computing capability quantitative indicators, device storage capacity indicators and device network condition indicators are mapped to resource weights to obtain device hardware resource evaluation data.
4. The multi-source data privacy computing method based on federated learning according to claim 3 is characterized in that: In step S2, the federated learning devices are divided into different resource levels based on the device hardware resource evaluation data, including: The computing power of a device is divided into three levels based on its computing power index. If the device has 4 or more CPU cores and a main frequency of 2.5 GHz or higher, it is classified as high computing power. If the device has 2 or 3 CPU cores and a main frequency between 1.5 GHz and 2.5 GHz, it is classified as medium computing power. If the device has 1 CPU core and a main frequency below 1.5 GHz, it is classified as low computing power. The storage capacity of a device is divided into three levels based on its storage capacity index. If the total storage capacity of the device is greater than or equal to 128GB, it is classified as high storage capacity level; if the total storage capacity of the device is between 32GB and 128GB, it is classified as medium storage capacity level; if the total storage capacity of the device is less than 32GB, it is classified as low storage capacity level; The device's network conditions are divided into three levels based on the device's network condition indicators. If the device supports 5G network or Wi-Fi 6 standard and the network bandwidth is not less than 100Mbps, it is classified as a high network condition level; if the device supports 4G network or Wi-Fi 5 standard and the network bandwidth is between 50Mbps and 100Mbps, it is classified as a medium network condition level; if the device only supports 3G network or Wi-Fi 4 standard and the network bandwidth is less than 50Mbps, it is classified as a low network condition level.
5. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: In step S2, federated learning tasks are distributed to devices at different levels according to the device hardware resource level, and a unique random identifier is assigned to each device, including: Evaluate the resource load of federated learning tasks based on the device hardware resource level to obtain learning task load data; Divide the learning task load data into high-load learning tasks, medium-load learning tasks, and low-load learning tasks; For high-load learning tasks, high-resource-level devices are used, and multiple subtasks are assigned to the devices simultaneously through parallel processing strategies; For medium-load learning tasks, medium-resource-level devices are used, and subtasks are assigned to devices in chronological order through a time-sharing processing strategy; For low-load learning tasks, low-resource-level devices are used, and subtasks are assigned to devices in chronological order through a time-sharing processing strategy.
6. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: In step S3, each device participating in federated learning performs privacy calculations on its multi-source private data locally and performs information obfuscation on the local data through privacy-enhancing perturbation processing, including: Classify the multi-source privacy data of devices into static data and dynamic data. Static data includes basic configuration information and historical operation data of the device, while dynamic data includes real-time operation status and real-time data during the production process. Encrypt static data using a symmetric encryption algorithm to generate encrypted static data segments; Perform hash processing on the dynamic data to generate a hash value, and bind the hash value to the dynamic data to generate a bound dynamic data segment; Random noise is added to the encrypted static data segment, and the hash value is replaced on the bound dynamic data segment to generate device obfuscated data.
7. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: In step S3, homomorphic encryption of the device obfuscated data using the multi-source private key fragments includes: Format and integrity check the multi-source private key fragments held by each device to generate a check value. If the check value does not meet the preset standard check value, regenerate the key fragment and verify it until the verification passes. Dividing the obfuscated data of the device into multiple obfuscated data segments, each obfuscated data segment corresponds to a specific key fragment; Marking each obfuscated data segment, where the marking content includes the data type, data purpose and corresponding key segment number; For each obfuscated data segment, homomorphic encryption is performed using the corresponding key fragment; During the encryption process, the fuzzified data segment is divided into multiple fuzzified blocks, and each fuzzified block is encrypted independently to obtain device homomorphic encrypted data.
8. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: In step S3, the device homomorphically encrypted data is transmitted to the edge computing node, and the device homomorphically encrypted data from multiple devices is aggregated, including: Encapsulate the device homomorphically encrypted data, adding a device identifier, timestamp, and data checksum to generate encapsulated data. The encapsulated data is divided into multiple encapsulated data packets, each of which contains part of the homomorphically encrypted data; Before transmitting each encapsulated data packet, the checksum of the encapsulated data packet is calculated, the checksum is appended to the encapsulated data packet, and a redundant transmission mechanism is adopted to transmit each encapsulated data packet multiple times; The edge computing node receives encapsulated data packets from the device and verifies each received encapsulated data packet; If a packet fails verification, a retransmission request is sent to the device; the device retransmits the packet until it passes verification. The edge computing node reassembles all received encapsulated data packets and, based on the device identifier and timestamp in the encapsulated data packets, reassembles the encapsulated data packets in sequence into complete device homomorphically encrypted data to obtain encapsulated and reassembled data; Decapsulate the encapsulated and reassembled data and extract the device identifier and timestamp of the device; The homomorphically encrypted data from multiple devices are aggregated according to the device identifier and timestamp. During the aggregation process, the homomorphically encrypted data of each device is added to generate device aggregate data.
9. The multi-source data privacy computing method based on federated learning according to claim 1 is characterized in that: Step S4 includes the following steps: Step S41: Using an asymmetric encryption algorithm at the edge computing node, the device aggregated data is encrypted with the public key of the cloud server to obtain a device aggregated encrypted data block; Step S42: Transmit the device aggregated encrypted data blocks one by one to the cloud server. Before each device aggregated encrypted data block is transmitted, a unique serial number is generated and attached to the data block; Step S43: When receiving the device aggregated encrypted data blocks through the cloud server, the device aggregated encrypted data blocks are sorted and integrity checked according to the serial numbers. The verification content includes the integrity of the data blocks and the correctness of the encryption format; Step S44: The cloud server decrypts the received device aggregated data, decrypts the data block using the private key, and recovers the device aggregated data; Step S45: incrementally update the device aggregated data, compare the newly received aggregated data with the historical data stored in the cloud server, and update only the changed parts to generate the updated data of the industrial Internet of Things device; Step S46: Digitally sign the IIoT device update data through the cloud server, use the private key to generate a signature value for the update data, and append the signature value to the update data; Step S47: Broadcast the signed IIoT device update data to each device, and transmit the data in segments during the broadcast process.
10. A multi-source data privacy computing system based on federated learning, characterized in that: For executing the multi-source data privacy computing method based on federated learning as claimed in claim 1, the multi-source data privacy computing system based on federated learning comprises: The multi-source data acquisition module is used to collect device operating status information, device production data, and device environmental parameters from multiple heterogeneous devices in an industrial IoT device environment to obtain multi-source private data of the devices. The module also performs distributed key processing on the multi-source private data of the devices to generate multi-source private key fragments, with each device holding a portion of the key fragments. The federated learning resource allocation module is used to obtain devices participating in federated learning and federated learning tasks; evaluate the computing power, storage capacity, and network condition resources of federated learning devices to obtain device hardware resource evaluation data; classify federated learning devices into different resource tiers based on the device hardware resource evaluation data to generate a device hardware resource hierarchy; allocate federated learning tasks to devices at different tiers based on the device hardware resource hierarchy, and assign a unique randomized identifier to each device; The privacy computing module is used by each device participating in federated learning to perform privacy computing on its multi-source private data locally, and to obfuscate the local data through privacy-enhancing perturbation processing to generate device obfuscated data; the device obfuscated data is homomorphically encrypted using multi-source private key fragments to obtain device homomorphically encrypted data; the device homomorphically encrypted data is transmitted to the edge computing node, and the device homomorphically encrypted data from multiple devices is aggregated to generate device aggregated data; The data update module is used to use edge computing nodes to transmit device aggregated data to the cloud server, and incrementally update the device aggregated data through the cloud server to obtain industrial IoT device update data; the industrial IoT device update data is verified and signed and broadcast to each device for use in the next round of privacy computing.
Citation Information
Cited By
Longitudinal federal data privacy calculation method and system based on artificial intelligence
CN121261970A