Device data privacy cooperative training method and system based on artificial intelligence
Through artificial intelligence dynamically setting acquisition windows and quantifying privacy needs, the privacy protection problem of collaborative training of multiple devices in the Internet of Things system is solved, the optimization allocation of resources and the differentiation of encryption strategies are realized, and security efficiency is improved.
Patent Information
- Application Number
- CN202510689452.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the Internet of Things system, multi-device collaborative training faces the contradiction between privacy protection costs and security effectiveness. The existing solutions have problems such as resource waste, static time window mismatch, and the lack of differentiated management of cross-device privacy collaborative training.
Based on artificial intelligence, by obtaining device security levels and data sensitivity, dynamically set acquisition windows, quantify privacy needs, divide training sets, and adopt differentiated encryption strategies to achieve dynamic hierarchical collaborative training.
Optimize resource allocation, improve the protection strength of key data, reduce equipment energy consumption, and improve the efficiency and stability of privacy collaborative training.
Smart Images

Figure CN120498674A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an artificial intelligence-based device data privacy collaborative training method and system. Background Art
[0002] In current IoT systems, the conflict between privacy protection costs and security effectiveness faced by multi-device collaborative training is becoming increasingly prominent. Existing solutions suffer from the following major technical flaws: First, when a globally unified encryption strategy is adopted, devices with high security requirements and those with low security requirements are forced to use the same strength for privacy collaborative training. For example, in smart healthcare scenarios, high-frequency vital sign data from a monitoring device and low-frequency data from a ward temperature and humidity sensor may use the same AES-256 encryption. This results in insufficient encryption strength for the former, which is exposed to the risk of arbitrary data access and tampering, while excessive encryption wastes computing resources for the latter. Second, when low-security devices use complex key rotation mechanisms, inefficient communication overhead is generated. If low-security sensors use the same key strategy as high-security industrial control equipment, the key energy consumption increases significantly, shortening device battery life. Third, static time windowing cannot adapt to the dynamic characteristics of data, resulting in misallocation of encryption resources. For example, in smart home systems, security cameras experience a surge in data sensitivity when detecting an anomaly, but existing fixed encryption windows make it difficult to increase protection strength in a timely manner. As a result, critical event data cannot be included in privacy-critical computations, and ultimately, the low-strength encryption is still used, creating security vulnerabilities. Fourth, there's a lack of differentiated management during cross-device privacy-focused collaborative training. Data from devices with high privacy requirements are forced to wait for data from devices with low privacy requirements to complete encryption, increasing overall latency and failing to meet required response standards in some usage environments. In summary, the root cause of these issues lies in the failure to establish a dynamic balance between a quantitative assessment system for privacy needs and encryption costs. Summary of the Invention
[0003] In response to the deficiencies in the prior art, the present invention provides an artificial intelligence-based device data privacy collaborative training method and system.
[0004] A device data privacy collaborative training method based on artificial intelligence includes: obtaining multiple source devices, and obtaining corresponding device security levels according to the source devices, and setting a pending time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices, and obtaining device operation data of the source devices at multiple time points within the pending time period, arranging the multiple device operation data in chronological order and forming a device data sequence corresponding to the source devices; obtaining the data sensitivity corresponding to the source devices according to the device data sequence corresponding to the source devices, and obtaining the fixed privacy requirement index corresponding to the source devices according to the data sensitivity corresponding to the source devices, correcting the fixed privacy requirement index according to the device security levels corresponding to the source devices and forming a target privacy requirement index corresponding to the source devices; dividing the multiple source devices into multiple training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training sets is less than a preset collaborative threshold; and completing the privacy collaborative training between the device data sequences corresponding to the multiple source devices in each training set in turn.
[0005] Optionally, completing the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set in sequence includes: obtaining the mean value of the privacy requirement indicator corresponding to each training set based on the target privacy requirement indicator corresponding to the multiple source devices in each training set; obtaining multiple different value ranges, where different value ranges correspond to different encryption levels; obtaining the encryption level corresponding to each training set in sequence based on the value range that each training set falls into, and completing the privacy collaborative training between the device data sequences corresponding to the multiple source devices in each training set according to the encryption level corresponding to each training set.
[0006] Optionally, completing privacy collaborative training between device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set includes: encrypting the device data sequences in the training set segment by segment using encryption logic corresponding to the encryption level, wherein the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; generating and distributing encryption keys based on the encryption level, wherein the device data sequences in each training set share key generation rules.
[0007] Optionally, the data sensitivity corresponding to the source device is obtained according to the device data sequence corresponding to the source device as follows: ;in, is the data sensitivity corresponding to the j-th source device, is the adjustment coefficient, is the jth device data sequence Equipment operation data at a point in time, is the device operation data at the first time point in the j-th device data sequence, is the number of time points in the j-th device data sequence, is the device operation data at the i+1th time point in the jth device data sequence, is the device operation data at the i-th time point in the j-th device data sequence, is the standard maximum operating data of the j-th source device, is the standard minimum operating data of the j-th source device, is a conditional parameter.
[0008] Optionally, the fixed privacy requirement indicator corresponding to the source device is obtained according to the data sensitivity corresponding to the source device and is expressed as: ;in, is the fixed privacy requirement indicator corresponding to the j-th source device, is the data sensitivity corresponding to the j-th source device.
[0009] Optionally, correcting the fixed privacy requirement indicator according to the device security level corresponding to the source device and forming the target privacy requirement indicator corresponding to the source device includes: obtaining the scaling factor corresponding to the source device according to the device security level corresponding to the source device; correcting the fixed privacy requirement indicator according to the scaling factor and forming the target privacy requirement indicator corresponding to the source device.
[0010] Optionally, the scaling factor corresponding to the source device is obtained according to the device security level corresponding to the source device and is expressed as: ;in, is the scaling factor corresponding to the j-th source device, is the device security level corresponding to the j-th source device, This is the maximum equipment safety level.
[0011] Also provided is an artificial intelligence-based device data privacy collaborative training system, which includes: a data acquisition module for acquiring multiple source devices, and acquiring corresponding device security levels based on the source devices, and setting a pending time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices, and acquiring device operation data of the source devices at multiple time points within the pending time period, arranging the multiple device operation data in chronological order and forming a device data sequence corresponding to the source devices; a data processing module for acquiring the data sensitivity corresponding to the source devices based on the device data sequence corresponding to the source devices, and acquiring the fixed privacy requirement index corresponding to the source devices based on the data sensitivity corresponding to the source devices, correcting the fixed privacy requirement index according to the device security level corresponding to the source devices and forming a target privacy requirement index corresponding to the source devices; a data partitioning module for dividing multiple source devices into multiple training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training set is less than a preset collaborative threshold; and a collaborative training module for sequentially completing privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set.
[0012] Optionally, the collaborative training module is also used to: obtain the mean privacy requirement indicator corresponding to each training set based on the target privacy requirement indicator corresponding to multiple source devices in each training set; obtain multiple different value ranges, where different value ranges correspond to different encryption levels; obtain the encryption level corresponding to each training set in turn according to the value range that each training set falls into, and complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set.
[0013] Optionally, the collaborative training module is also used to: encrypt the device data sequences in the training set segment by segment using the encryption logic corresponding to the encryption level, wherein the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; generate and distribute encryption keys based on the encryption level, wherein the device data sequences in each training set share key generation rules.
[0014] The beneficial effects of the present invention are embodied in: In the entire AI-based device data privacy collaborative training method, first, the dynamic data collection window division based on the device security level solves the resource mismatch problem of the static time window in the existing scheme. Through differentiated collection strategies, high-security devices can obtain more comprehensive timing feature support, while avoiding redundant data collection of low-security devices; further, the joint correction mechanism of data sensitivity quantification and device security level breaks through the limitations of single-dimensional evaluation. By integrating data volatility characteristics and inherent security attributes of the device, the device can automatically trigger enhanced privacy protection during burst data transmission, which greatly improves the protection strength of key data compared with the existing scheme; further, based on the technology of threshold adjustment and privacy requirement similarity, the device set is dynamically divided to facilitate subsequent privacy collaborative training of the sub-sets, thereby ensuring the stability of the grouping and accelerating the efficiency of subsequent layered privacy training; further, the layered key derivation and elastic encryption strategy are isolated through key rules at the training set level, which reduces the energy consumption of computing resources of the device data set while reducing the risk of arbitrary acquisition and tampering of device data. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.
[0016] Figure 1 A schematic diagram of the steps of an embodiment of the artificial intelligence-based device data privacy collaborative training method of the present invention; Figure 2 This is a schematic diagram of a portion of step S4 in the artificial intelligence-based device data privacy collaborative training method of the present invention; Figure 3 This is a schematic diagram of a portion of step S43 in the artificial intelligence-based device data privacy collaborative training method of the present invention; Figure 4 This is a schematic diagram of part of the steps in S2 of the artificial intelligence-based device data privacy collaborative training method of the present invention. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0018] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0019] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. In addition, the terms "first," "second," etc. are used only to distinguish the descriptions and are not to be understood as indicating or implying relative importance.
[0020] like Figure 1 As shown, a device data privacy collaborative training method based on artificial intelligence is provided, including: S1. Acquire multiple source devices, and obtain corresponding device security levels based on the source devices. Set a pending time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices. Acquire device operation data of the source devices at multiple time points within the pending time period, arrange the multiple device operation data in chronological order, and form a device data sequence corresponding to the source devices. S2. Obtaining the data sensitivity corresponding to the source device based on the device data sequence corresponding to the source device, and obtaining a fixed privacy requirement index corresponding to the source device based on the data sensitivity corresponding to the source device, and modifying the fixed privacy requirement index based on the device security level corresponding to the source device to form a target privacy requirement index corresponding to the source device; S3. Divide the multiple source devices into multiple training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training sets is less than a preset collaboration threshold; S4. Complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set in sequence.
[0021] In this embodiment, it should be noted that in S1, a dynamic mapping relationship between device security levels and data collection strategies is established. First, the device security level is predefined based on device identification (such as device type, communication security protocol, and historical security events). A higher security level indicates a greater potential risk to the data carried by the device. Generally, device security levels range from 1 to 10, with 1 being the lowest and 10 being the highest. Then, based on artificial intelligence (AI), device security levels are analyzed to dynamically set differentiated processing time periods. For high-security devices, the data collection window is extended to increase the number of samples, ensuring full-cycle characteristics of the device's operating status are captured. For low-security devices, the data collection window is shortened to reduce the number of samples, ensuring the integrity of basic data while reducing resource consumption. The intelligent setting of the processing time period essentially seeks the optimal balance between data comprehensiveness and collection cost. For example, a smart security camera (high security level) might be set to a 30-minute continuous monitoring period, while an air purifier (low security level) might only require 5-minute intervals for sampling, avoiding storage redundancy caused by excessive data collection from low-security devices.
[0022] Furthermore, during the device data sequence generation phase, a timestamped chain structure is employed for data organization. For example, a high-security medical testing device generates a time series consisting of high-frequency vital sign data that varies over time within a set two-minute acquisition window; whereas a low-security treatment room temperature sensor generates a data series consisting solely of temperature values that vary over time. Each data node's storage structure includes fields for device ID, acquisition time, value, and data checksum, providing a reliable foundation for time series features for subsequent sensitivity analysis. This hierarchical acquisition strategy optimizes edge computing resource allocation while ensuring the accuracy of privacy analysis.
[0023] In S2, a dynamic privacy needs assessment is constructed, achieving precise adaptation of protection strength through dual analysis of data characteristics and device attributes. First, data sensitivity is quantified based on the temporal characteristics of the device data sequence. By analyzing dimensions such as data fluctuation amplitude, change trends, and sudden outliers, a comprehensive assessment of the data's anti-tampering and anti-theft needs is made. For example, in a smart factory scenario, the real-time position data of a robotic arm is highly sensitive due to frequent mutations, while the on / off status data of warehouse lighting exhibits low volatility. By calculating the cumulative effect of data differences between adjacent time points and combining the degree of deviation of the latest data from the historical mean, the device types requiring key protection are dynamically identified. Data sensitivity is then mapped to a fixed privacy need index, employing a nonlinear function to ensure exponentially enhanced protection requirements for highly sensitive data. For example, real-time streaming data from medical imaging equipment, due to its bursty transmission characteristics, has a significantly higher privacy need index than the slowly changing ward environment monitoring data.
[0024] Furthermore, the device security level is introduced as a correction basis to achieve the joint optimization of security attributes and data characteristics. Through predefined device security levels, for example, the device security level of 10 for nuclear power plant controllers is higher than the device security level of 3 for peripheral meteorological sensors. This security level dynamically adjusts the fixed privacy requirement index through a certain proportional scaling mechanism, so that high-security devices can obtain a higher level of protection strategy even if the data sensitivity is the same. Taking autonomous driving as an example, the obstacle recognition data of the lidar and the volume data of the in-vehicle entertainment may have similar fluctuation characteristics. However, because the former involves driving safety, its target privacy requirement index will be significantly improved after security level weighting, driving the allocation of more complex encryption algorithms and more frequent key update strategies, thereby achieving security enhancement of key devices under resource constraints.
[0025] In S3, device clusters with homogeneous privacy requirements are constructed, and encryption policies are adapted through a dynamic grouping mechanism. Based on the target privacy requirement indicators of each device, adaptive clustering rules are used to group devices with similar indicator values into the same training set. For example, the difference between the target privacy requirement indicators corresponding to any two source devices in the training set must be less than a preset coordination threshold. The preset coordination threshold can be determined based on historical data or simulation training. For example, the standard deviation or quantile (such as the 95th percentile) of the target privacy requirement indicators of all devices can be calculated as a benchmark value. Unsupervised learning (such as K-means or DBSCAN) can also be used to cluster the target privacy requirement indicators and reversely calculate the preset coordination threshold based on the maximum distance or silhouette score within the cluster. For example, if the clustering results show that the most frequently occurring difference within the cluster is D, the preset coordination threshold can be set to 80% to 90% of D to balance the grouping granularity. Furthermore, in smart city traffic management scenarios, traffic light controllers and electronic police cameras, which have high privacy requirements and similar fluctuation ranges, are automatically grouped into a high-strength encryption cluster. Meanwhile, devices with lower privacy requirements, such as streetlight brightness sensors and green belt humidity monitors, are placed into a lightweight encryption cluster. This differentiated grouping mechanism ensures that devices within each cluster have the same privacy protection requirements, avoiding resource mismatches that can arise from traditional solutions where medical monitors and hospital ward curtain controllers are forced to use the same encryption policy.
[0026] Furthermore, sliding window threshold adjustment technology can be used to dynamically optimize cluster division accuracy based on real-time device status. When a CNC machine tool in the Industrial Internet of Things experiences a sudden increase in data sensitivity due to a sudden failure, its target privacy requirement indicator will exceed the coordination threshold of the original cluster, triggering a regrouping mechanism and migration to a cluster with a higher protection level. At the same time, to prevent cluster fragmentation, a weight decay function is introduced to smooth historical indicator data. For example, instantaneous data fluctuations of electric meter equipment in a smart grid during peak power consumption will not immediately cause cluster reorganization. Instead, time-weighted calculations are used to maintain grouping stability. This flexible grouping strategy ensures that security can quickly increase the encryption level of related equipment when an intrusion is identified, while maintaining the long-term stability of the factory production line equipment cluster, achieving an optimal balance between security and computing efficiency.
[0027] In S4, a dynamic, hierarchical privacy collaborative training mechanism is implemented to optimize security performance by precisely matching encryption strategies with device cluster characteristics. First, based on the mean privacy requirement indicators of each training set, device clusters are mapped to preset encryption strength levels. For example, in a smart grid scenario, the substation monitoring device cluster, due to its higher mean requirement indicator, is assigned an asymmetric encryption algorithm with a short-cycle key rotation strategy; while the meter reading cluster, due to its lower mean, uses lightweight symmetric encryption with long-cycle key updates. This hierarchical mechanism ensures that the dynamic video data of the traffic flow monitoring cluster in a smart city is protected by real-time encryption, while the static data of the street light status cluster only requires basic encryption, effectively balancing the security requirements of critical equipment with the energy efficiency constraints of ordinary equipment.
[0028] Furthermore, hierarchical key derivation technology is used in specific implementations to achieve cross-cluster isolation and protection. Each training set generates a unique key seed based on the encryption level. For example, the industrial robot cluster uses key derivation based on a physically unclonable function (PUF), while the warehouse temperature control cluster uses a traditional hash chain key. When the smart home security cluster detects an intrusion, it dynamically upgrades the encryption level to AES-256 and enables two-factor authentication, while maintaining AES-128 encryption for the door and window sensor cluster. This elastic mechanism enables the emergency braking data stream in autonomous driving to instantly switch to a quantum-safe encryption protocol, while the in-vehicle infotainment data maintains standard encryption, achieving millisecond-level enhanced protection for critical data channels in the event of sudden security threats, while avoiding resource overload caused by full encryption upgrades.
[0029] To sum up, in the entire AI-based device data privacy collaborative training method, first, the dynamic data collection window division based on device security level solves the resource mismatch problem of static time windows in existing solutions. Through differentiated collection strategies, high-security devices can obtain more comprehensive time series feature support, while avoiding redundant data collection of low-security devices; further, the joint correction mechanism of data sensitivity quantification and device security level breaks through the limitations of single-dimensional evaluation. By integrating data volatility characteristics and device inherent security attributes, the device can automatically trigger enhanced privacy protection during burst data transmission, which greatly improves the protection strength of key data compared with existing solutions; further, based on threshold adjustment and privacy requirement similarity technology, the device set is dynamically divided to facilitate subsequent privacy collaborative training of the sub-sets, thereby ensuring group stability and accelerating the efficiency of subsequent layered privacy training; further, the layered key derivation and elastic encryption strategy are isolated through training set-level key rules, which reduces the risk of arbitrary acquisition and tampering of device data while reducing the computing resource energy consumption of the device data set.
[0030] like Figure 2 As shown, in one embodiment, completing the privacy collaborative training between the device data sequences corresponding to the multiple source devices in each training set in sequence in S4 includes: S41. Obtaining the mean of the privacy requirement indicators corresponding to each training set based on the target privacy requirement indicators corresponding to multiple source devices in each training set; S42. Acquire multiple different value ranges, where different value ranges correspond to different encryption levels; S43. Obtain the encryption level corresponding to each training set according to the value range that each training set falls into, and complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set.
[0031] In this embodiment, it should be noted that in S41, a cluster-level security benchmark is established by aggregating the privacy requirement indicators of the devices in the training set. The target privacy requirement indicators of all devices in each training set are averaged, and weighted processing can also be performed during the calculation. The weights can be dynamically adjusted according to the amount of device data or the security level. For example, in a smart medical scenario, the training set containing cardiac monitors and ventilators will have a significantly higher mean index than a cluster containing only ward lighting equipment, because both types of equipment involve patient vital signs. This mean calculation can reflect the overall security situation of the cluster and provide a quantitative basis for the subsequent allocation of encryption levels. In the industrial Internet of Things, if a training set contains CNC machine tools and quality inspection cameras, the proportion of high-security level data of machine tools will be given priority, so that the mean is more biased towards high-intensity encryption requirements.
[0032] In S42, multiple levels of encryption strength are predefined, each corresponding to a specific algorithm combination and key management strategy. The interval division adopts a non-uniform design, with more fine-grained levels set in areas with high privacy requirements. For example, the privacy requirement indicator is divided into three main intervals: basic level (lightweight symmetric encryption), enhanced level (hybrid encryption), and critical level (asymmetric encryption + quantum security protocol). The enhanced level is further subdivided into two sub-intervals to adapt to different real-time requirements. In an autonomous driving scenario, the vehicle control cluster may be mapped to the critical level interval, triggering real-time data protection based on elliptic curve encryption; while the passenger entertainment cluster falls into the basic level and uses AES-128 encryption to save on-board computing resources. This asymmetric interval design ensures more precise protection adaptation for high-security devices.
[0033] In S43, the corresponding privacy collaborative training pipeline is started according to the encryption level interval that the training set falls into. For the high-intensity encryption interval, a phased encryption mechanism is adopted: the device data sequence is first time-sliced, real-time encryption is enabled for the burst-sensitive data segment, and batch processing is delayed for the stable data segment. For example, when the smart home security cluster detects abnormal movement, it immediately applies AES-256 encryption to the current video stream slice, while the historical normal segments are encrypted with lightweight encryption. At the same time, a heterogeneous strategy is adopted for cross-cluster key management: high-security clusters use key derivation based on physically unclonable functions (PUFs), and medium and low-security clusters use key distribution based on hash trees. In industrial control scenarios, when a training set jumps to a higher encryption interval due to a change in device status, the encryption acceleration module is automatically injected, and the key cache of the edge node is synchronously updated to achieve a seamless upgrade of encryption strength.
[0034] like Figure 3 As shown, in one embodiment, completing privacy collaborative training between device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set in S43 includes: S431. Encrypt the device data sequence in the training set segment by segment using encryption logic corresponding to the encryption level, where the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; S432. Generate and distribute encryption keys based on the encryption level, wherein the device data sequences in each training set share a key generation rule.
[0035] In this embodiment, it should be noted that in S431, a dynamic sharding encryption mechanism is used to achieve precise data granularity protection. Based on the logic corresponding to the encryption level, device data sequences are time-sliced and analyzed to identify highly sensitive data segments (such as sudden outliers and key operation nodes). High-strength encryption is prioritized for these segments, while basic encryption is used for stable data segments. For example, in an industrial equipment monitoring scenario, when a pressure sensor detects a sudden spike in data, asymmetric encryption is immediately applied to the data stream within that time period, while lightweight symmetric encryption is used for normal data. For non-continuous sensitive data (such as motion detection segments from smart home security cameras), multiple encryption algorithms can be mixed within a single data stream, ensuring quantum-level protection for key event frames while avoiding the waste of high-strength encryption throughout the entire process. Furthermore, the dynamic selection mechanism for encryption logic adapts to the varying computing power of edge devices. For example, low-power sensors may only use AES-128's fast encryption mode, while industrial control equipment equipped with security chips may run the nationally recognized SM4 algorithm.
[0036] In S432, this step uses hierarchical key derivation technology to achieve cross-cluster security isolation and efficiency optimization. Each training set generates a unique key seed based on the encryption level. High-security clusters combine physically unclonable features (such as device fingerprints) to generate quantum-resistant keys, while low-security clusters use lightweight hash chains to derive short-term keys. For example, in the Internet of Vehicles scenario, the autonomous driving control cluster uses a dynamic key based on a chaotic map, updated every 100ms; while the in-vehicle entertainment cluster uses a session key derived from a pre-shared master key, rotated every 10 minutes. At the same time, through a key generation rule sharing mechanism, devices within the same training set can quickly synchronize key status during encrypted communication. For example, a cluster of robots on the same production line in a smart factory achieves millisecond-level key synchronization through a key broadcast tree, reducing communication latency by 75% compared to the existing unicast mode. This design not only prevents the chain reaction risk of cross-cluster key leakage, but also improves the collaborative efficiency of devices in the same cluster through rule standardization.
[0037] In one embodiment, the data sensitivity corresponding to the source device is obtained according to the device data sequence corresponding to the source device in S2 as follows: ;in, is the data sensitivity corresponding to the j-th source device, is the adjustment coefficient, is the jth device data sequence Equipment operation data at a point in time, is the device operation data at the first time point in the j-th device data sequence, is the number of time points in the j-th device data sequence, is the device operation data at the i+1th time point in the jth device data sequence, is the device operation data at the i-th time point in the j-th device data sequence, is the standard maximum operating data of the j-th source device, is the standard minimum operating data of the j-th source device, is a conditional parameter.
[0038] In this embodiment, it should be noted that It is a bursty item (deviation from the mean item) used to capture the degree of deviation of the latest data point from the historical mean and identify sudden anomalies (such as a sudden increase in heart rate detected by a medical device). It can solve the problem that static time windows cannot detect dynamic sensitivity changes. The larger the burst value, the more likely the current data contains critical events, and the encryption strength needs to be improved.
[0039] Further, It is a volatility term (adjacent difference term) used to calculate the overall fluctuation amplitude of the data series, reflecting the stability of the device data. It can distinguish between key equipment with high-frequency changes (such as industrial robotic arms) and ordinary equipment with smooth changes (such as temperature and humidity sensors), avoiding the waste of resources caused by unified encryption across the entire domain.
[0040] Further, Normalize the data series to eliminate the impact of the dimension of the device data collected by different devices on the sensitivity, and prevent the problem of the inconsistency of the dimensions of different data types leading to the inability to guarantee the validity of subsequent data indicators; among them, the standard maximum operating data Refers to the upper limit of the normal operating state range, standard minimum operating data Refers to the lower limit of the normal operating state range. In most cases Greater than ;in, is a conditional parameter. When the condition is When it is zero, is not zero, and is generally a value of the same dimension as the numerator, which ensures that the denominator is not zero and the validity of the calculation result; when the condition is When it is not zero, Directly equal to 0 to ensure the validity of the calculation results.
[0041] Furthermore, an adjustment coefficient greater than or equal to 1 is set , used in When it is 0, it ensures that the volatility item can still affect the sensitivity, avoiding the situation where the sensitivity of the equipment with severe fluctuation is zero due to the sudden item being zero, and can still improve its sensitivity according to the volatility; at the same time, the adjustment coefficient Can be preset according to industry standards or equipment safety levels; =1, ensuring that the suddenness term dominates the sensitivity; =2, the overall sensitivity is improved, and at the same time, the dependence on volatility quantification is increased, which indirectly weakens the contribution of sudden detection and avoids abnormal false triggering of high-intensity encryption.
[0042] For example, the jth source device is a medical monitor, and the heart rate data sequence collected within 5 seconds is [70, 72, 75, 80, 85] (unit: beats / minute). The adjustment coefficient =2, =100, =60. Substitute into the expression for calculation, data sensitivity =0.99. In summary, the suddenness term (8.6) shows that the latest heart rate is significantly higher than the mean, and the volatility term (3.75) represents the fluctuation intensity of the device data series; the sensitivity is finally obtained. =0.99 is relatively high, thus making a greater contribution to the privacy requirement index, facilitating high-intensity privacy collaborative training and preventing vital sign data from being tampered with.
[0043] For example, the jth source device is a temperature sensor, and the temperature data sequence collected within 1 minute is [27, 28.6, 28.4, 28]. The adjustment coefficient =1, =35, =20. Substitute into the expression for calculation, data sensitivity =0.049. In summary, the suddenness term (0) shows that the latest temperature is not higher than the mean, and the volatility term (0.73) represents the fluctuation intensity of the device data series; the final sensitivity is =0.049 is too low, and the subsequent privacy requirement index is small, so there is no need for high-intensity privacy collaborative training.
[0044] In one embodiment, in S2, the fixed privacy requirement index corresponding to the source device is obtained according to the data sensitivity corresponding to the source device, which is expressed as: ;in, is the fixed privacy requirement indicator corresponding to the j-th source device, is the data sensitivity corresponding to the j-th source device.
[0045] In this embodiment, it should be noted that To achieve nonlinear growth suppression, when The growth rate slows down when the sensitivity increases; prevents the privacy requirements of highly sensitive devices from growing too fast and avoids over-allocation of encryption resources; ensures that the privacy requirements of low-sensitivity devices can still be reasonably improved to ensure basic protection. The linear value range (such as 0~2) is compressed to a logarithmic scale (such as 0~1.1) to facilitate the unified division of encryption level intervals.
[0046] Example: If the encryption level is Divided into: low (between 0 and 0.3), medium (between 0.3 and 0.6), high (above 0.6), then Reaching 0.82 triggers high-intensity data privacy training, Reaching 0.35 and below 0.82 triggers medium-intensity data privacy training. A value below 0.35 triggers low-intensity data privacy training to avoid misjudgment due to short-term fluctuations.
[0047] like Figure 4 As shown, in one embodiment, in S2, modifying the fixed privacy requirement indicator according to the device security level corresponding to the source device and forming the target privacy requirement indicator corresponding to the source device includes: S21. Obtaining a scaling factor corresponding to the source device according to the device security level corresponding to the source device; S22. Modify the fixed privacy requirement indicator according to the scaling factor and form a target privacy requirement indicator corresponding to the source device.
[0048] In this embodiment, it should be noted that in S21, the device security level is quantified into a scaling factor through linear proportional conversion to achieve quantitative fusion of security attributes. The current level is normalized according to the maximum value of the device security level to ensure that devices at different security levels receive differentiated weight adjustments. For example, in a system with a security level range of 1 to 10, the scaling factor of a nuclear power plant controller (security level 10) is 2.0, while the coefficient of a meteorological sensor (security level 3) is 1.3. This design ensures that even if the data sensitivity of a high-security device is the same as that of a low-security device, its privacy requirement index can still be significantly improved through coefficient weighting, thereby driving the allocation of more complex encryption algorithms and more frequent key update strategies, fundamentally avoiding the problem of resource waste caused by the same security strategy for industrial control equipment and ordinary sensors.
[0049] In S22, this step dynamically enhances the basic privacy requirements through a scaling factor, achieving coordinated optimization of device security attributes and data characteristics. By multiplying the fixed privacy requirement index by the scaling factor, the privacy requirements of high-security devices are linearly amplified, while those of low-security devices are only slightly increased. For example, if the fixed privacy requirement index of both medical imaging equipment and ward temperature and humidity sensors is 0.4, the former's index rises to 0.72 after correction due to its high security level (scaling factor 1.8), triggering high-intensity data privacy training; the latter's index is 0.48 after correction due to its low security level (scaling factor 1.2), requiring only medium-intensity data privacy training. This mechanism effectively solves the problem of vital sign data and environmental data being forced to use the same encryption strength in smart healthcare scenarios, preventing both insufficient protection of critical data and overloading the computing power of edge devices.
[0050] In one embodiment, in S21, the scaling factor corresponding to the source device is obtained according to the device security level corresponding to the source device, which is expressed as: ;in, is the scaling factor corresponding to the j-th source device, is the device security level corresponding to the j-th source device, This is the maximum equipment safety level.
[0051] In this embodiment, it should be noted that the “1” in the expression is used as a baseline to ensure that the scaling factor of the lowest security level device is 1, avoiding zero contribution or negative adjustment. This achieves a linear mapping of security levels, with higher-level devices receiving larger coefficient increases. For example, a device with security level 10 has a coefficient of 2.0, while a device with security level 3 has a coefficient of 1.3. Furthermore, security levels of varying ranges (e.g., 1-10 and 1-100) are uniformly mapped to the interval [1, 2] to ensure cross-system compatibility. For example, if a system's security level ranges from 1 to 5, a device with security level 5 has a coefficient of 2, consistent with the coefficient of a device with security level 10 in the 1-10 system, facilitating policy standardization. Furthermore, security attributes are decoupled from data characteristics. The security level acts as a correction factor independent of data sensitivity, preventing static attributes of high-security devices (such as device type) from being obscured by dynamic data fluctuations. This addresses the traditional problem of industrial control equipment and sensors being forced to adopt the same encryption strategy due to data similarity, and effectively increases the protection level of critical equipment through security level enforcement.
[0052] For example, equipment 1 (medical CT machine, safety level =8, maximum level =10); Device 2 (Ward temperature and humidity sensor, safety level =2, maximum level =10). Fixed privacy requirement index: The two devices are calculated by S2 =0.4, =0.2.
[0053] Substitute this into the expression to calculate the scaling factor of device 1. ; Scaling factor for device 2 ; The revised target privacy requirement indicators are , .
[0054] In summary, if encryption levels are categorized as follows: Low (between 0 and 0.3) for AES128, with daily key updates; Medium (between 0.3 and 0.6) for AES256, with hourly key updates; and High (above 0.6) for RSA3072, with updates every 10 minutes, for Device 1, a key of 0.4 before the correction triggered medium-intensity data privacy training, indicating a high security level. After the correction, the change was significant, while a key of 0.72 triggered high-intensity data privacy training. For Device 2, a key of 0.2 before the correction triggered low-intensity data privacy training, indicating a low security level. After the correction, the change was not significant, while a key of 0.24 still triggered low-intensity data privacy training.
[0055] A device data privacy collaborative training system based on artificial intelligence is also provided, which includes: A data acquisition module is used to acquire multiple source devices, obtain corresponding device security levels based on the source devices, set a pending time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices, acquire device operation data of the source devices at multiple time points within the pending time period, arrange the multiple device operation data in chronological order, and form a device data sequence corresponding to the source devices; a data processing module, configured to obtain a data sensitivity corresponding to the source device based on a device data sequence corresponding to the source device, obtain a fixed privacy requirement indicator corresponding to the source device based on the data sensitivity corresponding to the source device, and modify the fixed privacy requirement indicator based on a device security level corresponding to the source device to form a target privacy requirement indicator corresponding to the source device; a data partitioning module, configured to partition a plurality of source devices into a plurality of training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training set is less than a preset collaboration threshold; The collaborative training module is used to sequentially complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set.
[0056] In one embodiment, the collaborative training module is also used to: obtain the mean privacy requirement indicator corresponding to each training set based on the target privacy requirement indicator corresponding to multiple source devices in each training set; obtain multiple different value ranges, where different value ranges correspond to different encryption levels; obtain the encryption level corresponding to each training set in turn according to the value range that each training set falls into, and complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set.
[0057] In one embodiment, the collaborative training module is also used to: encrypt the device data sequence in the training set segment by segment using the encryption logic corresponding to the encryption level, wherein the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; generate and distribute encryption keys based on the encryption level, wherein the device data sequences in each training set share key generation rules.
[0058] In this embodiment, it should be noted that, regarding the above-mentioned artificial intelligence-based device data privacy collaborative training system, the specific method of performing operations has been described in detail in the implementation of the artificial intelligence-based device data privacy collaborative training method, and will not be elaborated here.
[0059] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.
[0060] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0061] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.
Claims
1. A device data privacy collaborative training method based on artificial intelligence, characterized in that: include: Acquire multiple source devices, and obtain corresponding device security levels based on the source devices, set a waiting time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices, and obtain device operation data of the source devices at multiple time points within the waiting time period, arrange the multiple device operation data in chronological order and form a device data sequence corresponding to the source devices; Obtaining the data sensitivity corresponding to the source device based on the device data sequence corresponding to the source device, and obtaining the fixed privacy requirement index corresponding to the source device based on the data sensitivity corresponding to the source device, and modifying the fixed privacy requirement index based on the device security level corresponding to the source device to form a target privacy requirement index corresponding to the source device; Dividing the plurality of source devices into a plurality of training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training set is less than a preset collaboration threshold; Complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set in sequence.
2. The device data privacy collaborative training method based on artificial intelligence according to claim 1 is characterized in that: The step of sequentially completing privacy collaborative training between device data sequences corresponding to multiple source devices in each training set includes: Obtain the mean of the privacy requirement indicators corresponding to each training set based on the target privacy requirement indicators corresponding to multiple source devices in each training set; Obtain multiple different value ranges, where different value ranges correspond to different encryption levels; The encryption level corresponding to each training set is obtained in turn according to the value range that each training set falls into, and the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set is completed according to the encryption level corresponding to each training set.
3. The device data privacy collaborative training method based on artificial intelligence according to claim 2 is characterized in that: The completing privacy collaborative training between device data sequences corresponding to multiple source devices in each training set according to the encryption level corresponding to each training set includes: Encrypting the device data sequence in the training set segment by segment using encryption logic corresponding to the encryption level, wherein the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; Encryption keys are generated and distributed based on the encryption level, wherein the device data sequences in each training set share a key generation rule.
4. The device data privacy collaborative training method based on artificial intelligence according to claim 1 is characterized in that: The data sensitivity corresponding to the source device is obtained according to the device data sequence corresponding to the source device as follows: ;in, is the data sensitivity corresponding to the j-th source device, is the adjustment coefficient, is the jth device data sequence Equipment operation data at a point in time, is the device operation data at the first time point in the j-th device data sequence, is the number of time points in the j-th device data sequence, is the device operation data at the i+1th time point in the jth device data sequence, is the device operation data at the i-th time point in the j-th device data sequence, is the standard maximum operating data of the j-th source device, is the standard minimum operating data of the j-th source device, is a conditional parameter.
5. The device data privacy collaborative training method based on artificial intelligence according to claim 1 is characterized in that: The fixed privacy requirement index corresponding to the source device is obtained according to the data sensitivity corresponding to the source device as follows: ;in, is the fixed privacy requirement indicator corresponding to the j-th source device, is the data sensitivity corresponding to the j-th source device.
6. The device data privacy collaborative training method based on artificial intelligence according to claim 1 is characterized in that: The step of modifying the fixed privacy requirement indicator according to the device security level corresponding to the source device and forming the target privacy requirement indicator corresponding to the source device includes: Obtaining a scaling factor corresponding to the source device according to a device security level corresponding to the source device; The fixed privacy requirement index is modified according to the scaling factor to form the target privacy requirement index corresponding to the source device.
7. The device data privacy collaborative training method based on artificial intelligence according to claim 6 is characterized in that: The scaling factor corresponding to the source device is obtained according to the device security level corresponding to the source device as follows: ;in, is the scaling factor corresponding to the j-th source device, is the device security level corresponding to the j-th source device, This is the maximum equipment safety level.
8. An artificial intelligence-based device data privacy collaborative training system, characterized in that: The system comprises: A data acquisition module is used to acquire multiple source devices, obtain corresponding device security levels based on the source devices, set a pending time period corresponding to the source devices based on artificial intelligence and the device security levels corresponding to the source devices, acquire device operation data of the source devices at multiple time points within the pending time period, arrange the multiple device operation data in chronological order, and form a device data sequence corresponding to the source devices; a data processing module, configured to obtain a data sensitivity corresponding to the source device based on a device data sequence corresponding to the source device, obtain a fixed privacy requirement indicator corresponding to the source device based on the data sensitivity corresponding to the source device, and modify the fixed privacy requirement indicator based on a device security level corresponding to the source device to form a target privacy requirement indicator corresponding to the source device; a data partitioning module, configured to partition a plurality of source devices into a plurality of training sets, wherein the difference between the target privacy requirement indicators corresponding to any two source devices in the training set is less than a preset collaboration threshold; The collaborative training module is used to sequentially complete the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set.
9. The artificial intelligence-based device data privacy collaborative training system according to claim 8, characterized in that: The collaborative training module is also used to: Obtain the mean of the privacy requirement indicators corresponding to each training set based on the target privacy requirement indicators corresponding to multiple source devices in each training set; Obtain multiple different value ranges, where different value ranges correspond to different encryption levels; The encryption level corresponding to each training set is obtained in turn according to the value range that each training set falls into, and the privacy collaborative training between the device data sequences corresponding to multiple source devices in each training set is completed according to the encryption level corresponding to each training set.
10. The artificial intelligence-based device data privacy collaborative training system according to claim 9, characterized in that: The collaborative training module is also used to: Encrypting the device data sequence in the training set segment by segment using encryption logic corresponding to the encryption level, wherein the encryption logic includes at least one of a symmetric encryption algorithm or an asymmetric encryption algorithm; Encryption keys are generated and distributed based on the encryption level, wherein the device data sequences in each training set share a key generation rule.