Internet of Things equipment behavior monitoring method, computer device, medium and product
By distribute the training of discrete denoising and diffusion model on local nodes of IoT devices, combined with activation value deviation and distribution characteristics, the problems of small data coverage and poor adaptability in IoT device behavior monitoring are solved, and efficient and flexible device behavior monitoring is achieved.
Patent Information
- Application Number
- CN202510313930.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing IoT device behavior monitoring methods have problems such as small data coverage, poor adaptability and poor flexibility in diverse scenarios, and it is difficult to effectively detect unknown attacks or abnormal behaviors.
A distributed training strategy is adopted to train discrete denoising diffusion models on the nodes of each IoT local device, extract the activation value to determine the trusted nodes, calculate the distribution characteristics and similarity based on the local data of the trusted nodes, configure the weight coefficients, perform weighted average fusion of model parameters, generate global model parameters, and perform iterative optimization.
It realizes a wider coverage of equipment behavior patterns, enhances adaptability to diverse scenarios, improves the flexibility and accuracy of the monitoring system, can adaptively adjust model parameters, reduces the need for manual adjustment of rules, and reduces the system maintenance cost.
Smart Images

Figure CN120223376A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of intelligent monitoring, and particularly to a method for monitoring the behavior of Internet of Things devices, a computer device, a medium, and a product. Background Art
[0002] The rapid development of Internet of Things technology has made it possible for a large number of devices to be interconnected, but at the same time, it has also brought challenges to the security and credibility of device behavior. Device behavior monitoring and authentication in the Internet of Things environment are key technologies to ensure the safe and stable operation of the system.
[0003] Currently, the common methods for monitoring the behavior of Internet of Things devices mainly include two types: rule-based abnormal behavior detection and using a centralized deep learning method to model and authenticate device behavior. Among them, the rule-based abnormal detection method usually identifies abnormal behavior by presetting thresholds and rules. Due to the complex and changeable Internet of Things environment, new device types and application scenarios are constantly emerging, and the publicly disclosed preset rules are difficult to cover all possible abnormal situation rules. When the Internet of Things system is upgraded or updated, the behavior pattern of the device may change. At this time, manual adjustment of the rules is required. If the adjustment is not timely, it may lead to normal behavior being misjudged as abnormal, or abnormal behavior not being detected. At the same time, rule-based abnormal detection can only detect those predefined abnormal patterns and is often powerless against unknown attacks or abnormal behaviors.
[0004] For the method of using a centralized deep learning method to model and authenticate device behavior, first, the behavior data of Internet of Things devices needs to be collected, then deep neural networks are used for feature extraction and behavior pattern learning, and finally, model reasoning is used to judge the legality of device behavior. However, this method has the following technical defects: 1) Centralized learning requires uploading all device data to the central server, which poses a risk of data privacy leakage. At the same time, there is also a problem of insufficient data diversity and inability to comprehensively cover all device behavior patterns; 2) The characteristics of uneven data distribution in the Internet of Things environment are not considered, and the model performance is easily affected; 3) It is vulnerable to Byzantine attacks, and attackers may contaminate the model training process by injecting malicious samples. Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide a method for monitoring the behavior of Internet of Things devices, a computer device, a medium, and a product, which can solve the problems of small data coverage, poor adaptability to diverse scenarios, and poor flexibility existing in the publicly disclosed monitoring methods in the prior art.
[0006] In a first aspect, the embodiments of the present disclosure provide a method for monitoring the behavior of Internet of Things devices, including:
[0007] According to the distribution of Internet of Things (IoT) devices, train a discrete denoising diffusion model at each node of the IoT local device, and denote the trained discrete denoising diffusion model as the pre-trained local model corresponding to the node;
[0008] Extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviation of each node;
[0009] Based on the local data of the trusted nodes, calculate the distribution characteristics of each node and the distribution similarity of the data between nodes;
[0010] Based on the distribution characteristics and the distribution similarity, determine the data heterogeneity evaluation index, and configure a weight coefficient for each node according to the data heterogeneity evaluation index;
[0011] Perform weighted average fusion on the model parameters of the pre-trained local model corresponding to each node uploaded by each node and the weight coefficient of each node to generate global model parameters;
[0012] Send the global model parameters as initial parameters to each node, and perform iterative optimization of the corresponding pre-trained local model at each node, and denote the pre-trained local model that meets the iteration condition as the target model;
[0013] Deploy all the target models to the monitoring nodes of the corresponding local devices to generate an integrated monitoring system;
[0014] Analyze the real-time behavior data of the target device collected based on the integrated monitoring system to generate a device behavior authentication result.
[0015] Optionally, the step of training a discrete denoising diffusion model at each node of the IoT local device according to the distribution of IoT devices and denoting the trained discrete denoising diffusion model as the pre-trained local model corresponding to the node includes:
[0016] According to the distribution of IoT devices, configure the Director node and several Envoy nodes through the OpenFL framework and establish a secure communication channel to obtain the constructed federated learning network topology;
[0017] Based on the federated learning network topology, collect the corresponding device operation data at each Envoy node, and after preprocessing the device operation data, obtain a standardized behavior data set;
[0018] Configure the architecture of the discrete denoising diffusion model based on the behavior data set, and denote the discrete denoising diffusion model with the configured architecture as the first model; the architecture includes an encoder, a noise predictor, and a decoder;
[0019] Perform local training of forward diffusion and reverse denoising on the first model corresponding to the Envoy node, and denote the trained first model as the pre-trained local model corresponding to the Envoy node.
[0020] Optionally, extract the activation values of each node during the training of the discrete denoising diffusion model, and obtain all trusted nodes according to the calculated activation value deviations of each node, including:
[0021] Extract the activation values of neurons in each layer during the training of the discrete denoising diffusion model, and generate an activation value statistical feature set;
[0022] Based on the activation value statistical feature set, calculate the activation value distribution parameters of normal nodes, and construct a standard activation mapping pattern;
[0023] Use the standard activation mapping pattern to calculate the activation value deviation of each node and perform anomaly detection to generate a list of Byzantine nodes;
[0024] Obtain the filtered trusted nodes according to the list of Byzantine nodes and the filtering rules.
[0025] Optionally, based on the pre-trained local model, extract the activation values of neurons in each layer during the training process to generate an activation value statistical feature set, including:
[0026] During the training of the discrete denoising diffusion model, add observation points in each layer of the neural network to collect the activation values of each neuron in real time;
[0027] Collect the activation value data of each neuron by layer to form a multi-dimensional data set, and each data point in the data set represents the activation state of a neuron at a specific training step;
[0028] Perform statistical analysis on the collected activation value data, calculate the activation value statistical features, and the activation value statistical features include activation data mean, activation data variance, activation data skewness, and activation data kurtosis;
[0029] Classify the activation value statistical features by layer and neuron organization to generate a multi-dimensional activation value statistical feature set.
[0030] Optionally, use the standard activation mapping pattern to calculate the activation value deviation of each node and perform anomaly detection to generate a list of Byzantine nodes, including:
[0031] Calculate the difference between the activation value distribution of each node and the standard activation mapping pattern;
[0032] Obtain the degree of shape deviation of the activation value distribution and its change trend over time series based on the differences, and mark the corresponding nodes that meet any of the conditions that the degree of shape deviation exceeds a preset threshold and the change trend reaches a preset condition as potential Byzantine nodes;
[0033] Divide the entire time series into Q time windows;
[0034] Calculate the degree of shape deviation of the activation value distribution of each node within each of the time windows and its change trend over time series. If a node is marked as a potential Byzantine node P times within Q time windows, then mark the node as a target Byzantine node;
[0035] Q≥2;
[0036] Summarize the IDs of the nodes marked as the target Byzantine nodes to generate a list of Byzantine nodes.
[0037] Optionally, calculating the distribution characteristics of each node and the distribution similarity of data between nodes based on the local data of the trusted nodes includes:
[0038] Calculate the distribution characteristics of each trusted node according to the obtained local data of each trusted node, where the distribution characteristics include statistical characteristics and distribution shape characteristics;
[0039] Calculate the distribution similarity of data between nodes using a statistical distance metric.
[0040] Optionally, analyzing the real-time behavior data of the target device collected by the comprehensive monitoring system to generate a device behavior authentication result includes:
[0041] Continuously collect the real-time behavior data of the target device through the comprehensive monitoring system to form a behavior sequence sorted by time;
[0042] Analyze the behavior sequence based on the target model to generate a behavior evaluation result;
[0043] Output a device behavior authentication result according to the behavior evaluation result and a preset anomaly threshold.
[0044] In a second aspect, an embodiment of the present disclosure also provides a computer device, adopting the following technical solution:
[0045] The computer device includes:
[0046] At least one processor; and,
[0047] A memory communicatively connected to the at least one processor; wherein,
[0048] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the Internet of Things device behavior monitoring method described in any one of the above.
[0049] In a third aspect, an embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions for causing a computer to execute the Internet of Things device behavior monitoring method described in any one of the above.
[0050] In a fourth aspect, an embodiment of the present disclosure further provides a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in any one of the above are implemented.
[0051] For the Internet of Things device behavior monitoring method disclosed in this application, according to the distribution of Internet of Things devices, a discrete denoising diffusion model is trained at each node of each Internet of Things local device, and the trained discrete denoising diffusion model is denoted as the pre-trained local model of the corresponding node. This distributed training strategy can better adapt to the heterogeneity of various devices; extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviations of each node; calculate the distribution characteristics of each node and the distribution similarity of data between nodes based on the local data of the trusted nodes; determine the data heterogeneity evaluation index based on the distribution characteristics and distribution similarity, and configure a weight coefficient for each node according to the data heterogeneity evaluation index; perform weighted average fusion on the model parameters of the corresponding pre-trained local model uploaded by each node and the weight coefficient of each node to generate global model parameters; send the global model parameters as initial parameters to each node, and perform iterative optimization of the corresponding pre-trained local model at each node, and denote the pre-trained local model that meets the iteration condition as the target model; deploy all the target models to the monitoring nodes of the corresponding local devices to generate a comprehensive monitoring system; analyze the real-time behavior data of the target device collected based on the comprehensive monitoring system to generate a device behavior authentication result, make full use of the local data of distributed devices, avoid the data limitation of a single monitoring center, can effectively integrate the data of different devices, expand the data coverage of the monitoring system, can dynamically adjust the model parameters according to the actual operation situation of the device, improve the flexibility of the monitoring system, can generate a customized monitoring system for different scenarios, enhance the adaptability to diverse scenarios, and at the same time can accurately and quickly output the monitoring result, with strong self-adaptability, does not require manual frequent adjustment of rules to adapt to the upgrade and change of the system, and effectively reduces the maintenance cost of the system.
[0052] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. In order to make the above and other objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, is described in detail as follows. Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic flowchart of the method for monitoring the behavior of Internet of Things devices provided by the embodiments of the present disclosure.
[0055] Figure 2 For Figure 1 It is a schematic flowchart of the method for obtaining the pre-trained local model in
[0056] Figure 3 For Figure 1 It is a schematic flowchart of the method for obtaining all trusted nodes in
[0057] Figure 4 For Figure 3 It is a schematic flowchart of the method for generating the activation value statistical feature set in
[0058] Figure 5 For Figure 3 It is a schematic flowchart of the method for constructing the standard activation mapping mode in
[0059] Figure 6 For Figure 3 It is a schematic flowchart of the method for generating the list of Byzantine nodes in
[0060] Figure 7 For Figure 3 It is a schematic flowchart of the method for filtering trusted nodes in
[0061] Figure 8 It is a schematic flowchart of the method for calculating the distribution characteristics of each node and the distribution similarity of data between nodes based on trusted nodes provided by the embodiments of the present disclosure.
[0062] Figure 9 For Figure 1 It is a schematic flowchart of the method for generating the device behavior authentication result in
[0063] Figure 10 It is a schematic structural diagram of a computer device provided by the embodiments of the present disclosure. Detailed implementation manners
[0064] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0065] It should be clear that the embodiments of the present disclosure are described below through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0066] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0067] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0068] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0069] Referring to Figure 1 , this application discloses an Internet of Things device behavior monitoring method, including:
[0070] S100. According to the distribution of IoT devices, train a discrete denoising diffusion model at each node of the IoT local device, and record the trained discrete denoising diffusion model as the pre-trained local model corresponding to the node.
[0071] Suppose there is an IoT system in a smart factory, which has multiple different types of production devices, such as machine tools, robots, etc. Each device can be regarded as a local node. Collect the historical behavior data of each device in the normal operating state, such as the rotation speed of the machine tool, the feed speed of the tool, etc., and the movement trajectory and joint angles of the robot.
[0072] Build a discrete denoising diffusion model for each node separately, and use the historical data of the node to train the model. During the training process, the model will learn the distribution characteristics of the data, such as the rotation speed change law of the machine tool under different production tasks.
[0073] In this step, the local data of each node is fully utilized, and the model can be built according to the unique data characteristics of each node, improving the adaptability and accuracy of the model to local data; training at the local node effectively reduces the pressure and cost of data transmission, while protecting the privacy of data, and realizes the distributed modeling of device behavior data.
[0074] S200. Extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviations of each node.
[0075] By analyzing the activation value distribution during the model training process, according to the calculated activation value deviations of each node, it is possible to identify nodes that may be abnormal or incorrect, eliminate the interference of untrusted nodes on subsequent model fusion and analysis, and improve the reliability of the overall system. At the same time, through the analysis of the activation value deviations, the differences between nodes can be found, which helps to further understand the operating state of the nodes.
[0076] In step S200, extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviations of each node. Only the model parameters of trusted nodes will participate in the subsequent fusion process, which can effectively resist Byzantine attacks and prevent attackers from contaminating the model training process by injecting malicious samples.
[0077] S300. Based on the local data of trusted nodes, calculate the distribution characteristics of each node and the distribution similarity of data between nodes.
[0078] Understanding the distribution characteristics of data at each node helps to deeply understand the operation rules of the nodes and provide a basis for subsequent data processing and model optimization; calculating the distribution similarity of data between nodes can discover the correlation relationships between nodes and provide a reference for model fusion and data heterogeneity evaluation.
[0079] S400. Based on the distribution characteristics and distribution similarity, determine the data heterogeneity evaluation index, and configure a weight coefficient for each node according to the data heterogeneity evaluation index.
[0080] In the scenario of federated learning or distributed machine learning, data heterogeneity (i.e., the data distribution difference between different nodes) is an important issue. To evaluate data heterogeneity, in this step, based on the local data of trusted nodes, calculate the distribution characteristics of each node and the distribution similarity of data between nodes, and finally determine the data heterogeneity evaluation index.
[0081] According to the data heterogeneity evaluation index, configure a weight coefficient for each node. Nodes with lower data heterogeneity can be assigned higher weights because the data of these nodes is more representative; while nodes with higher data heterogeneity are assigned lower weights.
[0082] Through the configuration of the data heterogeneity evaluation index and weight coefficients, it is possible to better consider the differences in data of each node during the model fusion process, improve the performance and generalization ability of the global model, effectively avoid model bias caused by data heterogeneity, and enable the global model to better adapt to the data characteristics of different nodes.
[0083] In step S300, based on the local data of trusted nodes, calculate the distribution characteristics of each node and the distribution similarity of data between nodes. Then, in step S400, determine the data heterogeneity evaluation index according to these distribution characteristics and similarities, and configure a weight coefficient for each node. In this way, when the model is fused, the differences in data distributions of different nodes can be fully considered, avoiding the impact on model performance caused by uneven data distribution.
[0084] S500. Perform weighted average fusion on the model parameters of the corresponding pre-trained local models uploaded by each node and the weight coefficients of each node to generate global model parameters.
[0085] Among them, configuring a weight coefficient for each node means configuring a weight coefficient for the node of each Internet of Things local device.
[0086] In this step, the local model information of each node can be fused, making full use of the advantages of different nodes to improve the performance and generalization ability of the global model; through weighted average fusion, weights can be reasonably allocated according to the importance of each node, enabling the global model to more accurately reflect the characteristics of the entire Internet of Things system.
[0087] In S600, the global model parameters are sent to each node as initial parameters, and iterative optimization of the corresponding pre-trained local models is performed separately at each node. The pre-trained local models that meet the iteration conditions are denoted as target models.
[0088] Specifically, the generated global model parameters can be sent to each trusted node through a central server; each node uses the global model parameters as the initial value and iteratively optimizes the pre-trained local model with local data; the iteration conditions can be reaching a certain number of training rounds or the loss function of the model converging to a smaller value.
[0089] For example, set the number of training rounds to 100 times. After the model training of each node reaches 100 times, check whether the loss function of the model is less than a preset threshold. If the condition is met, mark this model as the target model.
[0090] Through iterative optimization, the performance of each node's model can be further improved, enabling it to better adapt to local data; using the global model parameters as the initial value speeds up the convergence rate of the model and reduces the training time.
[0091] In S700, all the target models are respectively deployed to the monitoring nodes of the corresponding local devices to generate an integrated monitoring system.
[0092] Specifically, this deployment process takes into account the hardware conditions and operating environments of different monitoring nodes; first, an environment check is performed on each node, including computing resources, storage space, network bandwidth, etc.; according to the check results, the system automatically adjusts the running parameters of the model, such as batch size, inference frequency, etc., to ensure that the model can run stably on each node; at the same time, we also need to establish a data caching mechanism and a fault recovery mechanism to cope with possible network fluctuations or hardware failures; after the deployment is completed, each monitoring node will perform an initialization test to confirm that all functions of the monitoring system are running properly.
[0093] Furthermore, the target model of each node is deployed to the monitoring node of the corresponding local device. For example, the target model of the machine tool is deployed to the monitoring system of the machine tool, and the target model of the robot is deployed to the monitoring module of the robot. Data interaction and collaborative work can be carried out between each monitoring node to form an integrated monitoring system.
[0094] Through this step, real-time monitoring of each local device can be achieved, and abnormal behaviors of the devices can be detected in a timely manner. The integrated monitoring system can integrate the information of each node and provide more comprehensive and accurate monitoring and analysis of device behaviors.
[0095] In S800, based on the integrated monitoring system, the real-time behavior data of the target device collected is analyzed to generate a device behavior authentication result.
[0096] Specifically, the integrated monitoring system collects the behavioral data of the target device in real time, inputs the collected real-time data into the corresponding target model for analysis, and determines whether the behavior of the device is normal according to the output result of the model. For example, if the deviation between the output result of the model and the normal behavior pattern exceeds a certain threshold, it is considered that the device has abnormal behavior, and the corresponding authentication result is generated.
[0097] Through this step, the abnormal behavior of the device can be detected in time, providing a basis for the maintenance and management of the device, reducing device failures and production losses; through the device behavior authentication result, the security and reliability of the Internet of Things system can be improved.
[0098] The method for monitoring the behavior of Internet of Things devices disclosed in this application trains a discrete denoising diffusion model at each node of each Internet of Things local device according to the distribution of Internet of Things devices. The trained discrete denoising diffusion model is recorded as the pre-trained local model of the corresponding node. This distributed training strategy can better adapt to the heterogeneity of various devices; extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviations of each node; calculate the distribution characteristics of each node and the distribution similarity of the data between nodes based on the local data of the trusted nodes; determine the data heterogeneity evaluation index based on the distribution characteristics and distribution similarity, and configure the weight coefficient for each node according to the data heterogeneity evaluation index; perform weighted average fusion on the model parameters of the corresponding pre-trained local model uploaded by each node and the weight coefficient of each node to generate global model parameters; use the global model parameters as the initial parameters and send them to each node, and perform iterative optimization of the corresponding pre-trained local model at each node. The pre-trained local model that meets the iterative conditions is recorded as the target model; deploy all the target models to the monitoring nodes of the corresponding local devices to generate an integrated monitoring system; analyze the real-time behavior data of the target device collected based on the integrated monitoring system to generate a device behavior authentication result, making full use of the local data of distributed devices, avoiding the data limitation of a single monitoring center, being able to effectively integrate the data of different devices, expanding the data coverage of the monitoring system, being able to dynamically adjust the model parameters according to the actual operation situation of the device, improving the flexibility of the monitoring system, being able to generate a customized monitoring system for different scenarios, enhancing the adaptability to diverse scenarios, and at the same time being able to accurately and quickly output the monitoring result, with strong self-adaptability, not requiring manual frequent manual adjustment of rules to adapt to the upgrade and change of the system, and effectively reducing the maintenance cost of the system.
[0099] The method for monitoring the behavior of Internet of Things devices disclosed in this application makes full use of the local data and advantages of each node through local training, model fusion, and iterative optimization, improving the performance and generalization ability of the model; most of the data processing and model training are carried out on local nodes, effectively reducing data transmission and sharing and protecting the privacy of data; through the screening of trusted nodes and the evaluation of data heterogeneity, the interference of abnormal nodes and data is effectively excluded, improving the reliability and stability of the system; the target model is deployed to the local device monitoring node to achieve real-time monitoring and behavior authentication of the device, and abnormal behaviors of the device can be detected in a timely manner.
[0100] Referring to Figure 2 , for S100 "According to the distribution of Internet of Things devices, train a discrete denoising diffusion model at each node of the Internet of Things local device, and record the trained discrete denoising diffusion model as the pre-trained local model corresponding to the node", that is, the method for obtaining the pre-trained local model specifically includes:
[0101] S110, according to the distribution of Internet of Things devices, configure the Director node and several Envoy nodes through the OpenFL framework and establish a secure communication channel to obtain the constructed federated learning network topology.
[0102] The federated learning network topology constructed through the OpenFL framework allows each Internet of Things device node to perform data processing and model training locally, and then coordinate and exchange information through the Director node, realizing distributed collaborative learning and improving the scalability and flexibility of the system. The established secure communication channel uses encryption technology to ensure security during data transmission and protect the sensitive data of Internet of Things devices from being leaked.
[0103] In actual deployment, it is first necessary to investigate and analyze the distribution of Internet of Things devices. For example, assume that in an intelligent factory scenario, there are 100 production devices distributed in 5 different workshops. Each workshop is equipped with an edge server as an Envoy node, and a high-performance server is deployed in the central computer room of the factory as the Director node.
[0104] When configuring using the OpenFL framework, first install the Director component of OpenFL on the central server and configure its listening address as 192.168.1.100:50051. Then deploy the Envoy component of OpenFL on the edge server in each workshop and configure its connection address to point to the Director node. To ensure communication security, OpenFL uses TLS encryption and two-way authentication mechanisms, and SSL certificates need to be configured for each node. The specific configuration includes:
[0105] After the configuration is completed, each Envoy node will automatically establish a secure gRPC communication channel with the Director. The final formed network topology is a star topology, with the Director node at the center, connected to the surrounding 5 Envoy nodes through encrypted channels, constituting a reliable federated learning network infrastructure.
[0106] S120, based on the federated learning network topology, collect the corresponding device operation data at each Envoy node, and after preprocessing the device operation data, obtain a standardized behavior data set.
[0107] Among them, the device operation data includes data such as network traffic characteristics, resource usage, and operation logs. The network traffic characteristics include: the number of TCP / UDP connections, packet size distribution, communication frequency, etc.; the resource usage includes: CPU usage, memory occupancy, disk I / O, network bandwidth, etc.; the operation logs include: device start / stop records, configuration modifications, access controls, etc.
[0108] By preprocessing the device operation data, noise and error data are removed, improving the quality and usability of the data, providing a more reliable basis for subsequent model training. The standardized behavior data set makes the data of different Envoy nodes have a unified format and feature scale, facilitating subsequent model training and fusion on each node.
[0109] S130, configure the architecture of the discrete denoising diffusion model based on the behavior data set, and denote the configured discrete denoising diffusion model as the first model; the architecture includes an encoder, a noise predictor, and a decoder.
[0110] Among them, the encoder uses a multi-layer convolutional neural network to convert the input behavior data into a latent space representation; the noise predictor adopts a U-Net architecture to predict the noise added during the diffusion process; the decoder restores the denoised latent representation to behavior data through a transposed convolutional network.
[0111] Specifically, there is a multi-layer convolutional neural network in the encoder, that is, it includes a sequence of convolutional layers and a fully connected layer. Among them, the sequence of convolutional layers contains three convolutional layers, and each convolutional layer is followed by a ReLU activation function. The first convolutional layer converts the number of input channels to 64, the second convolutional layer increases the number of channels from 64 to 128, and the third convolutional layer increases the number of channels from 128 to 256; the output of the convolutional layer is converted into a latent space representation of a specified dimension through a fully connected layer; the input dimension of the fully connected layer is the total dimension of the feature map output by the convolutional layer, and the output dimension is latent_dim.
[0112] The noise predictor adopts a U-Net architecture to predict the noise added during the diffusion process;
[0113] Specifically, the noise predictor includes an encoding part and a decoding part. The encoding part consists of three DoubleConv modules and two max pooling layers. The DoubleConv module contains two convolutional layers and two ReLU activation functions for feature extraction. The max pooling layer is used for downsampling to reduce the size of the feature map.
[0114] The decoding part consists of three DoubleConv modules and two transposed convolutional layers. The transposed convolutional layer is used for upsampling to restore the size of the feature map. During the decoding process, the feature map of the encoding part is concatenated with the feature map of the decoding part to retain more detailed information.
[0115] The role of the decoder is to restore the denoised latent representation to behavioral data. Specifically, the decoder includes a fully connected layer and transposed convolutional layers. In specific implementation, a fully connected layer is used to convert the latent space representation into a larger feature vector. The sequence of transposed convolutional layers contains three transposed convolutional layers, each followed by a ReLU activation function. The number of output channels of the last transposed convolutional layer is the same as the number of channels of the input behavioral data, and the Sigmoid activation function is used to limit the output value between 0 and 1.
[0116] The specific initialization method includes: dividing the dataset into a training set and a test set. The training set is used to let the model learn the features and patterns of the data, while the test set is used to evaluate the performance of the model on unseen data. In the encoder, a multi-layer convolutional neural network is used to convert the behavioral data in the input behavioral dataset into a latent space representation. The decoder restores the denoised latent representation to behavioral data through a transposed convolutional network.
[0117] In this step, the model architecture is configured according to the characteristics of the behavioral dataset, enabling the model to better adapt to the behavioral data characteristics of IoT devices, improving the accuracy and performance of the model. The architectures of the encoder, noise predictor, and decoder of the discrete denoising diffusion model can be flexibly adjusted according to different IoT devices and data characteristics, with strong generality and adaptability.
[0118] S140, perform local training on the first model for forward diffusion and backward denoising at the corresponding Envoy node, and record the trained first model as the pre-trained local model of the corresponding Envoy node.
[0119] After the model architecture is determined, the local training process begins. During the forward diffusion process, Gaussian noise is gradually added to the original behavioral data. The backward denoising process restores the original data by predicting and removing the noise. After training, each Envoy node obtains a pre-trained local model that can model the behavior of local devices. These local models will be aggregated and optimized in subsequent steps.
[0120] Specifically, during the forward diffusion process, we gradually add Gaussian noise to the original behavioral data to simulate the degradation process of the data, making the data gradually become more noisy. The specific steps are as follows:
[0121] 1. Initialization: For the standardized behavioral data set x0 obtained by collecting and preprocessing from local devices, where each sample x0 represents a specific device behavior data instance.
[0122] 2. Gradual noise addition: According to the preset time step T, at each time step t (1 ≤ t ≤ T), add Gaussian noise ∈ t-1 to the current data x t-1 to obtain the noisy data x t . This process can be represented by the following formula: where, α t is a predefined attenuation coefficient used to control the proportion of noise added at each time step, and ∈ t-1 is a noise vector sampled from the standard normal distribution.
[0123] 3. Encoder processing: During the forward propagation of the model, the noisy data x t is processed through the encoder to convert it into a latent space representation z t . The encoder usually consists of a series of neural network layers, such as convolutional layers or fully connected layers, which are used to extract the features of the data and map them to the latent space.
[0124] Detailed description of the reverse denoising process: The goal of the reverse denoising process is to recover the original behavioral data from the noisy data. This is an iterative process that starts from time step T and gradually removes the noise until the original data x0 is recovered. The specific steps are as follows: 1. Noise prediction: At each time step t (1 ≤ t ≤ T), input the current latent representation z t into the noise predictor to predict the noise added at this time step . The noise predictor is also a neural network that predicts the noise at each time step by learning the noise patterns in the data. 2. Denoising operation: According to the predicted noise perform a denoising operation on the current latent representation z t to obtain the denoised latent representation z t-1 .
[0125] The denoising operation can be implemented by the following formula: ∈’ t-1 . Where, ∈’ t-1is a new noise vector sampled from the standard normal distribution, which is used to introduce a certain degree of randomness during the denoising process to help the model learn a more robust denoising strategy.
[0126] 3. Decoder processing: Process the denoised latent representation z t-1 through the decoder to restore it to behavioral data The structure of the decoder is usually the opposite of that of the encoder. Through a series of transposed convolutional layers or fully connected layers, the latent space representation is mapped back to the original data space.
[0127] 4. Iterative denoising: Repeat steps 1 - 3 until t = 1 to finally obtain the restored original behavioral data
[0128] Training process and loss calculation: During the training process, we calculate the loss function by comparing the restored behavioral data with the original behavioral data x0. The commonly used loss function is the mean squared error (MSE). By minimizing the loss function and using an optimization algorithm (such as stochastic gradient descent) to update the model's parameters, the model can perform better forward diffusion and backward denoising operations.
[0129] After multiple iterations of training, when the loss function converges to a small value, the training process ends. At this time, each Envoy node has obtained a pre - trained local model that can model the behavior of local devices. These local models will be aggregated and optimized in subsequent steps to obtain a more powerful and general global model.
[0130] Performing local training at Envoy nodes makes full use of the local data of each node, avoids the centralized transmission and storage of data, and protects data privacy; each Envoy node trains according to local data, enabling the pre - trained local model to better adapt to the behavior characteristics of local devices, improving the pertinence and accuracy of the model.
[0131] Refer to Figure 3 , for "extracting the activation values of each node during the process of training the discrete denoising diffusion model, and obtaining all trusted nodes according to the calculated activation value deviations of each node" in S200, that is, the method for obtaining all trusted nodes specifically includes:
[0132] S210, extracting the activation values of neurons in each layer during the training process of the discrete denoising diffusion model to generate an activation value statistical feature set.
[0133] Specifically refer to Figure 4 , the method for generating the activation value statistical feature set specifically includes:
[0134] S211. During the training of the discrete denoising diffusion model, observation points are added to each layer of the neural network to collect the activation values of each neuron in real time.
[0135] Specifically, during the training of the discrete denoising diffusion model, a hook function is used to add observation points at the output of each layer. For example, in the PyTorch framework, this can be achieved by registering a forward hook. The hook function will be triggered after the forward propagation of each layer of the neural network, and the activation values of each neuron in that layer will be recorded in real time. Suppose the model has L layers, each layer has N neurons, and it is trained for T steps. Then the collected activation value data will form a three-dimensional array with a size of L×N×T.
[0136] In this step, the activation values are collected in real time during training to ensure the timeliness and accuracy of the data; the activation values of all neurons in all layers are collected, providing a comprehensive data basis for subsequent analysis; using the hook function does not require modifying the model structure, maintaining the integrity of the model.
[0137] S212. The activation value data of each neuron is collected layer by layer to form a multi-dimensional data set, and each data point in the data set represents the activation state of a neuron at a specific training step.
[0138] Specifically, the activation value data collected in S211 is classified layer by layer to form a structured multi-dimensional data set. For example, use the NumPy library to create a three-dimensional array activation_values, where activation_values[layer][neuron][step] represents the activation value of the neuron at the neuron position in the layer layer at the training step step. Suppose the model has 10 layers, each layer has 100 neurons, and it is trained for 1000 steps. Then the size of activation_values is 10×100×1000.
[0139] Organizing the data by layer and neuron facilitates subsequent statistical analysis and feature extraction; each data point corresponds to a specific layer, neuron, and training step, facilitating backtracking and analysis; the structure of the multi-dimensional array facilitates batch processing and calculation, improving the efficiency of data processing.
[0140] S213. Perform statistical analysis on the collected activation value data, calculate the statistical features of the activation values, and the statistical features of the activation values include the mean of the activation data, the variance of the activation data, the skewness of the activation data, and the kurtosis of the activation data.
[0141] Specifically, for the mean: np.mean can be used to calculate the average activation value of each neuron over all training steps. For the variance: np.var can be used to calculate the variance of the activation values of each neuron over all training steps. For the skewness: scipy.stats.skew can be used to calculate the skewness of the activation values of each neuron over all training steps. For the kurtosis: scipy.stats.kurtosis can be used to calculate the kurtosis of the activation values of each neuron over all training steps.
[0142] Assume that the activation value sequence of a certain neuron is [a1, a2,..., aT]. Then its statistical features can be expressed as: mean: μ = mean([a1, a2,..., aT]); variance: σ 2 = var([a1, a2,..., aT]); skewness: γ1 = skew([a1, a2,..., aT]); kurtosis: γ2 = kurtosis([a1, a2,..., aT]).
[0143] By calculating the mean, variance, skewness, and kurtosis, the activation behavior of neurons is quantified into specific numerical values, facilitating subsequent analysis and comparison; these statistical features can describe the distribution characteristics of activation values, such as symmetry, kurtosis, etc., providing an in-depth perspective for understanding the behavior of neurons; abnormal changes in statistical features can be used as a basis for detecting abnormal nodes, helping to identify abnormal behaviors in the model.
[0144] S214. Classify the statistical features of activation values by layer and neuron organization to generate a multi-dimensional statistical feature set of activation values.
[0145] Specifically, classify the statistical features calculated in S213 by layer and neuron to form a structured multi-dimensional statistical feature set of activation values. For example, use the Pandas library to create a DataFrame, where the index is the combination of layer and neuron, and the columns are the names of the statistical features (mean, variance, skewness, kurtosis). Assume the model has L layers and each layer has N neurons. Then the size of the DataFrame is L×N rows and 4 columns.
[0146] Classifying the statistical features by layer and neuron facilitates subsequent analysis and application; using the DataFrame structure, specific statistical features can be quickly accessed by layer and neuron; the generated feature set can be directly used to construct a standard activation mapping pattern or perform anomaly detection, providing strong support for subsequent model optimization.
[0147] The generation method of the activation value statistical feature set disclosed in S213 - S214 ensures the comprehensiveness and timeliness of data by collecting the activation values of each neuron in each layer in real time; organizes the data by layer and neuron to form a structured data set, facilitating subsequent statistical analysis and feature extraction; quantifies the activation behavior of neurons into specific numerical values by calculating the mean, variance, skewness, and kurtosis, providing an in - depth perspective for understanding the behavior of the model; the abnormal changes in statistical features can be used as a basis for detecting abnormal nodes, helping to identify abnormal behaviors in the model and enhancing the robustness and reliability of the model; by analyzing the statistical features of activation values, abnormal nodes can be identified and removed to optimize the performance and generation quality of the model; by understanding the activation behavior of neurons, the interpretability of the model can be improved, helping researchers better understand the working principle of the model.
[0148] S220. Based on the activation value statistical feature set, calculate the activation value distribution parameters of normal nodes and construct a standard activation mapping pattern.
[0149] Specifically refer to Figure 5 , the construction method of the standard activation mapping pattern includes:
[0150] S221. Determine normal nodes.
[0151] Specifically, during the training process, by observing the activation value behavior of nodes, initially screen out nodes with stable performance. For example, select nodes with small fluctuations in activation values and no obvious abnormalities. In a federated learning network, most nodes are honest.
[0152] S222. Conduct a clustering analysis on the activation value statistical features of normal nodes to identify the target clustering centers.
[0153] Among them, clustering methods such as K - means and DBSCAN can be used to conduct a clustering analysis on the activation value statistical features of normal nodes to obtain the center points of each cluster as the target clustering centers.
[0154] Through clustering analysis, normal nodes are divided into several groups, facilitating subsequent calculation of distribution parameters; the clustering centers represent the typical features of each group of nodes, providing a basis for constructing the standard activation mapping pattern.
[0155] S223. Calculate the activation value distribution parameters of each target clustering center.
[0156] Among them, the activation value distribution parameters include the mean, variance, skewness, kurtosis, etc., and these parameters represent the typical activation value distribution characteristics of normal nodes.
[0157] Quantify the features of the cluster centers into specific distribution parameters for subsequent pattern construction; the distribution parameters can describe the typical activation value distribution characteristics of each group of nodes, providing an in-depth perspective for understanding the behavior of the model.
[0158] S224. Combine the activation value distribution parameters of all the cluster centers into a standard activation mapping pattern.
[0159] In this embodiment, the standard activation mapping pattern represents the behavioral characteristics that normal nodes should exhibit during the training process and will be dynamically updated as the training progresses.
[0160] Organize the distribution parameters by cluster center to form a structured standard activation mapping pattern; using the DataFrame structure, specific distribution parameters can be quickly accessed through the cluster center; the generated standard activation mapping pattern can be directly used for anomaly detection or model optimization.
[0161] The method for constructing the standard activation mapping pattern disclosed in S221 - S224 ensures the data quality for subsequent clustering analysis and pattern construction by screening and detecting normal nodes; through clustering analysis, normal nodes are divided into several groups, and the cluster centers are used as representatives, providing a basis for constructing the standard activation mapping pattern; by calculating the distribution parameters, the features of the cluster centers are quantified into specific values, providing an in-depth perspective for understanding the behavior of the model; the standard activation mapping pattern is stored in a structured manner, facilitating subsequent access and analysis, enhancing the practicality of the overall solution; the standard activation mapping pattern can be used as a benchmark for detecting abnormal nodes, improving the robustness and reliability of the model; by understanding the typical activation value distribution characteristics of normal nodes, the training and generation process of the model can be optimized, improving the performance of the model; by constructing the standard activation mapping pattern, the internal mechanism of the model can be better understood, improving the interpretability of the model.
[0162] This solution systematically screens normal nodes, conducts clustering analysis, calculates distribution parameters, and finally constructs a standard activation mapping pattern, which not only provides strong support for anomaly detection and model optimization, but also improves the robustness, performance, and interpretability of the model, providing an important technical guarantee for the training and application of the discrete denoising diffusion model.
[0163] S230. Use the standard activation mapping pattern to calculate the activation value deviation of each node and perform anomaly detection to generate a list of Byzantine nodes.
[0164] Refer to Figure 6 , the method for generating the list of Byzantine nodes specifically includes:
[0165] S231. Calculate the difference between the activation value distribution of each node and the standard activation mapping pattern.
[0166] Specifically, the Euclidean distance, KL divergence, Wasserstein distance, etc. can be used for difference calculation.
[0167] Suppose we have a neural network system with n nodes. At each time step, each node generates an activation value. Taking the Euclidean distance as an example, we calculate the difference as follows: 1) Collect activation values: Record the activation values of each node over a period of time to form an activation value sequence. For example, for node i, its activation value sequence is A i =[a i1 ,a i2 ,…,a it , where t is the number of time steps. 2) Construct the standard activation mapping pattern: Through statistical analysis of the activation values of a large number of normal nodes, obtain the standard activation value distribution. Suppose the standard activation value distribution is a normal distribution with a mean of μ and a standard deviation of σ. 3) Calculate the difference: For node i, we can calculate the mean and standard deviation s i of its activation value sequence, and then calculate its Euclidean distance from the standard activation value distribution:
[0168] By calculating the difference, we can quantitatively compare the activation value distribution of the node with the standard pattern, so as to more objectively judge whether the behavior of the node is normal; using common difference metrics (such as Euclidean distance, KL divergence, Wasserstein distance, etc.) makes the calculation results have a certain interpretability, facilitating subsequent analysis and processing.
[0169] S232, based on the difference, obtain the degree of shape deviation of the activation value distribution and the change trend in the time series, and mark the corresponding nodes that meet either the degree of shape deviation exceeding the preset threshold or the change trend reaching the preset condition as potential Byzantine nodes.
[0170] Among them, the degree of shape deviation can be reflected by skewness and kurtosis. The change trend in the time series can be reflected by the mean change rate of the sliding window.
[0171] Specifically, suppose the skewness of the standard activation value distribution is skew std , and the kurtosis is kurt std . For node i, its skewness is skew i , and the kurtosis is kurt i , then the degree of shape deviation can be expressed as:
[0172] Change trend in time series: Use time series analysis methods, such as the difference method or the moving average method, to detect the change trend of the node activation value. For example, calculate the difference Δa of the activation values of node i at adjacent time steps ij = a i,j+1 - a ij , if the absolute value of the difference at a certain time step exceeds a preset threshold τ, it is considered that the behavior pattern of the node has suddenly changed drastically.
[0173] Or, if the deviation degree of the shape of node i exceeds a preset threshold, or its behavior pattern has suddenly changed drastically, then mark the node as a potential Byzantine node.
[0174] By comprehensively considering the deviation degree of the shape of the activation value distribution and the change trend in the time series, the behavior of the node can be evaluated more comprehensively, reducing the possibility of misjudgment; by detecting the change trend in the time series, sudden changes in the node behavior pattern can be discovered in a timely manner, so as to discover potential Byzantine nodes earlier.
[0175] Among them, further, the deviation degree of the shape exceeding the preset threshold can be understood as the activation value distribution significantly deviating from the standard pattern; the change trend reaching the preset condition can be understood as the behavior pattern suddenly changing drastically.
[0176] S233, divide the entire time series into Q time windows;
[0177] Calculate the deviation degree of the shape of the activation value distribution and the change trend in the time series for each node in each time window. If each node is marked as a potential Byzantine node P times in Q time windows, then mark the node as a target Byzantine node.
[0178] Among them, Q ≥ 2.
[0179] By comprehensively considering the behavior characteristics in multiple time windows, misjudgment can be effectively avoided while ensuring the reliability of the detection results.
[0180] For example, the entire time series can be divided into multiple time windows, and the length of each time window is w. For example, the time series is [a1, a2, …, a T , then it can be divided into time windows, and the jth time window is [a (j-1)w+1 , a (j-1)w+2 , …, a jw .
[0181] Calculate the behavioral characteristics within each time window: For each time window, repeat steps S231 and S232 to calculate the difference between the activation value distribution of each node within this time window and the standard activation mapping pattern, as well as the degree of deviation of the shape of the activation value distribution and the change trend in the time series. If a certain node is marked as a potential Byzantine node in multiple time windows, then this node is considered to be a real potential Byzantine node; otherwise, mark it as a normal node.
[0182] By comprehensively considering the behavioral characteristics within multiple time windows, misjudgments caused by accidental factors can be reduced, and the reliability of the detection results can be improved. The division of time windows can smooth the noise in the node activation values, making the detection results more stable.
[0183] S234. Summarize the IDs of the marked target Byzantine nodes to generate a list of Byzantine nodes.
[0184] After completing step S233, summarize the IDs of all the marked potential Byzantine nodes into a list.
[0185] For example, assume that the IDs of the marked target Byzantine nodes are [id1, id2, …, id k , then the list of Byzantine nodes is L = [id1, id2, …, id k .
[0186] After generating the list of Byzantine nodes, it can be conveniently used as input to apply subsequent filtering rules, such as removing these nodes from the system or conducting further reviews on them; the list of Byzantine nodes can be used for visualizing and monitoring the security of the system, helping administrators to timely discover and handle potential security threats.
[0187] S240. Obtain the filtered trusted nodes according to the list of Byzantine nodes and the filtering rules.
[0188] Refer to Figure 7 , for S240 “Obtain the filtered trusted nodes according to the list of Byzantine nodes and the filtering rules”, that is, the filtering method for trusted nodes, specifically includes:
[0189] S241. Configure corresponding filtering rules according to the historical records of each node;
[0190] The historical records include historical abnormal information, historical behavior scores, and historical credit records.
[0191] First, we need to collect the historical records of each node, which include historical anomaly information, historical behavior scores, and historical credit records. Historical anomaly information reflects the degree of anomalies that have occurred in the node in the past, such as the frequency or severity of anomaly events like data transmission errors or connection interruptions in the node. The historical behavior score is a quantitative assessment of the node's past behavior performance, for example, whether the node completes tasks on time and whether it adheres to network rules. The historical credit record is a record of the node's long-term credit status, which may be comprehensively derived based on factors such as the node's cooperation in the network and whether there are any violations.
[0192] Based on these historical records, we configure the filtering rules. The filtering rules need to comprehensively consider multiple aspects. For example, set an anomaly degree threshold. When the anomaly degree of a node exceeds this threshold, the node may be filtered out. At the same time, for nodes with high historical behavior scores and good credit records, even if the current anomaly degree is relatively high, the anomaly degree threshold can be appropriately relaxed to keep them. For example, a node has always performed well and has a high credit, but has occasionally had some anomalies. Then it should not be easily judged as an untrusted node.
[0193] The historical performance of each node is different. By configuring personalized filtering rules for each node, its credibility can be evaluated more accurately. The roles and behavior patterns of different nodes in the network may vary greatly. Using personalized rules can avoid judging all nodes with a unified standard, thereby improving the accuracy of screening. Configuring rules by combining multiple dimensions of historical anomaly information, historical behavior scores, and historical credit records can comprehensively evaluate the situation of the node. A single evaluation index may be one-sided, while combining multiple indicators can more comprehensively reflect the true situation of the node and make the filtering rules more reasonable.
[0194] S242, adaptively adjust the filtering rules according to the overall situation of the network.
[0195] For example, when the overall anomaly degree of the network is relatively high, the filtering criteria can be appropriately relaxed.
[0196] Specifically, to measure the overall situation of the network, we can evaluate it through some indicators. For example, calculate the average anomaly degree of all nodes. When the overall anomaly degree of the network is relatively high, it indicates that the network may be in an unstable state. At this time, the probability of nodes having anomalies is relatively high. In this case, we can appropriately relax the filtering criteria. For example, the original anomaly degree threshold is 0.5. When the overall anomaly degree of the network is relatively high, the threshold can be increased to 0.6, so that more nodes can be retained to ensure the basic operation of the network. On the contrary, when the overall anomaly degree of the network is relatively low, it indicates that the network is operating relatively stably, and the filtering criteria can be maintained or tightened to ensure that the selected nodes have a high credibility.
[0197] The network environment is constantly changing. Under different time periods and different network loads, the abnormal conditions of nodes may vary greatly. Adaptive adjustment of filtering rules can make the filtering mechanism change with the change of network conditions, and always maintain a good filtering effect; when the overall abnormal degree of the network is relatively high, relaxing the filtering criteria can ensure that a certain number of nodes participate in the network operation and maintain the basic performance of the network. When the network condition is good, strict filtering can improve the security of the network and avoid the harm caused by untrusted nodes to the network.
[0198] S243. Apply the filtering rules to the list of Byzantine nodes, score and determine each node, and mark the nodes retained after being screened by the filtering rules as trusted nodes.
[0199] For a given list of Byzantine nodes, we apply the previously configured and adjusted filtering rules to each node. For each node, we compare its current abnormal degree with the abnormal degree threshold in the corresponding filtering rules. If the current abnormal degree of the node is lower than the threshold and meets the requirements of other filtering rules, the node is determined to be a trusted node and marked. For example, if the current abnormal degree of a node is 0.3 and the adjusted abnormal degree threshold corresponding to it is 0.5, then the node passes the screening and is marked as a trusted node.
[0200] By applying the filtering rules to the list of Byzantine nodes, it is possible to clearly determine which nodes are trusted and which are not. This provides a clear basis for subsequent network operations. For example, when performing data transmission, task allocation and other operations, trusted nodes can be preferentially selected to improve the reliability of the operations. Filtering out untrusted nodes and only retaining trusted nodes to participate in the network operation can reduce potential risks caused by untrusted nodes, such as data leakage, network attacks, etc., thereby improving the reliability and stability of the entire network.
[0201] The filtering method for trusted nodes disclosed in S241 - S243 comprehensively considers the historical records of nodes and the overall network situation, configures personalized and dynamically adjusted filtering rules, and can accurately screen out trusted nodes from the list of Byzantine nodes. This helps improve the security of the network and reduce the threat of untrusted nodes to the network. The adaptive adjustment of filtering rules enables the solution to adapt to different network environments and changes in network conditions. Whether in an environment with high network load and many abnormal situations or in a stable network operation environment, it can maintain a good filtering effect, with strong flexibility and adaptability. On the premise of ensuring network security, by reasonably adjusting the filtering criteria, it is possible to balance the security and performance of the network to a certain extent. It avoids excessive filtering resulting in too few available nodes in the network, affecting the normal operation of the network, and also avoids overly broad filtering criteria leading to untrusted nodes mixing into the network and reducing the security of the network.
[0202] Through the above steps, the system can effectively identify and filter malicious nodes while maintaining inclusiveness for the behavioral differences of normal nodes, providing a reliable guarantee for the secure operation of the entire federated learning system.
[0203] The Byzantine defense mechanism is a mechanism used to address possible faults or malicious behaviors of nodes in a distributed system. Its core is to ensure that the system can still achieve consistency and correctness in the presence of unreliable or malicious nodes. The Byzantine defense mechanism originated from the "Byzantine Generals Problem", which is a classic distributed system problem proposed by Leslie Lamport et al. in 1982. The problem describes that the generals of the Byzantine Empire need to pass messages through messengers to reach a consistent battle plan, but some generals may be traitors and send false information to interfere with the decisions of other generals. The core objectives of the Byzantine defense mechanism include: 1) Fault tolerance: Even if some nodes (such as generals or computers) fail or behave abnormally, the system can still operate normally; 2) Consistency: All normal nodes can reach a consistent decision without being affected by malicious nodes.
[0204] In federated learning, the Byzantine defense mechanism is used to prevent malicious clients from interfering with the training of the global model by sending incorrect model updates.
[0205] Furthermore, in the federated learning network, to identify and guard against attacks from Byzantine nodes, we first need to extract the activation value features of neurons from the pre-trained local models. During the model training process, the system collects the activation status of each neuron by adding observation points at each layer. After processing these raw activation value data, we can obtain a statistical feature set including mean, variance, skewness, kurtosis, etc. These feature sets precisely characterize the behavioral features of each node during the training process, laying the foundation for subsequent anomaly detection.
[0206] With these activation value statistical features, the next step is to construct the standard behavioral patterns of normal nodes. This process is based on an important assumption: in the federated learning network, most nodes are honest. We use the method of cluster analysis to identify the main activation value distribution patterns. By analyzing the characteristics of these distribution patterns, we can establish a standard activation mapping pattern, which represents the behavioral features that normal nodes should exhibit during the training process. This standard pattern will be dynamically updated as the training progresses to adapt to the evolution process of the model.
[0207] Based on the established standard activation mapping pattern, we can start to identify abnormal Byzantine nodes. The system calculates the degree of difference between the activation value distribution of each node and the standard pattern. This calculation of the difference mainly considers two aspects: one is the degree of deviation of the distribution shape, and the other is the change trend in the time series. If the activation value distribution of a certain node significantly deviates from the standard pattern, or its behavioral pattern shows a sudden drastic change, this node will be marked as a potential Byzantine node. To avoid misjudgment, the system comprehensively considers the behavioral features within multiple time windows to ensure the reliability of the detection results.
[0208] Finally, based on the list of identified Byzantine nodes, we need to design and apply corresponding filtering rules. These rules not only consider the current degree of abnormality of the node but also refer to its historical performance and reputation record. The filtering rules adopt a dynamically adjusted mechanism, which can adaptively adjust the judgment criteria according to the overall situation of the network. Through this filtering mechanism, we finally obtain a set of trusted nodes. This set of trusted nodes will be used in the subsequent model aggregation process to ensure the security and reliability of the federated learning system.
[0209] It should be noted that the entire activation mapping misidentification mechanism is a continuously running process. The system will regularly update the standard pattern and adjust the judgment rules to adapt to the dynamic changes in the network environment. At the same time, to handle the challenges brought by the non-IID environment, the design of the filtering rules specifically considers the situation of uneven data distribution, ensuring that the system can effectively identify malicious nodes without misjudging normal nodes that only show differences due to different data characteristics.
[0210] This Byzantine defense mechanism based on activation mapping establishes an accurate and flexible security protection system by analyzing the behavioral characteristics within deep learning models. It can effectively identify and filter malicious nodes while maintaining inclusiveness towards the behavioral differences of normal nodes, providing reliable guarantee for the secure operation of the entire federated learning system.
[0211] Referring to Figure 8 , the method of S300 "calculating the distribution characteristics of each node and the distribution similarity of data between nodes based on the local data of trusted nodes" includes:
[0212] S310, calculating the distribution characteristics of each trusted node according to the obtained local data of each trusted node, where the distribution characteristics include statistical characteristics and distribution shape characteristics.
[0213] Among them, the statistical characteristics can include one or more of mean, variance, and quantile, which can reflect the basic distribution of the data; the distribution shape characteristics can be one or more of skewness and kurtosis, which help to understand the shape and central tendency of the data distribution.
[0214] In this embodiment, skewness is used to reflect the asymmetry of the data distribution. Positive skewness indicates that the data is right-skewed, and negative skewness indicates that the data is left-skewed. Kurtosis is used to reflect the sharpness of the data distribution. High kurtosis indicates that the data distribution is more concentrated, and low kurtosis indicates that the data distribution is flatter.
[0215] Suppose we have multiple trusted nodes, and each node stores a set of local data. Taking nodes A, B, and C as examples, they store different numerical data sets respectively.
[0216] The mean is the average value of a set of data, which reflects the central tendency of the data. For example, for the local data set [1, 2, 3, 4, 5] of node A, its mean is (1 + 2 + 3 + 4 + 5) / 5 = 3.
[0217] The variance measures the degree of dispersion of the data relative to the mean. For the data of node A, first calculate the square of the difference between each data point and the mean, and then find the average of these squared values. That is, [(1 - 3) 2 + (2 - 3) 2 + (3 - 3) 2 + (4 - 3) 2 + (5 - 3) 2 / 5 = 2.
[0218] Median: After arranging the data in ascending or descending order, the value at the middle position. If the number of data is odd, the median is the middle number; if it is even, it is the average of the two middle numbers. For the data of node A, the median is 3.
[0219] Statistical software or formulas can be used to calculate skewness. For example, the scipy.stats.skew function in Python can be used to calculate the skewness of the data of Node A.
[0220] Kurtosis describes the peakedness of the data distribution, that is, the degree of concentration of the data near the mean and the thickness of the tails. The scipy.stats.kurtosis function can be used to calculate the kurtosis of the data of Node A.
[0221] Statistical features and distribution shape features can help us quickly understand the basic situation of the local data of each trusted node. For example, the mean and median can tell us the central position of the data, the variance can let us understand the degree of dispersion of the data, and skewness and kurtosis can let us understand the shape of the data distribution, so as to have a deeper understanding of the data. These distribution features provide a basis for subsequent data analysis and processing. For example, when performing operations such as data fusion and anomaly detection, these features can be used as important reference bases.
[0222] S320, calculate the distribution similarity between nodes using statistical distance metrics.
[0223] Specifically, the statistical distance metrics can be Euclidean distance, Manhattan distance, Kullback-Leibler (KL) divergence, etc.
[0224] For the data vectors of two nodes, the Euclidean distance is the square root of the sum of the squares of the differences of their corresponding elements. KL divergence is used to measure the difference between two probability distributions. Assuming that the data of Node A and Node B both conform to a certain probability distribution, we can use the scipy.stats.entropy function in Python to calculate the KL divergence between them.
[0225] Through the statistical distance metrics, we can quantify the distribution similarity between nodes. This can intuitively compare the degree of difference in data distribution between different nodes, facilitating subsequent clustering analysis, node grouping, etc. Understanding the distribution similarity between nodes helps to discover the data associations between nodes. For example, if the data distributions of two nodes are similar, it may mean that they are related in terms of data source, processing method, etc., which is very helpful for mining the potential information behind the data.
[0226] The method disclosed in S310 - S320 can deeply explore the internal features of the local data of trusted nodes and the relationships between nodes by calculating the distribution characteristics of nodes and the distribution similarity of data between nodes. This helps to discover the rules and patterns in the data and provides support for further data analysis and decision-making. In a distributed system, understanding the distribution similarity between nodes can provide a basis for data fusion and collaboration. For example, for nodes with similar data distributions, more efficient data merging and sharing can be carried out to improve the overall performance of the system. By comparing the distribution characteristics and distribution similarity of nodes, it is easier to detect abnormal nodes. If the data distribution of a certain node is significantly different from that of other nodes, it may mean that there are abnormal situations in this node, such as data being tampered with or the node malfunctioning, so as to take timely measures to ensure the security of the system.
[0227] The target model obtained through S100 - S600 is the optimized model. The system will distribute the aggregated global model parameters to each node and let them perform verification and fine-tuning based on local data. In this process, we can set multiple convergence metrics, including the improvement amplitude of model performance, the amplitude of parameter updates, and the change trend of the global loss function, etc. The system will continuously monitor these metrics. When they all stabilize within a certain range, it is considered that the model has reached the convergence state. To ensure the generalization ability of the final model, we will also perform cross-validation on different nodes to ensure that the model can maintain good performance under various data distributions.
[0228] This optimization process is a cyclic iterative process, and better model parameters will be generated in each round of iteration. As the iteration progresses, the model will gradually adapt to the data characteristics of each node, while maintaining global consistency and being able to handle local special situations well. When the model finally converges, we obtain an optimized model that can adapt to non-IID data distributions and has good generalization ability.
[0229] The key to the whole process lies in balancing global consistency and local specificity. Through carefully designed weight calculation and aggregation strategies, the system can effectively handle the challenges brought by data heterogeneity while ensuring the stability and convergence of model training. This adaptive optimization scheme enables the federated learning system to perform excellently in complex actual environments and provides reliable model support for subsequent anomaly detection tasks.
[0230] Refer to Figure 9 , for S800, "Analyze the real-time behavior data of the target device collected by the integrated monitoring system to generate a device behavior authentication result", that is, the method for generating the device behavior authentication result specifically includes:
[0231] S810. Continuously collect the real-time behavior data of the target device through the integrated monitoring system to form a behavior sequence sorted by time.
[0232] Among them, the real-time behavior data can include information in multiple dimensions such as the device's network communication mode, resource usage, operation logs, etc.
[0233] Furthermore, a hierarchical data processing strategy can be adopted in the collection process, specifically including bottom-layer data processing, middle-layer data processing, and top-layer data processing. Among them, sensors or monitoring software can be used at the bottom layer to collect and preprocess the raw data. The preprocessing includes data cleaning, format conversion, and preliminary screening.
[0234] The middle layer is responsible for the temporal organization of the data, organizing the scattered data points into meaningful behavior sequences. Specifically, the data points processed at the bottom layer can be organized into behavior sequences according to the time order. For example, the data such as the number of network connections and CPU usage within every 5 minutes are combined into the data of a time window to form an element of a behavior sequence.
[0235] The top layer is responsible for feature extraction and preprocessing of the data to prepare for subsequent anomaly detection. Specifically, the top layer is responsible for extracting features from the behavior sequence, such as calculating statistical features such as the mean, standard deviation, maximum value, and minimum value of the data within each time window. At the same time, preprocess these features, such as normalization, to scale the feature values to a specific range for subsequent anomaly detection. The system can maintain a sliding window. For example, the window size is 10 time windows. As time goes by, the data within the window is continuously updated to ensure that the time continuity of the device behavior can be captured.
[0236] The hierarchical processing strategy can gradually improve the quality of the data, extracting valuable and analyzable information from the raw data; through the sliding window method, both the long-term trend of the device behavior can be captured, and sudden changes in the behavior can be detected in a timely manner, improving the accuracy of anomaly detection. Organizing the data into behavior sequences and performing feature extraction and preprocessing provides suitable inputs for subsequent analysis using models.
[0237] Continuously collecting real-time data and forming a behavior sequence in this step can comprehensively record the behavior changes of the target device over a period of time, providing a complete data basis for subsequent analysis; the data sequence sorted by time contains information in the time dimension, which helps to discover the changing rules of the device behavior over time, such as whether there are periodic peaks or troughs.
[0238] S820. Analyze the behavior sequence based on the target model to generate a behavior evaluation result.
[0239] Specifically, the behavior sequence can be preprocessed first to ensure the quality and applicability of the data; among them, the preprocessing can include data cleaning and feature standardization.
[0240] After completing the preprocessing of the behavior sequence, it can be input into the target model for prediction. Specifically, the preprocessed behavior sequence is used as the input of the target model; the target model is a discrete denoising diffusion model obtained through data fusion and iterative optimization of multiple IoT local device nodes. It can learn the characteristics and distribution laws of data from different nodes, so as to effectively analyze and predict the input behavior sequence; the discrete denoising diffusion model generates samples by gradually removing noise. When evaluating the behavior sequence, the model will predict the behavior state at the next moment according to the input behavior sequence, or give the anomaly probability of the current behavior state. For example, for the behavior sequence of a server, the model may predict the CPU usage rate, memory usage rate, etc. in the next minute, or directly give the probability value of whether the current behavior state is normal or abnormal; the model will output a prediction result vector, which contains the predicted values or anomaly scores of each feature in the behavior sequence. After obtaining the prediction result of the target model, it is necessary to convert it into a specific behavior evaluation result.
[0241] If the model directly outputs the anomaly score, this score can be directly used for subsequent evaluation. If the model outputs predicted values, the anomaly score needs to be obtained by calculating the difference between the predicted values and the actual values. For example, the mean squared error (MSE) can be calculated, that is, first find the square of the difference between the predicted value and the actual value of each data point, then sum these squared values and take the average. This mean squared error can be used as the score to measure the degree of behavior anomaly.
[0242] The target model can automatically evaluate the behavior sequence of the device, avoiding the cumbersome and subjective manual analysis, and improving the efficiency and accuracy of the evaluation; the model can learn normal behavior patterns and identify behaviors that deviate from the normal patterns, which helps to detect potential anomalies in a timely manner.
[0243] S830, according to the behavior evaluation result and the preset anomaly threshold, output the device behavior authentication result.
[0244] Specifically, the calculated anomaly scores (i.e., behavior assessment results) can be compared with preset thresholds. These thresholds are not fixed but dynamically adjusted based on historical data and expert experience. The determination process adopts a multi-level early warning mechanism: minor anomalies will generate warning messages for the operation and maintenance personnel to refer to; medium anomalies will trigger automatic protection measures, such as restricting certain operation permissions of the device; severe anomalies will immediately take isolation measures to prevent the spread of possible security incidents. The system will also record all determination results and corresponding handling measures, and these records will be used for subsequent policy optimization and system improvement.
[0245] The entire device behavior anomaly detection process is a real-time running closed-loop system. It can not only detect and respond to abnormal behaviors in a timely manner but also improve the detection accuracy through continuous learning. The system will regularly evaluate the accuracy of the detection results and adjust the detection strategy according to the feedback from the operation and maintenance personnel. This adaptive detection mechanism enables the system to gradually enhance its ability to identify various abnormal behaviors and provides reliable protection for the secure operation of Internet of Things devices.
[0246] It is worth noting that the design of this detection system particularly emphasizes practicality and interpretability. For each anomaly detection result, the system will provide a detailed analysis report, including the type of anomaly, severity, possible causes, and recommended handling solutions. This information helps the operation and maintenance personnel quickly understand the problem and make correct decisions. At the same time, the various parameters of the system can be flexibly adjusted according to actual needs to ensure that the detection system can adapt to the security requirements of different scenarios.
[0247] For example, an anomaly threshold can be preset, such as an anomaly score threshold of 0.8. Compare the anomaly score generated in step S820 with this threshold. If the anomaly score is greater than 0.8, it is determined that the device behavior is abnormal, and an authentication result of "abnormal" is output; if the anomaly score is less than or equal to 0.8, it is determined that the device behavior is normal, and an authentication result of "normal" is output.
[0248] By presetting the anomaly threshold, a clear standard is provided for the determination of device behavior, avoiding the uncertainty of subjective judgment; the simple threshold comparison operation can quickly obtain the authentication result, facilitating the system to take corresponding measures in a timely manner, such as issuing an alarm or conducting further investigations.
[0249] By analyzing and authenticating the real-time behavior data of devices, abnormal behaviors of devices can be detected in a timely manner, potential security threats such as malicious attacks and system failures can be prevented, and the safe and stable operation of devices and related systems can be ensured. Continuously monitoring device behaviors and performing authentication helps to identify performance bottlenecks and potential problems of devices, provides a basis for device maintenance and optimization, and improves the usage efficiency and reliability of devices. The results of device behavior authentication can provide decision-making support for managers, such as making more reasonable decisions in aspects such as device procurement and resource allocation.
[0250] Regarding the problems existing in the rule-based anomaly detection methods disclosed in the prior art, the IoT device behavior monitoring method disclosed in this application uses a discrete denoising diffusion model for training. This model has powerful learning ability and can learn complex behavior patterns from a large amount of IoT device data. Compared with the rule-based method, it does not require manually presetting all possible anomaly rules, but automatically discovers anomaly patterns in a data-driven manner. Even when the IoT system is upgraded or updated and the device behavior patterns change, the model can adapt to the new behavior patterns through iterative optimization, reducing the situation of normal behaviors being misjudged or abnormal behaviors being missed. The discrete denoising diffusion model can learn the latent distribution of the data. For unknown attacks or abnormal behaviors, the model can detect anomalies by identifying the deviation from the normal distribution. While the rule-based method can only detect pre-defined anomaly patterns and often cannot identify unknown anomalies.
[0251] The IoT device behavior monitoring method disclosed in this application adopts a distributed training method. The discrete denoising diffusion model is trained at each node of the IoT local device respectively to obtain a pre-trained local model. Each node only needs to upload the model parameters instead of the original data, thus avoiding the risk of data privacy leakage caused by uploading all device data to the central server. Since the training is carried out at each local device node respectively, each node can use the unique data of the local area for model training, and can make full use of the data diversity of different nodes. In the subsequent model fusion process, the model parameters of different nodes are fused by weighted averaging, so that the final global model can cover a wider range of device behavior patterns.
[0252] The method for monitoring the behavior of Internet of Things (IoT) devices disclosed in this application can better adapt to the complex and changeable IoT environment through data-driven model training and iterative optimization, accurately detect various abnormal behaviors, including unknown attack patterns, and improve the accuracy and comprehensiveness of device behavior monitoring. It makes full use of the data diversity of different nodes, enabling the model to learn a wider range of device behavior patterns and reducing the missed detection cases caused by insufficient data coverage. The distributed training method avoids the centralized upload of raw data, protects the data privacy of IoT device users, and complies with the increasingly strict data protection regulations. Considering the characteristics of uneven data distribution, by configuring weight coefficients for different nodes, the model can maintain good performance in the face of various data distribution situations. At the same time, by screening trusted nodes, it effectively resists Byzantine attacks and enhances the robustness of the model.
[0253] Compared with the rule-based anomaly detection method, this solution does not require manual and frequent adjustment of rules to adapt to system upgrades and changes, reducing the system maintenance cost.
[0254] In a second aspect, an IoT device behavior monitoring system disclosed in this application is used to execute the IoT device behavior monitoring method disclosed in the first aspect of this application, and specifically includes:
[0255] A distributed training module, which is used to train a discrete denoising diffusion model at each node of each IoT local device according to the distribution of IoT devices, and record the trained discrete denoising diffusion model as the pre-trained local model corresponding to the node;
[0256] A trusted node determination module, which is used to extract the activation values of each node during the process of training the discrete denoising diffusion model, and determine all trusted nodes according to the calculated activation value deviations of each node;
[0257] A calculation module, which is used to calculate the distribution characteristics of each node and the distribution similarity of data between nodes based on the local data of the trusted nodes;
[0258] A configuration module, which is used to determine the data heterogeneity evaluation index based on the distribution characteristics and the distribution similarity, and configure weight coefficients for each node according to the data heterogeneity evaluation index;
[0259] A parameter acquisition module, which is used to perform weighted average fusion on the model parameters of the pre-trained local model corresponding to each node uploaded by each node and the weight coefficients of each node to generate global model parameters;
[0260] An optimization module, which is used to send the global model parameters to each node as initial parameters, and perform iterative optimization of the corresponding pre-trained local model at each node, and record the pre-trained local model that meets the iteration conditions as the target model;
[0261] A deployment module, configured to deploy all the target models to the monitoring nodes of corresponding local devices respectively, so as to generate an integrated monitoring system;
[0262] An output module, configured to analyze the real-time behavior data of a target device collected based on the integrated monitoring system, so as to generate a device behavior authentication result.
[0263] According to an embodiment of the present disclosure, a computer device includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0264] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the above-mentioned Internet of Things device behavior monitoring method according to the embodiments of the present disclosure.
[0265] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.
[0266] Such as Figure 10 FIG. 10 is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiments of the present disclosure. The computer device shown in FIG. 10 is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0267] Such as Figure 10 As shown, the computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0268] Typically, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device can allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 10 a computer device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0269] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the method for monitoring the behavior of IoT devices according to the embodiments of the present disclosure are executed.
[0270] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.
[0271] A computer-readable storage medium according to an embodiment of the present disclosure stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the method for monitoring the behavior of IoT devices according to the foregoing embodiments of the present disclosure are executed.
[0272] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or removable hard disks), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0273] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.
[0274] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative purposes and for ease of understanding, rather than limitations. These details do not limit the present disclosure to necessarily implementing with the above specific details.
[0275] In the present disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0276] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing. So, for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.
[0277] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0278] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings of the technology defined by the appended claims. Additionally, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0279] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0280] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for monitoring the behavior of an Internet of Things device, characterized in that: include: According to the distribution of IoT devices, a discrete denoising diffusion model is trained at each node of a local IoT device, and the trained discrete denoising diffusion model is recorded as a pre-trained local model of the corresponding node; Extract the activation value of each node in the process of training the discrete denoising diffusion model, and determine all the trusted nodes based on the calculated activation value deviation of each node; Based on the local data of the trusted node, calculate the distribution characteristics of each node and the distribution similarity of data between nodes; Based on the distribution characteristics and the distribution similarity, determine a data heterogeneity evaluation index, and configure a weight coefficient for each node according to the data heterogeneity evaluation index; Perform weighted average fusion on the model parameters of the corresponding pre-trained local model uploaded by each node and the weight coefficient of each node to generate the global model parameters; Send the global model parameters as initial parameters to each node, perform iterative optimization of the corresponding pre-trained local model at each node, and record the pre-trained local model that meets the iteration conditions as the target model; Deploy all the target models to the monitoring nodes of the corresponding local devices to generate a comprehensive monitoring system; Based on the comprehensive monitoring system, the collected real-time behavior data of the target device is analyzed to generate a device behavior authentication result.
2. The method for monitoring the behavior of an Internet of Things device according to claim 1, characterized in that: According to the distribution of IoT devices, a discrete denoising diffusion model is trained at each node of a local IoT device, and the trained discrete denoising diffusion model is recorded as a pre-trained local model of the corresponding node, including: According to the distribution of IoT devices, the Director node and several Envoy nodes are configured through the OpenFL framework and a secure communication channel is established to obtain the constructed federated learning network topology structure; Based on the federated learning network topology, corresponding device operation data is collected at each Envoy node, and after preprocessing the device operation data, a standardized behavior data set is obtained; The architecture of the discrete denoising diffusion model is configured based on the behavior data set, and the discrete denoising diffusion model with the configured architecture is recorded as a first model; the architecture includes an encoder, a noise predictor and a decoder; Perform local training of forward diffusion and reverse denoising on the first model at the corresponding Envoy node, and record the trained first model as a pre-trained local model corresponding to the Envoy node.
3. The method for monitoring the behavior of an Internet of Things device according to claim 2, characterized in that: The extraction of the activation value of each node in the process of training the discrete denoising diffusion model, and obtaining all the trusted nodes according to the calculated activation value deviation of each node, include: Extract the activation values of neurons in each layer during the training of the discrete denoising diffusion model and generate a statistical feature set of activation values; Based on the activation value statistical feature set, the activation value distribution parameters of normal nodes are calculated to construct a standard activation mapping mode; Using the standard activation mapping mode, the activation value deviation of each node is calculated and anomaly detection is performed to generate a Byzantine node list; According to the Byzantine node list and filtering rules, the filtered trusted nodes are obtained.
4. The method for monitoring the behavior of an Internet of Things device according to claim 3, characterized in that: The extracting activation values of neurons in each layer during the training process based on the pre-trained local model to generate an activation value statistical feature set includes: In the process of training the discrete denoising diffusion model, the activation value of each neuron is collected in real time by adding observation points in each layer of the neural network; The activation value data of each neuron is collected layer by layer to form a multi-dimensional data set, wherein each data point in the data set represents the activation state of a neuron in a specific training step; Performing statistical analysis on the collected activation value data to calculate activation value statistical features, wherein the activation value statistical features include activation data mean, activation data variance, activation data skewness, and activation data kurtosis; The activation value statistical features are classified according to layers and neuron organizations to generate a multi-dimensional activation value statistical feature set.
5. The method for monitoring the behavior of an Internet of Things device according to claim 4, characterized in that: The method of using the standard activation mapping mode to calculate the activation value deviation of each node and perform anomaly detection to generate a Byzantine node list includes: Calculate the difference between the activation value distribution of each node and the standard activation mapping pattern; Based on the difference, the shape deviation degree of the activation value distribution and the change trend in the time series are obtained, and the corresponding nodes that meet any one of the following conditions, that is, the shape deviation degree exceeds a preset threshold and the change trend meets a preset condition, are marked as potential Byzantine nodes; Divide the entire time series into Q time windows; Calculate the shape deviation degree of the activation value distribution of each node in each time window and the change trend in the time series. If each node is marked as a potential Byzantine node P times in Q time windows, mark the node as a target Byzantine node; Q≥2; The IDs of the target Byzantine nodes are aggregated to generate a Byzantine node list.
6. The method for monitoring the behavior of an Internet of Things device according to claim 1, characterized in that: The calculating, based on the local data of the trusted node, the distribution characteristics of each node and the distribution similarity of data between nodes includes: Calculate the distribution characteristics of each of the trusted nodes according to the acquired local data of each of the trusted nodes, where the distribution characteristics include statistical characteristics and distribution shape characteristics; The statistical distance index is used to calculate the distribution similarity of data between nodes.
7. The method for monitoring the behavior of an Internet of Things device according to claim 1, characterized in that: The integrated monitoring system is used to analyze the collected real-time behavior data of the target device to generate a device behavior authentication result, including: Through the integrated monitoring system, real-time behavior data of the target device is continuously collected to form a behavior sequence sorted by time; Analyzing the behavior sequence based on the target model to generate a behavior evaluation result; According to the behavior evaluation result and the preset abnormal threshold, the device behavior authentication result is output.
8. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the Internet of Things device behavior monitoring method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the Internet of Things device behavior monitoring method described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.