A Threat Detection Method and System for Federated Learning Based on Modbus
By calculating the similarity relationships of important parameters of local model and Modbus hidden communication, a personalized federated learning model is generated, which solves the generalization of federated learning in industrial control networks and attacker concerns, and improves threat detection effect and collaboration capabilities.
Patent Information
- Application Number
- CN202411778751.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The application of existing federated learning in industrial control networks has problems such as poor generalization capabilities, lack of unified feature representation methods, and easy to attract attention from attackers, and it is difficult to effectively apply in fragmented industrial scenarios.
By calculating the similar relationships of important parameters of each local model, a personalized detection model is generated for each federated client, and multi-source data input is supported, while optimizing the personalized and global parts, the Modbus protocol is used for hidden communication.
It improves the effectiveness of cyber threat detection, enhances the collaboration capabilities between similar clients, has global scenario detection capabilities, and reduces attackers' attention to federated devices.
Smart Images

Figure CN119449467B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and in particular, to a threat detection method and system for federated learning based on Modbus. Background Art
[0002] As a common language in the field of industrial control, the Modbus protocol has been widely used in industries related to the national economic lifeline, such as energy, transportation, petrochemical, etc.
[0003] Federated learning is a distributed machine learning technology, and its core advantage is that it allows all participating parties to jointly train an ideal global model while protecting privacy. However, the fragmented industrial control scenarios pose great challenges to the direct application of federated learning. First of all, due to the diversity of network environments and the unpredictability of attacks, the data owned by federated learning clients is heterogeneous, which will lead to weight differences between local models and affect the performance of federated learning. Although there are some personalized federated learning methods claiming to be able to alleviate this problem, they all aim to narrow the deviation between the global model and the local model while ignoring the potential similarity relationships between clients. In addition, the goal of federated learning should balance global and local optimization. That is to say, clients should have the ability to detect local network threats and at the same time obtain the ability to detect unknown threats through federated learning.
[0004] Secondly, in order to easily quantify and control the degree of data imbalance, most federated threat detection methods allocate client data on a specific dataset. This means that all data actually comes from the same test platform. However, in practice, the data sources of network threats have various channels, which may come from the traffic of real industrial control systems or the traffic generated by security analysis tools related to the systems. Clients have the right to select appropriate data according to their own needs to build a detection model. To achieve this goal, a more efficient network traffic feature extraction method must be proposed.
[0005] Finally, the Modbus control network uses the Modbus protocol for data exchange between nodes, but currently the communication methods between the server and clients in most federated learning solutions are mainly based on traditional IT protocols. Although they apply various security encryption technologies to ensure privacy, they are also very likely to attract extra attention from attackers. Because attackers can clearly know that the devices involved in federated learning are performing a task outside of an industrial process. Therefore, this poses higher requirements for federated learning solutions in the Modbus control network, that is, the concealment of communication needs to be considered.
[0006] The paper "Personalized federated learning of execution&evaluation dual network for CPS intrusion detection." proposed an industrial federated threat detection scheme for an execution-evaluation dual network that can generate both global and local models simultaneously. The generation of its personalized model still depends on the global model created in the execution network. However, in the case of non-independent and identically distributed data, it is difficult for the execution network to generate a globally well-performing model for all clients. In addition, this method uses the cosine distance of the parameters of each local model to measure the similarity relationship between models. Due to the large scale of the parameters of the neural network, this approach will result in the curse of dimensionality problem.
[0007] The paper "An Effective Clustered Federated Learning Framework for Industrial Internet of Things Intrusion Detection" proposed an industrial network threat detection method based on clustered federated learning. This method finds federated clients with similar data distributions by performing temporal clustering on model evaluation metrics and collaboratively trains a common network threat detection model for them. Using the method of temporal clustering to mine the similarity relationship between clients requires the prerequisite that there are natural groupings among the clients. However, in the actual process, the data distributions among clients may be different, which will lead to the failure of this method. In addition, this method aims to improve the performance of local models on local data while ignoring the learning of global information, and the robustness of the model is poor.
[0008] In summary, the existing federated learning network threat detection methods applied to industrial control networks have problems such as poor model generalization ability, lack of a unified feature representation method, and being easily noticed by attackers, making it difficult for them to be applied in fragmented industrial scenarios, and their practicality and reference value are not high. Summary of the Invention
[0009] To solve the above-mentioned problems, the present invention provides an intelligent auction method and system. By calculating the similarity relationship of the important parameters of each local model during the federated training process, a personalized detection model is generated for each federated client, and federated learning is allowed to optimize both the personalized and global parts simultaneously. In addition, this method can support multi-source data input and ensure the concealment of federated communication.
[0010] In a first aspect, a Modbus-based federated learning threat detection method provided by the present invention adopts the following technical solution:
[0011] A threat detection method for federated learning based on Modbus, including:
[0012] Obtain binary network traffic data;
[0013] Preprocess the obtained binary network traffic data to obtain Json file data;
[0014] Extract features from the Json file data;
[0015] Initialize the federated learning model, and use the extracted features for model training to obtain several important parameters of the model;
[0016] By summarizing several important parameters of the model, obtain the global important parameter positions;
[0017] Update the model based on the global important parameter positions to obtain a personalized federated learning model;
[0018] Use the personalized federated learning model for network threat detection.
[0019] Furthermore, the obtaining of the binary network traffic data includes obtaining five data sets: industrial control traffic data, low-interaction industrial honeypot data, high-interaction industrial honeypot data, attack tool data, and mixed data. Among them, set the number of clients N = 20, the number of training rounds T = 150, the selection ratio of important parameters K% = 5%, and the balance factor , learning rate , 80% of the data in each client is used for training, and 20% of the data is used for local model evaluation.
[0020] Furthermore, the preprocessing of the obtained binary network traffic data to obtain Json file data includes preprocessing the binary network traffic into Json file data using the Tshark API, and dividing the Json file data into three groups: Session, Flow, and Packet according to different feature extraction granularities.
[0021] Furthermore, the feature extraction from the Json file data includes basic feature extraction and Modbus protocol feature extraction. Among them, basic feature extraction includes extracting features from the IT level, including quantity features, time features, byte features, protocol features, and port number features; Modbus protocol feature extraction includes extracting features related to Modbus data packets, including whether the request is responded, Trans-ID, Unit-Id, function code, start address of read / write operation, and address length of read / write operation.
[0022] Further, training the model using the extracted features to obtain several important parameters of the model, including training a federated learning model based on local data to obtain the trained model parameters, calculating the cumulative contribution of each model parameter using a formula to obtain the importance of each parameter, selecting several parameters as the important model parameters of the local model according to the importance, and marking the positions of all important parameters.
[0023] Further, summarizing several important parameters of the model to obtain the global important parameter positions, including extracting the marked important model parameters from the model, and calculating the correlation coefficient matrix of each model according to the important model parameters, expressed as:
[0024] (4.2)
[0025] where the correlation coefficient matrix ; N is the number of clients, represents the similarity coefficient vector of the important model parameters of client i with other important model parameters, represents the important model parameters of client i and the important model parameters of client j The similarity coefficient, softmax() is a normalization function that ensures the sum of each row in matrix A is 1, represents the important model parameters of client i and the important model parameters of client j The set of all joint probability distribution sets, is the probability that x appears in and y appears in The probability. represents the expected distance of all x and y, and its minimum value is and The EMD distance, that is , stores The EMD distance from and other important model parameters, and the function max() is the maximum value function.
[0026] Further, updating the model based on the global important parameter positions to obtain a personalized model, including updating the model through a correlation coefficient formula to form a personalized model for each client, and the correlation coefficient formula is expressed as:
[0027]
[0028] is the correlation coefficient between client i and other clients, is the local model of client i, = 0.7 is the balance factor.
[0029] In a second aspect, a Modbus-based federated learning threat detection system includes:
[0030] A data acquisition module configured to acquire binary network traffic data;
[0031] A preprocessing module configured to preprocess the acquired binary network traffic data to obtain Json file data;
[0032] A feature extraction module configured to extract features from the Json file data;
[0033] A model training module configured to initialize the federated learning model and use the extracted features to train the model to obtain several important parameters of the model;
[0034] A parameter module configured to summarize several important parameters of the model to obtain the global important parameter positions;
[0035] An update module configured to update the model based on the global important parameter positions to obtain a personalized federated learning model;
[0036] A detection module configured to perform network threat detection using the personalized federated learning model.
[0037] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the Modbus-based federated learning threat detection method.
[0038] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the Modbus-based federated learning threat detection method.
[0039] In summary, the present invention has the following beneficial technical effects:
[0040] (1) This application proposes a multi-granularity-based network traffic representation method, which can characterize network traffic from multiple dimensions. While providing support for multi-source input in the federated learning framework, it improves the network threat detection effect at the data level. (2) This application proposes a similarity update algorithm based on important model parameters, which enhances the cooperation effect between similar clients to better adapt to local application scenarios. At the same time, the personalized model can also integrate global knowledge, enabling the model to have the global scenario detection ability. (3) This application proposes a covert federated communication scheme for Modbus control networks. This scheme uses the standard Modbus protocol to realize the information exchange between the server and the client, making the devices performing federated learning behave more like real industrial control devices, thus masking the federated learning process and reducing the attacker's attention to federated devices. Description of the Drawings
[0041] Figure 1 It is a schematic diagram of the process of the covert personalized federated learning threat detection method in Embodiment 1 of the present invention;
[0042] Figure 2 It is a flowchart of the network traffic feature extraction method in Embodiment 1 of the present invention;
[0043] Figure 3 It is a flowchart of the client training in Embodiment 1 of the present invention; Figure 4 It is an F1 score graph of different feature extraction methods on each dataset in Embodiment 1 of the present invention;
[0044] Figure 5 It is a graph of the local and global test accuracy effect of Scenario 1 - Dirichlet~(0.3) in Embodiment 1 of the present invention;
[0045] Figure 6 It is a graph of the local and global test accuracy effect of Scenario 2 - Dirichlet~(0.3) in Embodiment 1 of the present invention;
[0046] Figure 7 It is a graph of the time overhead t of different methods under Scenario 2 - Dirichlet~(0.3) in Embodiment 1 of the present invention;
[0047] Figure 8 It is a Modbus covert communication deployment diagram in Embodiment 1 of the present invention;
[0048] Figure 9 It is a graph of the proportion of the session length between different protocol clients and attackers in Embodiment 1 of the present invention;
[0049] Figure 10 It is a system structure diagram in Embodiment 1 of the present invention. Detailed Implementation Modes
[0050] The present invention will be further described in detail below with reference to the accompanying drawings.
[0051] Embodiment 1
[0052] Referring to Figure 1 , a Modbus-based federated learning threat detection method in this embodiment includes:
[0053] Obtaining binary network traffic data;
[0054] Preprocessing the obtained binary network traffic data to obtain Json file data;
[0055] Performing feature extraction on the Json file data;
[0056] Initializing the federated learning model, and using the extracted features for model training to obtain several important parameters of the model;
[0057] By summarizing several important parameters of the model, obtaining the global important parameter positions;
[0058] Updating the model based on the global important parameter positions to obtain a personalized federated learning model;
[0059] Using the personalized federated learning model for network threat detection.
[0060] Specifically,
[0061] This embodiment uses five data sets, namely industrial control traffic data, low-interaction industrial honeypot data, high-interaction industrial honeypot data, attack tool data, and mixed data, to comprehensively evaluate the method proposed in the present invention. Among them, the industrial control traffic data is network traffic data collected in a small-scale power simulation system. In addition to normal system polling and manual operation traffic, the traffic also includes 6 types of attack traffic. The low-interaction industrial honeypot data is the traffic collected by a low-interaction Modbus honeypot from 10 different attack organizations from July 2017 to February 2023. Traffic from the same organization is considered to belong to the same type of attack. The high-interaction industrial honeypot data consists of 4 different types of attack traffic collected by a high-interaction Modbus honeypot. The attack tool data uses data collected by scanning 10 real Modbus industrial control devices exposed to the Internet with 5 public attack tools. The mixed data set is composed of the fusion of the above four data sets. The specific distribution of the data sets is shown in Table 1, and the total number of data packets represents the total amount of normal and attack packets. The number of clients in this embodiment is N = 20, the number of training rounds is T = 150, the selection ratio of important parameters is K% = 5%, the balance factor , learning rate 80% of the data in each client is used for training, and 20% of the data is used for local model evaluation.
[0062] Table 1
[0063]
[0064] A threat detection method for federated learning based on Modbus in this embodiment includes the following steps:
[0065] Step 1: The client executes a network traffic feature extraction method to convert network traffic into structured data. As shown in Figure 2 the specific steps are as follows:
[0066] (1) Use the Tshark API to preprocess binary network traffic into Json file data.
[0067] (2) The shunt module divides the Json file data into three groups: Session, Flow, and Packet according to different feature extraction granularities. Among them, Session is determined by consecutive IP data packets in the bidirectional triple {source IP, destination IP, transport layer protocol} within 10 seconds, which provides a comprehensive view of the network behavior between two IP addresses within a period of time. Flow is uniquely identified by the bidirectional quintuple {source IP, destination IP, source port, destination port, transport layer protocol} within 10 seconds, which provides key statistical features of network node interaction under a specific port. Packet consists of a request-response pair carrying a valid Modbus protocol payload, which can reflect the changes in the current physical process of the system.
[0068] (3) The basic feature extraction module is mainly responsible for extracting features from the IT level, including quantity features, time features, byte features, protocol features, and port number features, as shown in rows 2-6 of Table 2, a total of 16 features.
[0069] (4) The Modbus protocol feature extraction module is responsible for extracting features related to Modbus data packets, including whether the request is responded, Trans-ID, Unit-Id, function code, start address of read / write operation, and address length of read / write operation, as shown in row 7 of Table 2, a total of 6 features.
[0070] (5) Loop and execute (3)-(4) until all groups are processed.
[0071] Table 2
[0072] Feature type Feature description Quantity Time The durations of Session and Flow, the average interval between packet arrivals in Flow, and the time interval between request and response packets. 4 Quantity The number of Flows in Session, the number of industrial protocol packets in Flow, and the total number of packets in Flow 3 Protocol Transport layer and application layer protocols 2 Byte The number of bytes in Flow, the average byte size of packets in Flow, the transmission rate of Flow, and the number of bytes in request and response packets 5 Port Source port and destination port 2 Modbus Whether the request packet is responded, Trans-ID, Unit-Id, function code, start address of read / write operation, and address length of read / write operation 6
[0073] Step 2: Initialize the federated learning system. The specific steps are as follows:
[0074] (1) The client counts each feature range , where represents all feature ranges of client i, represents the range of the nth feature of client i, represents the minimum value of the nth feature of client i, represents the maximum value of the nth feature of client i.
[0075] (2) The server requests each feature range from the client through the read input register (0x04) function code of the Modbus protocol.
[0076] (3) The client encapsulates each feature range into the Modbus protocol through the Libmodbus API and sends it to the server.
[0077] (4) The server counts all feature ranges of each client and sends the global feature ranges of each feature , , , to the client through the write holding register (0x10) function code of the Modbus protocol. Among them, represents the global range of all features, represents the global range of the nth feature, represents the global minimum value of the nth feature, represents the global maximum value of the nth feature, min{} is the minimum value function, and max{} is the maximum value function.
[0078] (5) After receiving the global feature range, the client returns an acknowledgement data packet.
[0079] (6) The server sends the model structure and hyperparameters required for federated learning training through the write holding register (0x10) function code of the Modbus protocol. The model structure includes the number of neural network layers 5 and the number of neurons in each layer (22, 64, 128, 64, 26). The hyperparameters required for training include the number of federated training rounds , the learning rate , and the selection ratio of important parameters .
[0080] (7) The server sends the initial model parameters through the write holding register (0x10) function code of the Modbus protocol . Since the maximum number of registers that Modbus allows to operate in one request is 123, in this invention, when transmitting the model parameters each time, the parameters of each layer of the neural network are flattened into a one-dimensional vector and sliced according to formula (2.1).
[0081] (2.1)
[0082] where is the number of parameters of the j-th layer neural network, is the number of Modbus data packets required to transmit the parameters of the j-th layer neural network. The reason for multiplying by 2 is that for each model parameter in floating-point format, Modbus needs to store it in two registers. Therefore, to transmit all model parameters, Modbus data packets are required. In addition, during the parameter transmission process, the present invention sequentially increments and sends the Trans-ID field used to identify the Modbus request-response pair, so that the client can clearly understand the position of the parameters to be received or uploaded, and use it as the basis for retransmitting lost packets.
[0083] (7) After receiving the model parameters, the client returns an acknowledgment data packet.
[0084] (8) The server continuously monitors the client training status through the Modbus read discrete output register (0x01) function code.
[0085] Step 3: The client performs model training. The client uses the model sent by the server to train the model on local data and upload the trained model and the positions of important parameters to the server. Refer to Figure 3 , and its specific steps are as follows:
[0086] (1) The client uses formula (3.1) to train the model on local data and obtains the trained model parameters:
[0087] (3.1)
[0088] where represents the model parameters received by the i-th client in the t-th round of federated training, and the function represents performing neural network loss minimization training, = 0.001 represents the learning rate.
[0089] (2) Use formula (3.2) to calculate the cumulative contribution of each model parameter:
[0090] (3.2)
[0091] where represents the cumulative contribution of the m-th parameter of the client i model in the first t rounds, represents the gradient of the m-th parameter in the t-th round of training, is the change of the m-th parameter in the t-th round. represents the contribution of the m-th parameter in the t-th round.
[0092] (3)Select parameters, and calculate the importance of each parameter, as shown in formula (3.3).
[0093] (3.3)
[0094] where is the cumulative change of the m-th parameter of the model. According to the above formula, the present invention can quantitatively evaluate the importance of all model parameters, and the importance of all local model parameters can be expressed as , M = 19456 is the parameter scale of the neural network.
[0095] (4)Select the top K% of the maximum values from as the important model parameters of the i-th local model, and mark their positions through formula (3.4). The important model parameters will be marked as 1, and the unimportant model parameters will be marked as 0, forming an important parameter position vector :
[0096] (3.4)
[0097] (5)The client listens to step (8) of step two to generate an acknowledgment packet to identify the end of training.
[0098] (6)The server requests the local model parameters from the client through the read input register (0x04) function code of the Modbus protocol.
[0099] (7)The client slices and uploads the model parameters to the server according to formula (2.1).
[0100] (8)The server requests the positions of the important parameters from the client through the read input register (0x04) function code of the Modbus protocol.
[0101] (9)The client uploads the important parameter position vector .
[0102] Step Four: Server model update, the steps include
[0103] (1)The server aggregates the positions of the important parameters of all models to obtain the global important parameter position , where each element can be calculated by (4.1)
[0104] (4.1)
[0105] where represents the sum of the position vectors of the m-th model parameter.
[0106] (2) The server extracts those model parameters marked as important from the model , and calculates the correlation coefficient matrix of each model in through formula (4.2). .
[0107] (4.2)
[0108] where N is the number of clients, represents the similarity coefficient vector of the important model parameters of client i with other important model parameters, represents the important model parameters of client i and the important model parameters of client j The similarity coefficient, softmax() is a normalization function, which ensures that the sum of each row in matrix A is 1. represents the important model parameters of client i and the important model parameters of client j The set of all joint probability distribution sets, is the probability that x appears in and y appears in . represents the expected distance of all x and y, and its minimum value is and The EMD distance, that is, . stores the and the EMD distance of other important model parameters, and the function max() is the maximum value function.
[0109] (3) The server updates the model through formula (4.3) to form a personalized model for each client
[0110] (4.3)
[0111] is the correlation coefficient between client i and other clients, and its value is obtained in the previous step, is the local model of client i, = 0.7 is the balance factor.
[0112] (4) The server slices the model parameters according to formula (2.1), and issues the model parameters through the write holding register (0x10) function code of the Modbus protocol, and the training round t = t + 1.
[0113] Step 5: Repeat Step 3 and Step 4 until the maximum number of training rounds T = 150 is reached, and end the federated learning training.
[0114] Experimental verification
[0115] To verify the effectiveness of the method proposed in the present invention, we first evaluated the network traffic feature representation method proposed in the present invention and the commonly used CICFlowMeter and Honeyeye feature representation methods in terms of performance and detection effect respectively. Table 3 and Figure 4 respectively give the feature extraction rate results of different network traffic feature representation methods and the F1 scores on different datasets.
[0116] As shown in Table 3, CICFlowMeter statistically analyzes the features of a large number of TCP connection packets during network traffic characterization, which makes the time cost consumed by this method relatively high. In the case of a small amount of data, Honeyeye shows good performance, mainly because this method only extracts features from packets containing application layer protocols, and packets without application layer protocols are directly discarded. However, as the amount of data increases, this method will cause a large amount of memory space to be occupied, resulting in a significant decrease in its processing speed. In contrast, the method proposed in the present invention can centrally process the packets containing application layer protocols in each data stream, so it can still maintain good performance in a large data volume environment.
[0117] Table 3
[0118] Data volume CICFlowMeter (ms) Honeyeye (ms) The method of the present invention (ms) 432 442 9 75 4664 542 120 254 32299 667 883 3439 200300 5963 7896 4276
[0119] As Figure 4 shown, although the three feature extraction methods all achieved F1 scores above 0.75 on different datasets, the method of the present invention and the Honeyeye method achieved higher performance because they considered the features of the industrial protocol part during feature extraction. In addition, the present invention simultaneously considers the features at three levels of Session, Flow, and Packet, making its description of network behavior more comprehensive and more sensitive to the traffic differences between different types, thus achieving better detection effects.
[0120] Secondly, to verify the effectiveness of the federated learning system, the present invention set up two experimental scenarios to conduct a comparative experiment on the proposed federated learning method and the commonly used FedAVG, FedProx, Ditto, IFCA, CFL. The two experimental scenarios are described as follows:
[0121] (1)Scenario 1: In Scenario 1, each client only relies on the data provided by a specific functional entity for federated learning. Based on this assumption, 20 clients were created in this experiment and divided into 4 groups, with each group of nodes only using one of the four datasets. Therefore, the data among the 4 client groups presents the characteristic of non-independent and identically distributed. In addition, to simulate a more complex data distribution situation, 5 clients in the same group will be divided according to the Dirichlet ( distribution for non-independent and identically distributed partitioning, where is the concentration parameter, and the smaller its value, the more extreme the non-independent and identically distributed situation represents.
[0122] (2)Scenario 2: Scenario 2 refers to the situation where each client uses the data of multiple functional entities for federated learning training when available. At this time, each client will have a mixture of data from the four datasets at the same time. In this scenario, we also set 20 clients and simulated non-independent and identically distributed data partitioning according to distribution, and the data between clients does not overlap.
[0123] For each type of experimental scenario, the present invention respectively performs local and global evaluations on the model. In the local evaluation, the client will evaluate the model on the local data. In Scenario 1, the global evaluation refers to the client testing the model based on the shared data within the same group (one of the four datasets); while in Scenario 2, the global evaluation is to verify the model on the mixed dataset. This means that before data partitioning, 20% of the data from all five datasets will first be extracted as global evaluation data, and the remaining 80% of the data will be allocated to each client according to the data partitioning methods of the above different scenarios. For both scenarios, this experiment will the parameters in be set to 0.3 and 0.6 respectively to simulate different non-independent and identically distributed data partitioning situations.
[0124] The experiment respectively counted the accuracy and F1 score of each method in the global and local evaluations of each scenario, and used the average evaluation situation of the global and local to measure the overall performance of the model. The results are shown in Table 4. Figure 5 and Figure 6 respectively give the accuracy change trends of each method in Scenario 1 - Dirichlet~(0.3) and Scenario 2 - Dirichlet~(0.3). At the same time, we also recorded the additional computational overhead generated by each method in each round of communication in Scenario 2 - Dirichlet~(0.3) due to additional operations (compared with the traditional FedAVG method), as Figure 7 shown.
[0125] Table 4
[0126]
[0127] Since FedAVG does not consider the data heterogeneity problem, its global evaluation and local evaluation have similar performances, both of which are at the lowest level. The method proposed in the present invention achieves the optimal effect on the average index by enhancing the cooperation between similar nodes on the basis of integrating global information. Among them, the average accuracy is improved by 0.011, 0.028, 0.005 and 0.006 respectively under four different data partitioning cases, and this achievement is improved by 0.025, 0.031, 0.003 and 0.008 on the F1 score. The one that is closest to the performance of the present invention is Ditto, which shows better fairness and achieves good results in global evaluation. However, the two-stage optimization method requires the edge nodes to be trained twice in each round of training, which makes the time overhead per round 40 seconds higher than that of the method of the present invention. FedProx does not show obvious advantages in various indicators, which indicates that it is very difficult to train an excellent global model in the case of non-independent and identically distributed data. Combining Figure 5 and Figure 6 It can be seen that the method of the present invention has a faster convergence speed, which is beneficial to reducing the communication overhead of federated learning. In contrast, FedAVG, FedProx and IFCA have greater volatility.
[0128] As Figure 7 shown, in terms of time overhead, the present invention effectively reduces the computational overhead of client similarity through a personalized model update algorithm for an important parameter. In addition, the gradient information of the model in the present invention can be automatically saved during the local model training process, so it does not bring too much time overhead and achieves good results.
[0129] Finally, the concealment of the Modbus communication scheme was evaluated by the degree of attraction to external attackers. The experimental evaluation was carried out on a physical server in an Internet Data Center (IDC) equipped with 10 public IP addresses. The server was equipped with two E5-2698 V4 CPUs and had 128GB of memory. The 10 IP addresses were in the same C-class network segment. During the experiment, Modbus high-interaction honeypots were deployed on 6 of the IP addresses to simulate the behavior of real industrial control devices. The client program of the present invention was deployed on the remaining 4 IPs. Among them, 2 IPs communicated with the server using the Modbus communication scheme proposed by the present invention, and the remaining 2 IPs communicated with the federation server using the commonly used HTTP protocol in federated communication. The server program of the present invention was deployed on two other physical servers and communicated with the corresponding edge nodes using the Modbus and HTTP protocols respectively. At the same time, the two servers also acted as remote master devices and accessed the honeypot nodes in a Modbus polling manner. The overall deployment scheme is as Figure 8 shown.
[0130] For a device running in an industrial control system, the degree of attraction to external attackers was determined by calculating the number of communications between the attacker and the device within 10 seconds, i.e., SessionLength. The larger the SessionLength, the higher the degree of attraction of the device to the attacker, which also means that the attacker is willing to invest more effort to obtain the target information. In the first month of the experiment, only the two IP addresses using the HTTP protocol were enabled. In the second month of the experiment, all 10 IP addresses were enabled, and the Wireshark tool was used to capture the attack traffic. The average proportion of SessionLength for each entity is as Figure 9 shown. The label "HTTP-First month" in the figure represents the session distribution of the two IP addresses using the HTTP protocol for federated communication in the first month.
[0131] For the nodes using Modbus for federated communication, the performance of SessionLength was extremely similar to that of the honeypot. Especially and The proportions only differ by 3.2% and 0.5%, indicating that attackers do not show extra attention to nodes using Modbus for federated communication. Although the SessionLength of nodes using HTTP for federated communication in the second month also concentrated in 1 - 2 rounds, given that most of the attack behaviors captured during the experiment were of a scanning nature, this finding is normal. In addition, by comparing the session statistics of HTTP federated communication nodes in the first and second months, it can be found that there is an obvious increase in SessionLength in the data captured in the second month, especially The session of
[0132] Example 2
[0133] This example provides a Modbus - based federated learning threat detection system, including:
[0134] A data acquisition module, configured to acquire binary network traffic data;
[0135] A pre - processing module, configured to pre - process the acquired binary network traffic data to obtain Json file data;
[0136] A feature extraction module, configured to extract features from the Json file data;
[0137] A model training module, configured to initialize the federated learning model and use the extracted features for model training to obtain several important parameters of the model;
[0138] A parameter module, configured to summarize several important parameters of the model to obtain the global important parameter positions;
[0139] An update module, configured to update the model based on the global important parameter positions to obtain a personalized federated learning model;
[0140] A detection module, configured to perform network threat detection using the personalized federated learning model.
[0141] As a further implementation method,
[0142] The system structure of the present invention is as Figure 10 shown, and it mainly consists of a client and a server.
[0143] Server: The server is responsible for aggregating the uploaded local models in a round of federated training. First, the server calculates the similarity relationships among clients using the important model parameters of all local models, and creates personalized detection models for each client by combining the global average model during the aggregation process. , and then covertly transmits it to each client through an industrial protocol. Multiple rounds of interaction are required between the server and the clients to form the final personalized threat detection model.
[0144] Client: The client first converts network traffic into structured data using a feature extraction method. Secondly, the client trains the personalized detection model sent by the server on local data , to form a local model . Subsequently, the client filters out the most important K% of the model parameters in the local model . Finally, the client covertly transmits the local model and the positions of the important model parameters to the server through the Modbus protocol.
[0145] A computer-readable storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device for a Modbus-based federated learning threat detection method.
[0146] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for a Modbus-based federated learning threat detection method.
[0147] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A threat detection method for federated learning based on Modbus, characterized in that, Including: Obtain binary network traffic data; Preprocess the obtained binary network traffic data to obtain Json file data; Extract features from the Json file data; Initialize the federated learning model, and use the extracted features for model training to obtain several important parameters of the model; Summarize several important parameters of the model to obtain the global important parameter positions; Update the model based on the global important parameter positions to obtain a personalized federated learning model; Use the personalized federated learning model for network threat detection; The obtaining of binary network traffic data includes obtaining five data sets: industrial control traffic data, low-interaction industrial honeypot data, high-interaction industrial honeypot data, attack tool data, and mixed data. Among them, the number of clients is set to N = 20, the number of training rounds is T = 150, the selection ratio of important parameters is K% = 5%, and the balance factor , the learning rate , 80% of the data in each client is used for training, and 20% of the data is used for local model evaluation; The extracting features from the Json file data includes basic feature extraction and Modbus protocol feature extraction. Among them, basic feature extraction includes extracting features from the IT level, including quantity features, time features, byte features, protocol features, and port number features; Modbus protocol feature extraction includes extracting features related to Modbus data packets, including whether the request is responded, Trans-ID, Unit-Id, function code, start address of read / write operation, and address length of read / write operation; The using the extracted features for model training to obtain several important parameters of the model includes training the federated learning model based on local data to obtain the trained model parameters, calculating the cumulative contribution of each model parameter using the formula to obtain the importance of each parameter, selecting several parameters as the important model parameters of the local model according to the importance, and marking all important parameter positions.
2. The method for detecting threats in federated learning based on Modbus according to claim 1, characterized in that The preprocessing the obtained binary network traffic data to obtain Json file data includes preprocessing the binary network traffic into Json file data using the Tshark API, and dividing the Json file data into three groups: Session, Flow, and Packet according to different feature extraction granularities.
3. The method for detecting threats in federated learning based on Modbus according to claim 2, characterized in that, The obtaining the global important parameter positions by summarizing several important parameters of the model includes extracting the marked important model parameters from the model, and calculating the correlation coefficient matrix of each model according to the important model parameters, expressed as: , Among them, the correlation coefficient matrix ; N is the number of clients, represents the similarity coefficient vector between the important model parameters of client i and other important model parameters, represents the important model parameters of client i and the important model parameters of client j The similarity coefficient, softmax() is a normalization function, which ensures that the sum of each row in matrix A is 1, represents the important model parameters of client i and the important model parameters of client j The set of all joint probability distribution sets, is the probability that x appears in and y appears in The probability, represents the expected distance of all x and y, and its minimum value is and The EMD distance, that is , stores the The EMD distance between and other important model parameters, and the function max() is a maximum value function.
4. A Modbus-based federated learning threat detection method according to claim 3, characterized in that The updating the model based on the global important parameter positions to obtain a personalized model includes updating the model based on the correlation coefficient formula to form a personalized model for each client, and the correlation coefficient formula is expressed as: , is the correlation coefficient between client i and other clients, is the local model of client i, = 0.7 is the balance factor.
5. A Modbus-based federated learning threat detection system that executes a Modbus-based federated learning threat detection method as described in claim 1, characterized in that, Including: Data acquisition module, configured to obtain binary network traffic data; Preprocessing module, configured to preprocess the obtained binary network traffic data to obtain Json file data; Feature extraction module, configured to extract features from the Json file data; Model training module, configured to initialize the federated learning model and use the extracted features for model training to obtain several important parameters of the model; Parameter module, configured to summarize several important parameters of the model to obtain the global important parameter positions; Update module, configured to update the model based on the global important parameter positions to obtain a personalized federated learning model; Detection module, configured to use the personalized federated learning model for network threat detection.
6. A computer-readable storage medium storing multiple instructions, characterized in that, The instructions are adapted to be loaded and executed by a processor of a terminal device to perform the method according to claim 1.
7. A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement each instruction; the computer-readable storage medium being configured to store a plurality of instructions, characterized in that, The instructions are adapted to be loaded and executed by a processor to perform the method according to claim 1.
Citation Information
Patent Citations
Industrial equipment fault detection method based on graph neural network federated learning
CN115311205A
Federal learning method and device for industrial Internet of Things intrusion detection
CN117459299A