An adaptive asynchronous federated learning intrusion detection method and device

By employing an adaptive asynchronous federated learning method, combined with cosine similarity penalty terms and dynamic gradient optimization penalty terms, the model inconsistency problem caused by dynamic resource changes and device heterogeneity in edge computing networks is solved. This enables rapid response to intrusion detection capabilities for new types of attacks and improves network threat defense capabilities.

CN120750675BActive Publication Date: 2025-11-04JIANGXI NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511262904.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-04
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing federated learning algorithms struggle to cope with the highly dynamic changes in resources, device heterogeneity, and unbalanced data distribution in edge computing networks. This leads to time waste in synchronous aggregation algorithms and inconsistencies in asynchronous update algorithm models, making it impossible to quickly update and iterate network intrusion detection models and to quickly and adaptively identify new types of attacks.

Method used

An adaptive asynchronous federated learning method is adopted, which enables real-time model updates and adaptive weight adjustment by co-training cloud servers and edge servers, combining cosine similarity penalty term and dynamic gradient optimization penalty term. The model parameters are aggregated based on comprehensive weights.

Benefits of technology

It enhances the robustness and flexibility of the federated learning system, improves its defense capabilities against complex attack scenarios, enables rapid response to new attacks and cyber threats, and ensures the consistency and accuracy of the global model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750675B_ABST
    Figure CN120750675B_ABST
Patent Text Reader

Abstract

The application provides a kind of adaptive asynchronous federal learning intrusion detection method and device, it is related to federal learning intrusion detection field, cloud server first trains initial cloud server model and issues model parameter to edge server, edge server trains model using local data, introduce cosine similarity and dynamic gradient optimization penalty term dynamically adjust the update direction of edge server model, reduce the gradient conflict between local and global model, then upload information to cloud server, cloud server is based on the information uploaded by edge server Comprehensive weight is calculated, based on adaptive weight mechanism, according to data distribution and classification accuracy dynamically adjust the contribution of edge server to global model, aggregation generates new cloud server model and issues new model parameter, cycle to cloud server model is stable and issues cloud server model to carry out intrusion detection.The application can enhance the robustness and flexibility of the system, effectively deal with new attacks, and improve the network defense capability in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of federated learning intrusion detection, and particularly relates to a self-adaptive asynchronous federated learning intrusion detection method and device. BACKGROUND

[0002] Federated learning (FL) is a decentralized machine learning paradigm that enables distributed collaboration to complete model training without directly accessing local data of participants. The core idea is to sink the model training process to the data end, and to achieve cloud server model optimization under the premise of protecting user privacy and data security.

[0003] However, due to the high dynamic changes of resources in edge computing networks, significant device heterogeneity, and unbalanced data distribution, existing federated learning algorithms are difficult to be directly applied to such scenarios. First, the synchronous aggregation algorithm requires waiting for all clients to complete local training for each round of global model update, which can easily lead to waste of time and computing resources in a high dynamic environment, limiting its practicality. Second, although the asynchronous update algorithm can partially alleviate the waiting problem, independent updates by clients can easily lead to inconsistency in global model updates, affecting convergence speed and stability. In addition, existing aggregation algorithms usually assign weights based on the proportion of client data volume, but in scenarios where data quality, value difference, and client data distribution are extremely non-independent and identically distributed, this strategy may not be applicable.

[0004] Publication No. CN114640498A, entitled "Network intrusion collaborative detection method based on federated learning", proposes to realize distributed collaborative training by constructing a "initiator-coordinator-participant" three-level architecture, to improve model convergence efficiency by combining incremental average aggregation algorithm, to use homomorphic encryption transmission and arbitration mechanism dual protection to protect data privacy, to introduce noise confusion in local model training to resist inversion attacks, and to support personalized business layer customization, effectively solving the model leakage risk and collaboration efficiency problems in traditional solutions while ensuring detection accuracy. However, the synchronous aggregation algorithm has a slow update frequency and cannot quickly update and iterate the network intrusion detection model, and the security needs to be improved.

[0005] CN114785608A, an industrial control network intrusion detection method based on decentralized federated learning, proposes a residual convolutional neural network combined with online difficult example mining technology, which uses multi-party industrial control network flow data to collaboratively train the model, and uses a model parameter federated update mechanism under a decentralized architecture to effectively solve the problems of data acquisition difficulty and class imbalance. This method optimizes sample distribution through data preprocessing, strengthens key sample learning through OHEM, and breaks down data silos through decentralized federation cooperation. Without revealing privacy, it significantly improves detection efficiency and accuracy, especially for real-time intrusion detection needs in low-resource, high-heterogeneous environments in industrial scenarios. However, this method does not consider increasing the weight of newly emerging attacks, and cannot quickly and adaptively improve the recognition ability of newly emerging attacks. SUMMARY

[0006] To solve the above technical problems, one technical solution adopted by the present application is to provide an adaptive asynchronous federated learning intrusion detection method, which comprises:

[0007] S10. The cloud server trains an initial cloud server model, and the model parameters of the trained initial cloud server model are downloaded to each edge server connected to the cloud server at the current time;

[0008] S20. The edge server loads the model parameters downloaded by the cloud server as the initial parameters of the edge server model, and updates the edge server model through local training while calculating the gradient change of the model;

[0009] S30. Based on the cosine similarity penalty term and the dynamic gradient optimization penalty term, the edge server model is corrected to obtain a corrected edge server model, and the information of the corrected edge server model is uploaded to the cloud server for updating;

[0010] S40. When the cloud server receives the update information of the edge server, it calculates the comprehensive weight, aggregates the edge server model parameters and the current cloud server model parameters based on the comprehensive weight, and obtains new cloud server model parameters;

[0011] S50. If the cloud server model converges, the training is ended and S60 is turned, otherwise the new cloud server model parameters are downloaded to each edge server, and then S20 is turned;

[0012] S60. The converged cloud server model is downloaded to each edge server for intrusion detection.

[0013] Further, each edge server connected to the cloud server at the current time is represented as , represents the i-th edge server connected to the cloud server edge servers; each edge server has a local dataset, wherein represents the local dataset of the th edge server, each edge server is equipped with a communication base station, and the edge server interacts with the cloud server and the user terminal through the communication base station;

[0014] The local dataset is represented as:

[0015] ;

[0016] wherein, represents the input feature of the th sample, represents the label of , and the global dataset is represented as , and the global dataset is the union of the local datasets of the edge servers , wherein represents an index variable, represents the number of edge servers participating in federated learning.

[0017] Further, the gradient change of the calculation model includes:

[0018] calculating the gradient change of each layer between the th round of cloud server model and the th round of cloud server model , the gradient change is the subtraction of two SGD gradients;

[0019] The th round of cloud server model and the th round of cloud server model have been saved when the edge server receives the cloud server model parameters;

[0020] Each edge server maintains a version queue to store the model sequence issued by the cloud server in the receiving order , and strictly follows the first-in first-out (FIFO) principle for processing, and the version loading takes the current version to be processed from the head of the queue (initially ), and loads it as the initial parameters of the local model. After the edge server completes the processing of the current version , it is automatically removed from the queue , and starts processing the next version . Even if the cloud server has iterated to , the edge server must execute all intermediate version training in sequence and cannot skip any version;

[0021] Calculate the first Wheel cloud server model With the modified edge server model The gradient change between each layer is the difference between two SGD gradients.

[0022] Furthermore, the calculation formula for correcting the edge server model based on the cosine similarity penalty term and the dynamic gradient optimization penalty term is as follows:

[0023] ;

[0024] ;

[0025] in, Represents edge server The regularization objective function, This represents the gradient of the parameters of the regularization objective function. Indicates the first Wheel cloud server model With the modified edge server model Gradient change between For learning rate, and These represent the cosine similarity penalty coefficient and the dynamic gradient optimization penalty coefficient, respectively. This represents the cosine similarity penalty term. This represents the dynamic gradient optimization penalty term, with weights after the model parameters are updated. For simplicity, it will still be written as .

[0026] Furthermore, the cosine similarity penalty term , used to indicate the first Wheel cloud server model With the modified edge server model The similarity of gradients between different layers is calculated using the following formula:

[0027] ;

[0028] in, Indicates the first Wheel cloud server model With the modified edge server model No. Gradient changes in layers, Indicates the first Wheel edge server Model No. Layer model parameters, represents the first round cloud server model with the first round cloud server model the first layer gradient change, represents the first round cloud server model the first layer model parameter, represents the number of layers of the model.

[0029] Further, the dynamic gradient optimization penalty term , used to represent the consistency of the global gradient between the first round cloud server model and the modified edge server model , the calculation formula is as follows:

[0030] ;

[0031] wherein, represents the first round cloud server model and the modified edge server model between the gradient change, represents the first round cloud server model and the first round cloud server model global gradient change.

[0032] Further, the information of the modified edge server model includes: edge server model parameters, edge server model classification accuracy for each type of attack and the amount of data of each type of attack input during edge server local training.

[0033] Further, the integrated weight-based aggregation edge server model parameters and the current cloud server model parameters, the aggregation formula is as follows:

[0034] ;

[0035] ;

[0036] wherein, represents the current cloud server model parameter of the cloud server, represents the edge server the first round model parameters uploaded to the cloud server by the modified edge server model, denote the model parameters of the cloud server model after aggregation, denote the comprehensive weight calculated by the penalty weight of cosine similarity and the adaptive weight of classification accuracy of each type of attack, denote the penalty weight of cosine similarity, denote the adaptive weight of data volume and data importance of each round of edge server, denote the adjustment coefficient of the comprehensive weight for adjusting the weight preference, denote the update frequency fairness correction term.

[0037] Further, the penalty weight of cosine similarity , the calculation formula is as follows:

[0038] ;

[0039] ;

[0040] wherein, is the model parameter of the edge server The first round of the edge server model uploaded to the cloud server is constructed into a model The cosine similarity between the gradient of the current model of the cloud server and the gradient of the model of the first round of the cloud server, is the initial weight of the data volume of the edge server, denote the model parameter of the first round of the cloud server model The gradient change of the first layer of the modified model The first round of the edge server The model parameter of the first layer of the model of the edge server, denote the model of the cloud server running at the current time of the cloud server The gradient change of the first layer of the first round of the cloud server model The first layer of the model.

[0041] Further, the initial weight of the data volume of each edge server , used to represent the proportion of the data volume of each edge server, the calculation formula is as follows:

[0042] ;

[0043] wherein, represents the edge server data volume, represents the total data volume of all edge servers.

[0044] Further, the adaptive weight of each edge server in each round , the calculation formula is as follows:

[0045] ;

[0046] ;

[0047] wherein the edge server data category is represented as , represents the initial weight of the data volume of each edge server, represents a parameter that determines the influence degree of accuracy difference on weight adjustment, represents the edge server data volume belonging to data category , represents the total data volume of each edge server data category , represents a new data category set, represents the local accuracy of the edge server data category , indicates that when aggregated, the cloud server determines the importance degree of data category according to the classification accuracy uploaded by the edge server, when the average recognition accuracy of the edge server model for a new data category is less than 90%, the data category is called a new data category, which is added to the new data category set , increase its weight, improve the recognition ability of the model for the new data category, represents the current training round, represents an intermediate round of adjustment point greater than the current training round.

[0048] Further, the update frequency fairness correction term , the fairness correction term formula is as follows:

[0049] ;

[0050] wherein, represents the average update times of the edge server, the calculation formula is as follows:

[0051] ;

[0052] wherein, represents the current training round, represents the number of edge servers participating in federated learning, is the update frequency of the edge server.

[0053] The application also provides an adaptive asynchronous federated learning intrusion detection device, comprising:

[0054] A data collection and preprocessing module is configured to collect a known public DDoS dataset, and pre-process the data to ensure that the data format and quality meet the training requirements;

[0055] An initial model training module is configured to train an initial cloud server model by the cloud server, and to distribute model parameters of the initial cloud server model to each edge server;

[0056] An edge model training and correction module is configured to receive the model parameters of the cloud server model by the edge server, load the model parameters as initial parameters of a local model, perform local training to generate an updated model, calculate the gradient change of the model, correct the model trained by the edge server based on a cosine similarity penalty term and a dynamic gradient optimization penalty term, obtain a corrected edge server model, and upload information of the model to the cloud server for aggregation;

[0057] A model aggregation and distribution module is configured to use an asynchronous aggregation mechanism by the cloud server, calculate a comprehensive weight when the cloud server receives update information of the edge server, aggregate model parameters of the edge server and current model parameters of the cloud server based on the comprehensive weight to obtain new cloud server model parameters, keep the cloud server model unchanged if the cloud server does not receive the update information of the edge server, terminate the global training process of the federated learning if the cloud server model converges, and otherwise distribute the new cloud server model parameters to each edge server;

[0058] An intrusion detection module is configured to distribute the converged cloud server model to each edge server for intrusion detection, record detection results for further analysis and alarm processing, and feed back the detection results to the cloud server for further optimization of the model.

[0059] The application has the following advantages:

[0060] The method is an adaptive asynchronous federated learning intrusion detection method, which can cope with new attacks and changing network threats. The method adjusts the update direction of the edge server model dynamically, combines the cosine similarity and dynamic gradient optimization penalty term, reduces the gradient conflict between the local and global models, and ensures that the global model has a consistent update direction. When the edge server perceives abnormal conditions, the system can use local data to update the model in real time, and based on the adaptive weight mechanism, dynamically adjusts the contribution of the edge server to the global model according to the data distribution and classification accuracy. This innovative strategy enhances the robustness and flexibility of the federated learning system, significantly improving the defense capability against complex attack scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0061] Figure 1 is a flowchart of an adaptive asynchronous federated learning intrusion detection method provided by the present application.

[0062] Figure 2 is an edge federated computing architecture diagram of an adaptive asynchronous federated learning intrusion detection method provided by the present application.

[0063] Figure 3 is a model optimization direction correction schematic diagram of an adaptive asynchronous federated learning intrusion detection method provided by the present application.

[0064] Figure 4 is a structural diagram of an adaptive asynchronous federated learning intrusion detection device provided by the present application. DETAILED DESCRIPTION

[0065] The preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings, so that the advantages and features of the present application can be more easily understood by those skilled in the art, and the scope of protection of the present application can be more clearly defined.

[0066] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein; obviously, the examples in the specification are only a part of the embodiments of the present application, not all the embodiments.

[0067] Embodiment 1

[0068] Figure 1 is a flowchart of an adaptive asynchronous federated learning intrusion detection method provided by the present application, which comprises:

[0069] S10. The cloud server trains an initial cloud server model, and the model parameters of the trained initial cloud server model are distributed to each edge server connected to the cloud server at the current time;

[0070] S20. The edge server loads the model parameters issued by the cloud server as the initial parameters of the edge server model, updates the edge server model through local training, and calculates the gradient change of the model.

[0071] S30. The edge server model is corrected based on the cosine similarity penalty term and the dynamic gradient optimization penalty term to obtain the corrected edge server model. The information of the corrected edge server model is uploaded to the cloud server for updating.

[0072] S40. When the cloud server receives the update information from the edge server, it calculates the comprehensive weight, and aggregates the edge server model parameters and the current cloud server model parameters based on the comprehensive weight to obtain new cloud server model parameters;

[0073] S50. If the cloud server model converges, end the training and go to S60; otherwise, send the new cloud server model parameters to each edge server and then go to S20.

[0074] S60. Distribute the converged cloud server model to each edge server for intrusion detection.

[0075] In this embodiment, the cloud server first trains an initial cloud server model based on the publicly available CICDDoS2019 dataset, and then distributes the parameters of the initial cloud server model to each edge server. Then, in the federated learning phase... Wheel, edge server Receive cloud server model The model parameters are loaded as the initial parameters of the local model, and the model is trained locally to generate an updated model. Simultaneously, the gradient change of each layer is calculated. Then, the edge server model is corrected based on the cosine similarity penalty term and the dynamic gradient optimization penalty term to obtain the corrected edge server model. The information of the corrected edge server model is uploaded to the cloud server for updates. The cloud server uses an asynchronous aggregation mechanism. If the cloud server receives the updated information from the edge server, it calculates the comprehensive weight and aggregates the edge server model parameters with the current cloud server model parameters based on the comprehensive weight to obtain new cloud server model parameters. Otherwise, the cloud server model remains unchanged. If the cloud server parameters are no longer updated, the federated learning global training process is terminated. Otherwise, the new cloud server model parameters are distributed to each edge server. Finally, the distribution and uploading operations are repeated until the global model is no longer updated.

[0076] refer to Figure 2 The edge servers currently connected to the cloud server are represented as follows: , The first one indicates the connection to the cloud server edge servers; each edge server has a local dataset, wherein represents the local dataset of the i-th edge server, each edge server is equipped with a communication base station, and the edge server interacts with the cloud server and the user terminal through the communication base station;

[0077] The local dataset is represented as:

[0078]

[0079] wherein represents the input feature of the i-th sample, represents the label of , the global dataset is represented as , and the global dataset is the union of the datasets of the edge servers , wherein represents an index variable, represents the number of edge servers participating in federated learning.

[0080] In order to simulate the non-independent and identically distributed situation of data in the actual edge federated environment, the power-law distribution is used to generate the sample size ratio. There are clients, and the sample size allocation ratio obeying Zipf distribution is generated:

[0081]

[0082] wherein controls the degree of imbalance, represents the sample size ratio of the i-th edge server, represents an index, and the ratio is mapped to the actual sample size:

[0083]

[0084] wherein represents the total sample size, represents the actual sample size, represents the floor function, which avoids decimal samples, represents the number of samples of the i-th class, and the sample size of the last client needs to be adjusted to ensure , after the global data is randomly shuffled during data allocation, the sample size is allocated to each client, and the client dataset is:

[0085]

[0086] ​​​​​​​​Due to the complex and diverse feature set of network traffic data, and considering the non-sequential nature of data features after removing network suite characters (i.e. the features of network packets do not have obvious time dependence), the embodiment selects a convolutional neural network (CNN) as the core detection model. CNN has attracted much attention due to its ability to automatically learn the hierarchical structure of the feature space, especially when dealing with structured data, it has shown strong advantages. Its multi-layer convolution operation can capture local and global patterns in network traffic features, and has significant effect on multi-classification tasks. Table 1 gives the specific architecture design of the CNN used in the embodiment.

[0087] Table 1 CNN architecture of adaptive federated learning algorithm

[0088] Layer type Number of nodes / filters Other parameters Input layer - Input: number of features Convolutional layer 1 32 filters Kernel size: 3 x 3 Max pooling layer 1 - Pool size: 2 x 2 Convolutional layer 2 64 filters Kernel size: 3 x 3 Max pooling layer 2 - Pool size: 2 x 2 Flattening layer - - Dense layer 128 nodes Activation function: ReLU Dropout - Dropout probability: 0.5 Output layer 2 nodes Activation function: Softmax

[0089] Further, the gradient change of the computing model includes:

[0090] Calculate the gradient change between the first round cloud server model and the second round cloud server model , the gradient change being the subtraction of the two SGD gradients;

[0091] The first round cloud server model and the second round cloud server model have been saved when the edge server receives the cloud server model parameters;

[0092] Each edge server maintains a version queue to store the model sequence issued by the cloud server in the order of receipt , strictly following the first-in first-out (FIFO) principle for processing, and the version loading takes out the current version to be processed from the head of the queue (initially ), and loads it as the initial parameters of the local model. After the edge server completes the processing of the current version , it is automatically removed from the queue , and starts processing the next version . Even if the cloud server has iterated to , the edge server must execute all intermediate version training in order and cannot skip any version;

[0093] Calculate the gradient change between the first round cloud server model and the corrected edge server model The gradient change between each layer is the difference between two SGD gradients.

[0094] Furthermore, the calculation formula for correcting the edge server model based on the cosine similarity penalty term and the dynamic gradient optimization penalty term is as follows:

[0095] ;

[0096] ;

[0097] in, Represents edge server The regularization objective function, This represents the gradient of the parameters of the regularization objective function. Indicates the first Wheel cloud server model With the modified edge server model Gradient change between For learning rate, and These represent the cosine similarity penalty coefficient and the dynamic gradient optimization penalty coefficient, respectively. This represents the cosine similarity penalty term. This represents the dynamic gradient optimization penalty term, with weights after the model parameters are updated. For simplicity, it will still be written as .

[0098] Furthermore, the cosine similarity penalty term , used to indicate the first Wheel cloud server model With the modified edge server model The similarity of gradients between different layers is calculated using the following formula:

[0099] ;

[0100] in, Indicates the first Wheel cloud server model With the modified edge server model No. Gradient changes in layers Indicates the first Wheel edge server Model No. Layer model parameters, Indicates the first Wheel cloud server model With the Wheel cloud server model No. gradient change of layers, representing the number of layers of the model. round cloud server model the number of layers of the model. model parameters of layers, representing the number of layers of the model.

[0101] Further, the dynamic gradient optimization penalty term , used to represent the consistency of the global gradient between the round cloud server model and the modified edge server model , the calculation formula is as follows:

[0102] ;

[0103] wherein, representing the gradient change between the round cloud server model and the modified edge server model , representing the gradient change between the round cloud server model and the round cloud server model global gradient change.

[0104] Further, the information of the modified edge server model includes: edge server model parameters, classification accuracy of the edge server model for each type of attack, and data amount of each type of attack category input during edge server local training.

[0105] Further, in combination with Figure 3 , the edge server model parameters aggregated based on the comprehensive weight and the model parameters of the current cloud server model, the aggregation formula is as follows:

[0106] ;

[0107] ;

[0108] wherein, representing the cloud server model parameters of the cloud server at the current time, representing the edge server uploading the model parameters of the round modified edge server model to the cloud server, representing the aggregated cloud server model parameters, representing the comprehensive weight, which is calculated by the penalty weight of the cosine similarity and the adaptive weight penalty term weight of the classification accuracy of each type of attack corresponding to the uploaded edge server. The penalized weights representing cosine similarity. This represents the adaptive weights that represent the amount of data and its importance on the edge servers in each round. This represents the adjustment coefficient used to adjust weight preferences in the overall weighting, with a value between 0 and 1. This indicates the frequency of updates for the fairness correction item.

[0109] Furthermore, the penalized weights of the cosine similarity The calculation formula is as follows:

[0110] ;

[0111] ;

[0112] in, For the edge server No. The modified edge server model is uploaded to the cloud server to construct the model using the model parameters. The cosine similarity between the gradient of the model and the gradient of the cloud server at the current moment. This is the sign function, used to adjust the sign of the weights based on the consistency between the local and global update directions. As the initial weight for the amount of data on the edge servers, Indicates the first Wheel cloud server model With the modified model No. Gradient changes in layers Indicates the first Wheel edge server The model number Layer model parameters, This represents the model in which the cloud server is currently running. With the Wheel cloud server model No. Gradient changes in the layer.

[0113] Furthermore, the initial weights of the data volume of each edge server. This represents the proportion of data volume on each edge server, and the calculation formula is as follows:

[0114] ;

[0115] in, Represents edge server Medium data volume This represents the total amount of data across all edge servers.

[0116] Further, the adaptive weight based on the respective situation of each edge server in each round The specific calculation formula is as follows:

[0117] ;

[0118] ;

[0119] Wherein, the edge server The data category is represented as , represents the initial weight of the data size of each edge server, represents the parameter that determines the influence degree of the accuracy difference on the weight adjustment, represents the data size of the data category in the edge server , represents the total data size of the data category of each edge server, represents the new data category set, represents the local accuracy of the data category in the edge server , indicates that when aggregated, the cloud server determines the importance degree of the data category according to the classification accuracy uploaded by the edge server, when the average recognition accuracy of the edge server model to the new data category is less than 90%, the data category is called a new data category, which is added to the new data category set , increase its weight, improve the recognition ability of the model to the new data category, represents the current training round, represents an intermediate round of adjustment point greater than the current training round, equal to the set maximum training round divided by 2.

[0120] Further, the update frequency fairness correction term The fairness correction term formula is as follows:

[0121] ;

[0122] Wherein, represents the average update times of the edge server, and the calculation formula is:

[0123] ;

[0124] Wherein, represents the current training round, represents the number of edge servers participating in federated learning, For the update frequency of edge servers, when the device's update frequency is within a reasonable range, a fair correction item is applied. The weight is set to 1 to ensure that the weight of devices with normal update frequency remains unchanged. When the update frequency of a device is too high, the weight of the fairness correction item will decrease to prevent bias caused by excessive updates during the cloud server model optimization process. When the update frequency of a device is too low, the weight will also decrease, but the minimum weight is set to 0.05 to avoid the weight from becoming completely ineffective in extreme cases.

[0125] Example 2

[0126] The following describes an adaptive asynchronous federated learning intrusion detection device provided by an embodiment of the present invention. The adaptive asynchronous federated learning intrusion detection device described below can be referred to in correspondence with the adaptive asynchronous federated learning intrusion detection method described above.

[0127] refer to Figure 4 An adaptive asynchronous federated learning intrusion detection device includes:

[0128] Data acquisition and preprocessing module: used to acquire known and publicly available DDoS datasets, preprocess the data, and ensure that the data format and quality meet the training requirements;

[0129] Initial model training module: used to train the initial cloud server model on the cloud server and distribute the model parameters of the initial cloud server model to each edge server;

[0130] Edge model training and correction module: This module is used by the edge server to receive model parameters from the cloud server model, load them as the initial parameters of the local model, train the model locally, generate an updated model, calculate the gradient change of the model, correct the model trained on the edge server based on the cosine similarity penalty term and the dynamic gradient optimization penalty term, obtain the corrected edge server model, and upload the model information to the cloud server for aggregation.

[0131] Model aggregation and distribution module: When the cloud server adopts an asynchronous aggregation mechanism, if the cloud server receives the update information from the edge server, it calculates the comprehensive weight, aggregates the model parameters of the edge server with the model parameters of the current cloud server based on the comprehensive weight, and obtains the new cloud server model parameters. Otherwise, the cloud server model remains unchanged. If the cloud server model converges, the federated learning global training process is terminated; otherwise, the new cloud server model parameters are distributed to each edge server.

[0132] Intrusion detection module: Distributes the converged cloud server model to each edge server for intrusion detection, records the detection results for further analysis and alarm processing, and feeds them back to the cloud server for further optimization of the model.

[0133] In this embodiment, to simulate the non-independent and identically distributed nature of data in a real edge federation environment, a power-law distribution is used to generate the sample size ratio in the data preprocessing section. Since network traffic data has a complex and diverse feature set, and considering the non-sequential nature of data features after removing network packets (i.e., the features of network data packets do not have obvious time dependencies), a convolutional neural network (CNN) is chosen as the core detection model to train the initial cloud server model. Cloud server model The parameters are sent to each edge server, and then in the federated learning process... Wheel, edge server Receive cloud server model The model parameters are loaded as the initial parameters of the local model, and the model is trained locally to generate an updated model. Simultaneously, the gradient change of each layer is calculated. Then, the edge server model is corrected based on the cosine similarity penalty term and the dynamic gradient optimization penalty term to obtain the corrected edge server model. The information of the corrected edge server model is uploaded to the cloud server for updates. The cloud server uses an asynchronous aggregation mechanism. If the cloud server receives the updated information from the edge server, it calculates the comprehensive weight and aggregates the edge server model parameters with the current cloud server model parameters based on the comprehensive weight to obtain new cloud server model parameters. Otherwise, the cloud server model remains unchanged. If the cloud server parameters are no longer updated, the federated learning global training process is terminated. Otherwise, the new cloud server model parameters are distributed to each edge server. Finally, the distribution and uploading operations are repeated until the global model is no longer updated.

[0134] The embodiments in this specification are described in a progressive manner. The embodiments focus on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other.

[0135] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An adaptive asynchronous federated learning intrusion detection method, characterized in that, Comprise: S10. The cloud server trains an initial cloud server model, and model parameters of the trained initial cloud server model are sent to each edge server connected to the cloud server at the current moment; S20. The edge server loads the model parameters sent by the cloud server as initial parameters of an edge server model, and updates the edge server model through local training while calculating gradient changes of the model; S30. The edge server model is corrected based on a cosine similarity penalty term and a dynamic gradient optimization penalty term to obtain a corrected edge server model, and information of the corrected edge server model is uploaded to the cloud server for updating; S40. When the cloud server receives the update information of the edge server, a comprehensive weight is calculated, and edge server model parameters and current cloud server model parameters are aggregated based on the comprehensive weight to obtain new cloud server model parameters; S50. If the cloud server model converges, the training is ended, and S60 is turned to, otherwise the new cloud server model parameters are sent to each edge server, and then S20 is turned to; S60. The converged cloud server model is sent to each edge server for intrusion detection; The calculation formula of the edge server model corrected based on the cosine similarity penalty term and the dynamic gradient optimization penalty term is as follows: ; ; wherein, denotes the regularization objective function of the edge server , denotes the parameter gradient of the regularization objective function, denotes the model of the edge server in the i-th round, denotes the model of the edge server in the i-th round, denotes the input feature of the j-th sample, denotes the input feature of the j-th sample, denotes the gradient change between the modified edge server model and the cloud server model in the i-th round, is the learning rate, and respectively denote the cosine similarity penalty coefficient and the dynamic gradient optimization penalty coefficient, denotes the cosine similarity penalty term, denotes the dynamic gradient optimization penalty term, and the weight is after the model parameter update is completed, for simplicity of representation.

2. The adaptive asynchronous federated learning intrusion detection method of claim 1, characterized in that: The edge servers currently connected to the cloud server are represented as follows: , The first one indicates the connection to the cloud server There are 1 edge server; each edge server has a local dataset, of which Indicates the first The local dataset of each edge server is equipped with a communication base station, and the edge server interacts with the cloud server and user terminal through the communication base station. The local dataset is represented as: ; in, Indicates the first Input features of each sample express The labels, represented in the global dataset are as follows The global dataset is the union of the local datasets of each edge server. ,in Indicates an index variable. This indicates the number of edge servers participating in federated learning.

3. The method of claim 1, wherein the method further comprises: The calculation of the gradient changes of the model comprises: computing the first round robin server model with the first round robin server model gradient change between each layer, the gradient change being the subtraction of two SGD gradients; the first Round robin server model and the first Round robin server model at the edge server saved when receiving the cloud server model parameters; Compute the first Round robin server model The gradient change between each layer of the modified edge server model is the subtraction of two SGD gradients.

4. The adaptive asynchronous federated learning intrusion detection method of claim 1, characterized in that: The cosine similarity penalty term , for representing the first Wheel cloud server model The similarity between the layers of the modified edge server model Gradient, the calculation formula is: ; wherein, represents the round edge server model with the modified edge server model the gradient change of the layer, represents the round edge server the model parameters of the layer, represents the round cloud server model with the round cloud server model the gradient change of the layer, represents the round cloud server model the model parameters of the layer, represents the number of layers of the model.

5. The adaptive asynchronous federated learning intrusion detection method of claim 1, characterized in that: The dynamic gradient optimization penalty term , for representing the first Wheel cloud server model The consistency of the global gradient between the modified edge server model , the calculation formula is as follows: ; wherein, denotes the cloud server model with the modified edge server model gradient change, denotes the cloud server model with the cloud server model global gradient change.

6. The adaptive asynchronous federated learning intrusion detection method of claim 1, characterized in that: The information of the corrected edge server model comprises edge server model parameters, classification accuracy of the edge server model on each type of attack, and data volume of each type of attack category input during local training of the edge server.

7. The method of claim 1, wherein the method further comprises: The aggregation formula of the aggregation of the edge server model parameters and the model parameters of the current cloud server model based on the comprehensive weight is as follows: ; ; wherein, represents the cloud server model parameters of the cloud server at the current time, represents the edge server The first the modified edge server model uploaded to the model parameters of the cloud server in the round, represents the aggregated cloud server model model parameters, represents the comprehensive weight, which is calculated by the penalty weight of the cosine similarity corresponding to the uploaded edge server and the adaptive weight of the classification accuracy of each type of attack, represents the penalty weight of the cosine similarity, represents the adaptive weight of the data volume and data importance of each round of edge server, represents the adjustment coefficient in the comprehensive weight for adjusting the weight preference, represents the update frequency fairness correction term.

8. The adaptive asynchronous federated learning intrusion detection method of claim 7, characterized in that: a penalty weight of the cosine similarity The calculation formula is as follows: ; ; in, For the edge server No. The modified edge server model is uploaded to the cloud server to construct the model using the model parameters. The cosine similarity between the gradient of the model and the gradient of the cloud server at the current moment. This is the sign function, used to adjust the sign of the weights based on the consistency between the local and global update directions. As the initial weight for the amount of data on the edge servers, Indicates the first Wheel cloud server model With the modified model No. Gradient changes in layers, Indicates the first Wheel edge server The model number Layer model parameters, This represents the model in which the cloud server is currently running. With the Wheel cloud server model No. Gradient changes in layers, The model representing the current running state of the cloud server. No. Layer model parameters, Indicates the number of layers in the model.

9. The method of claim 8, wherein the method further comprises: an initial weight of a data volume size of the edge server , for representing the proportion of the data volume of the edge server, and the calculation formula is as follows: ; wherein, represents an edge server represents the data volume in the edge server, represents the total data volume of all edge servers.

10. The method of adaptive asynchronous federated learning intrusion detection of claim 7, wherein, The adaptive weight of data volume and data importance of each round edge server The calculation formula is as follows: ; ; Among them, the edge server Data categories are represented as , The initial weight representing the data size of each edge server, The parameter representing the influence degree of the decision accuracy difference on the weight adjustment, The data size of the data category in the edge server , The total data size of each edge server data category , The new data category set, The local accuracy of the data category in the edge server , The importance degree of the data category is determined by the cloud server according to the classification accuracy uploaded by the edge server when aggregating, when the average recognition accuracy of the edge server model to the new data category is less than 90%, the data category is called a new data category, which is added to the new data category set The current training round, The intermediate round of the node is greater than the current training round.

11. The method of adaptive asynchronous federated learning intrusion detection of claim 7, wherein, The update frequency fair correction term The calculation formula is as follows: ; wherein, represents the average number of updates of the edge server, and the calculation formula is: ; wherein, denotes the current training round, denotes the number of edge servers participating in federated learning, is the update frequency of the edge server.

12. An adaptive asynchronous federated learning intrusion detection apparatus, comprising: The adaptive asynchronous federated learning intrusion detection method of any one of claims 1-11 comprises: A data collection and preprocessing module for collecting known public DDoS data sets and preprocessing data to ensure that the data format and quality meet the training requirements; An initial model training module for training an initial cloud server model by the cloud server, and sending model parameters of the initial cloud server model to each edge server; An initial model training module for training an initial cloud server model by the cloud server, and sending model parameters of the initial cloud server model to each edge server; An edge model training and correction module is configured to receive model parameters of a cloud server model by an edge server, load the model parameters as initial parameters of a local model, perform local training, generate an updated model, calculate gradient changes of the model, correct the model trained by the edge server based on a cosine similarity penalty term and a dynamic gradient optimization penalty term, obtain a corrected edge server model, and upload information of the model to a cloud server for aggregation. A model aggregation and delivery module is configured to use an asynchronous aggregation mechanism by the cloud server. If the cloud server receives update information of the edge server, the cloud server calculates a comprehensive weight, aggregates model parameters of the edge server and current model parameters of the cloud server based on the comprehensive weight, obtains new model parameters of the cloud server, and otherwise, the model of the cloud server remains unchanged. If the model of the cloud server converges, the global training process of the federated learning is terminated, otherwise, the new model parameters of the cloud server are delivered to each edge server. An intrusion detection module is configured to deliver the converged model of the cloud server to each edge server for intrusion detection, record detection results for further analysis and alarm processing, and feed back to the cloud server for further optimization of the model.

Citation Information

Patent Citations

  • Network intrusion cooperative detection method based on federated learning

    CN114640498A

  • Industrial control network intrusion detection method based on decentralized federated learning

    CN114785608A

  • Federal learning load prediction method based on dynamic weighted aggregation

    CN114707765A

  • Industrial internet cloud edge model aggregation method based on federated learning

    CN116204793A