Federal clustering abnormal flow detection method based on feature importance

By computing feature importance and hierarchical clustering, the problem of heterogeneity of equipment and non-independent and homogeneous distribution of data in federated learning is solved, and lightweight and efficient abnormal traffic detection is achieved.

CN120378170APending Publication Date: 2025-07-25FUZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510540708.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In federated learning, device heterogeneity and data non-independent homogeneity lead to difficulty in training global models, and traditional feature selection algorithms cannot be applied, affecting detection performance and increasing training time.

Method used

By calculating the importance sorting of features based on JS divergence and maximum information coefficient, performing hierarchical clustering, selecting parallel federated learning of features within clusters, alleviating the difference in data distribution between devices, and reducing the amount of model parameters.

Benefits of technology

Lightweight abnormal traffic detection is realized, reducing the performance threshold for deploying equipment, and improving detection performance after feature selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378170A_ABST
    Figure CN120378170A_ABST
Patent Text Reader

Abstract

The invention provides a federal clustering abnormal traffic detection method based on feature importance, which comprises the following steps: S1, all clients collect traffic data from network data streams, and pre-process the traffic data to form a local data set with labels; s2, the feature importance sorting set is uploaded to a server # imgabs0 #; s3, the server # imgabs1 # executes hierarchical clustering based on the importance of the clients to obtain each client cluster, and the server # imgabs2 # calculates the average feature importance sequence of each client cluster; s4, the server sends the global model to the client cluster corresponding to the server, and finally, the local models are aggregated on the server to obtain the global model corresponding to the client cluster; repeating the steps until the global model converges; s5, performing prediction classification on the test sample by using the global model of the client cluster where the test sample is located; by applying the technical scheme, the performance threshold of the deployment equipment can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intrusion detection and machine learning, and in particular to a federated clustering abnormal traffic detection method based on feature importance. Background Art

[0002] With the continuous enrichment of the types of devices accessing the network, the situation of heterogeneous devices inside the network is becoming more and more common. At the same time, running multiple services on the devices will also cause the collected traffic data to show statistical heterogeneity. In federated anomaly detection, device heterogeneity may lead to certain differences in computing performance among the devices participating in the training, thus restricting the expansion of the global model parameter scale. And in typical federated learning algorithms, the time required to complete one round of communication is often limited by the device with the worst performance among the participating devices. Statistical heterogeneity means that the data on different devices is not independently and identically distributed, and the degree of Non-IID will affect the fitting degree of the global model to the data distribution on each participating device. In severe cases, it may lead to the failure of global model training to converge.

[0003] Network anomaly detection usually takes high-dimensional network traffic data as input, and the features of these data often have redundancy. At present, most research on network anomaly detection uses deep neural networks, and deep neural networks have an end-to-end workflow and black-box characteristics. Therefore, processing the input features and removing the redundant part through interpretable feature selection can improve the training and detection efficiency and guide the improvement direction for researchers. And in order to alleviate the impact brought by device heterogeneity, it is usually necessary to reduce the number of model parameters. Reducing the feature dimension of the input samples is one of the common methods. However, under the constraints of federated learning, the original data cannot leave the local device, so the complete sample set cannot be accessed, resulting in the difficulty of applying traditional feature selection algorithms to federated learning. Horizontal federated learning assumes that there is a large overlap in the data feature space among the participating clients in the training, but not all features of all clients are beneficial to the global model training. Using all features without selection for federated learning may damage the detection performance of the global model and increase the training time. Therefore, feature selection is also applicable to federated learning. However, there are related problems in performing feature selection in federated learning. First, since the original data in federated learning cannot leave the local, traditional feature selection schemes cannot be directly applied to all the data. Second, when each client performs feature selection locally alone, due to the lack of a global perspective, there are certain biases in the features selected among the clients. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a federated clustering abnormal traffic detection method based on feature importance. After clustering, the global feature importance can be obtained to perform feature selection on the samples within the cluster, making the proposed scheme more lightweight and reducing the performance threshold of the deployed devices.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A federated clustering-based abnormal traffic detection method based on feature importance, comprising the following steps:

[0006] Step S1: All clients collect traffic data from the network data stream, and form a labeled local data set after preprocessing;

[0007] Step S2: All clients obtain the feature importance ranking using JS divergence and the maximum information coefficient, and upload the feature importance ranking set to server A;

[0008] Step S3: Server A performs hierarchical clustering based on client importance to obtain each client cluster. Server A calculates the average feature importance ranking of each client cluster, and selects the first part of the features as the result of feature selection. Assign a server to each client cluster, and initialize the global model of each server, so as to alleviate the impact brought by the non-independent and identically distributed data among clients;

[0009] Step S4: In one round of global federated communication, each client cluster independently performs the following steps: The server sends the global model to its corresponding client cluster, then the client cluster trains and updates the local model for the sent global model, then the client cluster uploads the updated local model to the server, and finally aggregates the local models on the server to obtain the global model corresponding to the client cluster. Repeat the above steps until the global model converges;

[0010] Step S5: The test samples are predicted and classified using the global model of the client cluster where they are located.

[0011] Further, the specific steps of step S2 are as follows:

[0012] Step S21: Experimental parameter setting: In this method, set a total of K independent federated learning models according to the scenario requirements, including a total of K servers and C clients. Each client c i (i ∈ [1, C]) has a local data set X i , and all client data sets constitute the total data set X = X1 ∪ X2... ∪ X C , N = |X|, where N represents the number of samples in the total data set, and each sample x j ∈ X has features where M is the total number of features, the sample label set Y = {y1, y2,..., y N}, and the label category set L = {l n , l a1 , l a2 , …, l aT}(where l nIt is of the normal class normal, l ai It is the i-th type of attack, and there are a total of T types of attacks). Assume that each client's dataset has at least samples of normal traffic and samples of several random types of attacks.

[0013] Step S22: C clients respectively use the JS divergence for each feature f m Calculate their distributions in different attack classes (l' ∈ [1, T'], where T' represents the total number of T' attack types on the current client) and the distribution in normal traffic And perform a mixed calculation on the two according to the JS divergence formula, as shown in formula (1)

[0014]

[0015] Thus, the feature discrimination degree is obtained

[0016] Step S23: For the label set Y = {y1, y2,..., y N} of the sample set X and the values of the feature f m in X Calculate the mutual information I(F m , Y) and the maximum information coefficient MIC(F m , Y), as shown in formula (2) and formula (3) (where B m , B y are the number of grid partitions in the F m , Y directions respectively. Grid partitioning means dividing the value ranges of F m , Y into several intervals respectively to form a two-dimensional grid structure);

[0017]

[0018] Step S24: For Perform a descending order sorting to obtain the subscript set

[0019] Step S25: Perform an ascending order sorting on MIC(F m , Y) to obtain the subscript set R MIC ;

[0020] Step S26: According to the proportion of samples of attack type l', perform a weighted sum on , as shown in formula (4), to obtain the result R JSC ,

[0021]

[0022] Then calculate the feature importance ranking of client c Obtain the set of sorted feature importances I = {I1, I2, …, I C} for C clients;

[0023] Step S27: Upload the set of sorted feature importances I to server A.

[0024] Further, step S3 is specifically as follows:

[0025] Step S31: Server A initializes C clients as separate client clusters respectively, obtaining the initial client cluster set CLS = {CL1, CL2, …, CL C};

[0026] Step S32: Use the Spearman Footrule distance algorithm to calculate the distance between any two clusters, that is, multiply the difference in ranks in the two sorts for each feature by the mean of the ranks as the weight Spear(I a , I b ), as shown in formula (5) (where a, b ∈ [1, C] and a ≠ b), and store it in the pairwise client cluster distance set Q = {Spear1, Spear2... Spear V}, where D is the number of client clusters in the current round of hierarchical clustering, and the initial D is C;

[0027]

[0028] Step S33: Select the two CLs corresponding to the two clusters with the smallest distance in CLS according to Q i , CL j , and merge these two clusters into a new cluster CL' i ;

[0029] Step S34: Remove CL i , CL j from the client cluster set CLS, and insert CL' i into CLS;

[0030] Step S35: Using the average linkage algorithm, the distance between the new cluster CL' i and other clusters CL k (k ∈ D, k ≠ i, j) is the average of the Spearman Footrule distances between all elements in the two old clusters CL i , CL j (i.e., the elements in the new cluster CL' i ) and the cluster CL k , and update the set Q, as shown in formula (6);

[0031]

[0032] Step S36: Repeat S33 to S35 until the number of client clusters converges to the target number of clusters K, and obtain the client cluster set CLS = {CL1, CL2, …, CL K};

[0033] Step S37: Server A calculates the average feature importance within each of its client clusters CL k within.

[0034] Step S38: Obtain the in-cluster feature importance ranking in descending order according to Then select the top S features as the result of feature selection and apply them to all clients within the cluster;

[0035] Step S39: Assign a server to each client cluster, initialize each global model, and perform independent federated learning on different client clusters in parallel, so as to mitigate the impact brought by non-independent and identically distributed data among clients.

[0036] Furthermore, the specific steps of Step S4 are as follows:

[0037] Step S41: For each client cluster, the server sends its respective global model to it. The client cluster trains the received global model and updates the local model, and then uploads the local model parameters (r is the current round);

[0038] Step S42: The server inputs the model parameters i uploaded by the corresponding client cluster CL into the aggregation function to update the global model parameters W i,r+1 , as shown in formula (8).

[0039]

[0040] Where is the number of samples on the c j -th client, n is the sum of the number of samples of all |CL i | clients within the client cluster, and r is the communication round between the server and the client;

[0041] Step S43: Repeat S41 to S42 until all global models converge;

[0042] Step S44: Finally, obtain K global model parameter sets WS = {W1, W2, …, W K}, and distribute each global model to the corresponding client cluster.​

[0043] Compared with the prior art, the present invention has the following beneficial effects: By calculating the feature importance on the client for hierarchical clustering and performing independent federated learning in parallel on different client clusters, the impact brought by non-independent and identically distributed data among clients is alleviated. After clustering, the global feature importance can be obtained to perform feature selection on the samples within the cluster, making the proposed scheme more lightweight and reducing the performance threshold of the deployed devices. The experimental results show that when the proposed scheme is only used for clustering, the effect is close to that of the advanced scheme. When each scheme only selects 50% of the features, the detection performance of the proposed scheme is better than that of the advanced scheme. Description of the Drawings

[0044] Figure 1 It is a diagram of the federated clustering anomaly traffic detection method based on feature importance in an embodiment of the present invention.

[0045] Figure 2 It is a flowchart of feature importance calculation in an embodiment of the present invention.

[0046] Figure 3 It is a flowchart of client clustering based on hierarchical clustering in an embodiment of the present invention.

[0047] Figure 4 It is a comparison diagram with the baseline scheme on the UNSW-NB15 dataset in an embodiment of the present invention.

[0048] Figure 5 It is a comparison diagram with the baseline scheme on the CIC-IDS2018 dataset in an embodiment of the present invention.

[0049] Figure 6 It is a comparison diagram with the advanced scheme on the UNSW-NB15 dataset in an embodiment of the present invention.

[0050] Figure 7 It is a comparison diagram with the advanced scheme on the CIC-IDS2018 dataset in an embodiment of the present invention.

[0051] Figure 8 It is a performance diagram after feature selection in comparison with the baseline scheme on the UNSW-NB15 dataset in an embodiment of the present invention.

[0052] Figure 9 It is a performance diagram after feature selection in comparison with the baseline scheme on the CIC-IDS2018 dataset in an embodiment of the present invention.

[0053] Figure 10 It is a performance diagram after feature selection in comparison with the advanced scheme on the UNSW-NB15 dataset in an embodiment of the present invention.

[0054] Figure 11It is the performance graph after feature selection in the CIC-IDS2018 dataset in the embodiments of the present invention for comparison with advanced solutions. Detailed implementation manners

[0055] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0056] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0057] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0058] Reference Figure 1 , the present invention provides a federated clustering-based abnormal traffic detection method based on feature importance; based on the federated learning framework, each client calculates its own feature importance ranking based on JS divergence and the maximum information coefficient, and server A performs hierarchical clustering according to the feature importance of the clients, divides the clients into K client clusters, and thus forms K global models. In the initial stage, each client collects traffic data from the network data stream, preprocesses it to form a local dataset X c , and obtains the feature importance ranking I using JS divergence and the maximum information coefficient c ; all clients upload the feature importance ranking set to server A; server A performs hierarchical clustering based on client importance to obtain K client clusters; server A calculates the average feature importance of each client cluster, and selects the top S features as the result of feature selection; in each global communication round, each client cluster independently performs the following steps: the server sends the global model to its corresponding client cluster, then the client cluster trains and updates the local model for the downloaded global model, then the client cluster uploads the updated local model to the server, and finally aggregates the local models on the server to obtain the global model corresponding to the client cluster, and repeats the above steps until the global model converges. Finally, K global models are obtained. As Figure 1 shown, it specifically includes the following steps:

[0059] Step (1): Training dataset collection and preprocessing;

[0060] In this embodiment, the publicly available UNSW-NB15 dataset and the training and test datasets in CSE-CIC-IDS2018 are used. For the UNSW-NB15 dataset, the character-type features are converted into discrete numerical values through label encoding, and then standard normalization is performed on it; for the CSE-CIC-IDS2018 dataset, there are no character-type features on this dataset. After removing the useless timestamp features, min-max normalization is performed on it.

[0061] Step (2): Feature importance based on Jensen-Shannon divergence and maximum information coefficient;

[0062] All clients obtain the feature importance ranking based on Jensen-Shannon divergence and maximum information coefficient, and upload the feature importance ranking set to Server A;

[0063] The calculation process of feature importance based on Jensen-Shannon divergence and maximum information coefficient is as Figure 2 shown.

[0064] Furthermore, in the said step (2), experimental parameter setting: In this method, according to the scenario requirements, a set of K independent federated learning models is set up, which includes a total of K servers and C clients. Each client c i (i ∈ [1, C]) has a local dataset X i , and all client datasets constitute the total dataset X = X1 ∪ X2... ∪ X C , N = |X|, where N represents the number of samples in the total dataset. Each sample x j ∈ X has features where M is the total number of features. The sample label set Y = {y1, y2,..., y N}, and the label category set L = {l n , l a1 , l a2 , …, l aT}(where l n is the normal class normal, and l ai is the i-th attack type, and there are a total of T attack types). It is assumed that each client's dataset has at least samples of normal traffic and samples of several random attack types. The C clients respectively use Jensen-Shannon divergence to calculate the distribution of each feature f m in different attack classes (l' ∈ [1, T'], T' represents the total number of T' attack types on the current client) and the distribution in normal traffic and perform a mixed calculation on the two according to the Jensen-Shannon divergence formula to obtain the feature discrimination For the label set Y = {y1, y2, …, y N}, and feature f m The set of values in X Calculate the mutual information I(F m , Y) and the maximum information coefficient MIC(F m , Y); for Perform a descending sort to obtain the subscript set For MIC(F m , Y) perform a descending sort to obtain the subscript set R MIC ; According to the proportion of samples with attack type l' for Perform a weighted sum to obtain the result R JSC , calculate the feature importance ranking of client c Obtain the feature importance ranking set I = {I1, I2,..., I C} of C clients; Upload the feature importance ranking set I to server A.

[0065] Step (3): Hierarchical clustering based on the Spearman Footrule distance;

[0066] Server A performs hierarchical clustering based on client importance to obtain each client cluster. Server A calculates the average feature importance ranking of each client cluster, selects the first part of the features as the result of feature selection, assigns a server to each client cluster, and initializes the global model of each server, thereby alleviating the impact brought by non-independent and identically distributed data among clients;

[0067] The client clustering process based on hierarchical clustering is as Figure 3 shown.

[0068] Furthermore, in the said step (3), server A initializes C clients as separate client clusters respectively to obtain the initial client cluster set CLS = {CL1, CL2,..., CL C}; Using the Spearman Footrule distance algorithm, calculate the distance Spear(I a , I b ) between any two clusters, and store it in the pairwise client cluster distance set Q = {Spear1, Spear2...Spear V}, where D is the number of client clusters in the current round of hierarchical clustering, and the initial D is C; Select the CL i , CL j corresponding to the two clusters with the smallest distance in CLS according to Q, and merge these two clusters into a new cluster CL' i ; Remove CL i , CLj , and insert CL' i into CLS; Use the average-linkage algorithm to refine the set Q. Repeat the above steps until the number of client clusters converges to the target number of clusters K, and obtain the client cluster set CLS = {CL1, CL2, …, CL K}; Server A calculates the average feature importance k within each client cluster CL According to in descending order to obtain the in-cluster feature importance ranking Then take the top S features as the result of feature selection and apply them to all clients within the cluster; Assign a server to each client cluster, initialize each global model, and perform independent federated learning in parallel for different client clusters, thereby alleviating the impact brought by non-independent and identically distributed data among clients.

[0069] Step (4): In one round of global federated communication, each client cluster independently performs local training and model update, and then aggregates the global model on the corresponding server, updates the global model parameters, and repeats the above steps until the global model converges.

[0070] Furthermore, in step (4), for each client cluster, the server sends its respective global model to it, and the global model sent down is trained within the client cluster, the local model is updated, and the local model parameters are uploaded (r is the current round); The server aggregates the model parameters i uploaded by the corresponding client cluster CL into the aggregation function to update the global model parameters W i,r+1 , and repeat the above steps until all global models converge, and finally obtain K global model parameter sets WS = {W1, W2, …, W K}, and distribute each global model to the corresponding client cluster.

[0071] Preferably, in this embodiment, the global model adopts a two-layer CNN model, and each layer from input to output is 1D-CNN, ReLU, 1D-CNN, ReLU, FullConnect, ReLU, FullConnect in sequence. The parameters of the first 1D-CNN layer are set as: convolution kernel size: 6, number of convolution kernels: 128, the parameters of the second 1D-CNN layer are set as: number of hidden layer units: 6, number of layers: 256, the parameters of the first FullConnect layer are set as: input: (feature dimension - 6) * 256, output: 256, and the parameters of the second FullConnect layer are set as: input: 256, output: number of classes.

[0072] Step (5): Intrusion detection;

[0073] The test samples are predicted and classified using the global model of their client cluster.

[0074] Preferably, the default settings during the simulation experiment in this embodiment are as follows: Two training sets, UNSW-NB15 and CSE-CIC-IDS2018, are used for training, and the built-in test set is used for testing. To create a situation of label skew, the dataset on the client is divided based on the following settings in this example: The parameter α controls the proportion of clients among all C clients that contain all T types of attacks on the client, and the parameter β controls the proportion of the attack types contained in each client among the remaining clients to the total number of attack types, thereby creating a situation where the data distribution on the client is non-independent and identically distributed. By controlling the parameter β, the types of label distributions on the existing clients are approximately All clients contain normal traffic. In terms of hyperparameter settings, the cases where K takes 5 and 10 are considered for both datasets. The number of communication rounds for federated learning is 10 rounds, the local client iteratively trains 5 times, each batch during training has 256 samples, the loss function is cross-entropy, and the learning rate is 0.001.

[0075] The clients are divided into training clients and test clients, and both the training clients and test clients divide the data according to the above settings. When evaluating the performance of the model, the maximum value of each test client on the K global models is taken and then the average value of all test clients is calculated as the final result.

[0076] Figure 4 To compare with the baseline scheme on the UNSW-NB15 dataset, the performance of the proposed scheme in this paper and other baseline schemes on the UNSW-NB15 dataset is shown. For the number K of global models, increasing the number of K does not always improve the detection performance in the case of random clustering. In the cases where β takes 0.5 and 0.7, the performance of the random clustering scheme with k = 10 is significantly lower than that of the random clustering scheme with k = 5. For the scheme proposed in this paper, when k = 5, except that the precision is lower than k = 5 in the case of β = 0.7, the detection performance is better than k = 5 in other cases. In the comparison with the baseline scheme, the scheme proposed in this paper is better than the baseline scheme in all cases. Among them, when k = 10 for the scheme proposed in this paper, the F1 score leads other schemes by 1.5% to 16.58%. Among the baseline schemes, the FedAVG with only a single global model has the lowest performance, and the random clustering scheme has better F1 scores than the single global model FedAVG under two different values of k when β = 0.4. However, in β = 0.5 and β = 0.7, when k = 10, the F1 score of the random clustering scheme is lower than that of the single global model FedAVG. And for the scheme proposed in this paper, in the cases where k takes different values, the performance is improved compared with the single global model FedAVG.

[0077] Figure 5 To compare with the baseline scheme on the CIC-IDS2018 dataset, the performance of the proposed scheme and other baseline schemes on the CIC-IDS2018 dataset is shown. When β takes 0.4 and 0.5, the proposed scheme leads the baseline scheme by 0.42% to 2.69% in terms of F1-score when k = 10. However, when β takes 0.6, the random clustering scheme outperforms the proposed scheme in all metrics when k = 5, leading the proposed scheme by 0.78% in terms of F1-score.

[0078] From the above experiments, it can be seen that in most cases, the proposed scheme is better than the baseline scheme, indicating that it has a certain effect in client selection.

[0079] Figure 6 To compare with the advanced scheme on the UNSW-NB15 dataset, the performance of the proposed scheme and the advanced scheme on the UNSW-NB15 dataset is shown. When β takes 0.4 and 0.5, kfed is better than other schemes when k = 5 / 10, and the proposed scheme is relatively close to kfed when k = 5 / 10, with the difference in F1-score ranging from 0.28% to 0.41%. When β takes 0.7, the proposed scheme is better than other schemes when k = 10, leading IFCA by 11.68% and leading kfed with k = 10 by 0.04% in terms of F1-score.

[0080] Figure 7 To compare with the advanced scheme on the CIC-IDS2018 dataset, the performance of the proposed scheme and the advanced scheme on the CIC-IDS2018 dataset is shown. When β takes 0.4 and 0.6, in terms of F1-score, both schemes with k = 10 are better than k = 5, and kfed is better than other schemes, leading the proposed scheme with the same parameter settings by 0.35% to 0.65%. When β = 0.5, the proposed scheme is better than other schemes when k = 10, leading IFCA by 6.98% and leading kfed with k = 5 by 0.2%.

[0081] It can be seen that the proposed scheme is close to the advanced scheme in performance.

[0082] Figure 8 For the performance after feature selection when comparing with the baseline scheme under the UNSW-NB15 dataset, Figure 9 For the performance after feature selection when comparing with the baseline scheme under the CIC-IDS2018 dataset, Figure 8 、 9It shows the performance comparison after feature selection between the proposed scheme and the baseline scheme under different datasets. It can be seen that the performance of the scheme proposed in this paper is significantly better than that of the baseline scheme after feature selection. On the UNSW-NB15 dataset, the proposed scheme leads the scheme of a single global model by 10.05% to 25.35% in terms of F1 score, and leads the scheme of random clustering with the same k = 10 by 5.23% to 15.67%. On the CIC-IDS2018 dataset, the proposed scheme leads the scheme of a single global model by 2.23% to 10.91% in terms of F1 score, and leads the scheme of random clustering with the same k = 10 by 1.37% to 11.20%.

[0083] Figure 10 Performance after feature selection for comparison with advanced schemes under the UNSW-NB15 dataset Figure 11 Performance after feature selection for comparison with advanced schemes under the CIC-IDS2018 dataset Figure 10 、 11 It shows the performance comparison after feature selection between the proposed scheme and the advanced scheme under different datasets. It can be seen that, benefiting from the advantages that the scheme proposed in this paper can comply with the limitations of federated learning and perform feature selection, the performance of the scheme proposed in this paper is better than that of the advanced scheme after feature selection. On the UNSW-NB15 dataset, with the same k = 10, it leads kfed by 0.9% to 2.62%. On the CIC-IDS 2018 dataset, with the same k = 10, it leads kfed by 1.05% to 2.81%. Compared with random feature selection, the performance of the scheme proposed in this paper drops by less than 1% after filtering half of the features, while the performance drops by 1% to 5% in the case of randomly filtering half of the features. It can be seen that the feature selection scheme proposed in this paper can select valuable features while complying with the limitations of federated learning and minimizing the impact on detection performance as much as possible.

[0084] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A federated clustering-based abnormal traffic detection method based on feature importance, characterized in that, Including the following steps: Step S1: All clients collect traffic data from the network data stream, and form a labeled local data set after preprocessing; Step S2: All clients obtain the feature importance ranking using JS divergence and the maximum information coefficient, and upload the feature importance ranking set to Server A; Step S3: Server A performs hierarchical clustering based on client importance to obtain each client cluster. Server A calculates the average feature importance ranking of each client cluster, and selects the first part of the features as the result of feature selection; assigns a server to each client cluster, and initializes the global model of each server; Step S4: In one round of global federated communication, each client cluster independently performs the following steps: The server sends the global model to its corresponding client cluster, then the global model sent down is trained within the client cluster and the local model is updated. Then the client cluster uploads the updated local model to the server. Finally, the local models are aggregated on the server to obtain the global model corresponding to the client cluster; Repeat the above steps until the global model converges; Step S5: The test samples are predicted and classified using the global model of their corresponding client cluster.

2. The method for detecting abnormal traffic in federal clustering based on feature importance according to claim 1, wherein The specific content of Step S2 is: Step S21: Experimental parameter setting: Set up k independent federated learning models according to the scenario requirements, with a total of K servers and C clients; each client c i (i ∈ [1, C]) has a local dataset X i , and all client datasets form the total dataset X = X1 ∪ X2... ∪ X C , N = |X|, where N represents the number of samples in the total dataset, and each sample x j ∈ X has features where M is the total number of features, the sample label set Y = {y1, y2, …, y N} and the label category set L = {l n , l a1 , l a2 , …, l aT} where l n is the normal class normal, l ai is the i-th type of attack, and there are T types of attacks in total; assume that each client's dataset has at least samples of normal traffic and samples of several random types of attacks; Step S22: C clients respectively use JS divergence for each feature f m to calculate their distributions in different attack classes respectively and their distributions in normal traffic and perform a mixed calculation on the two according to the JS divergence formula, as shown in formula (1), so as to obtain the feature discrimination degree l' ∈ [1, T'], where T' represents the total number of T' attack types on the current client; Step S23: label set Y = {y1, y2, ..., y N } and feature f m The set of values in X Calculate mutual information I(F m ,Y) and the maximum information coefficient MIC(F m ,Y); Formula (2) and Formula (3) show: Among which B m , B y are respectively the number of grid divisions in the F m and Y directions. Grid division means dividing the value ranges of F m and Y into several intervals respectively to form a two-dimensional grid structure; Step S24: For perform a descending sort to obtain a subscript set Step S25: Ascendingly sort MIC(F m , Y) to obtain the subscript set R MIC ; Step S26: Perform weighted summation based on the proportion of samples with attack type l', as shown in Formula (4), to obtain the result R ; JSC ; Then calculate the sorted feature importance of client c Obtain the sorted feature importance set I = {I1, I2, …, I C} for C clients; Step S27: Upload the feature importance ranking set I to Server A.

3. The method for detecting abnormal traffic in a federated clustering based on feature importance according to claim 1, characterized in that, The specific content of Step S3 is: Step S31: Server A initializes C clients into separate client clusters respectively, obtaining an initial client cluster set CLS = {CL1, CL2, …, CL C}; Step S32: Calculate the distance Spear(I a , I b ) between any two clusters using the Spearman Footrule distance algorithm as shown in formula (5), and store it in the pairwise client cluster distance set Q = {Spear1, Spear2... Spear V}, where D is the number of client clusters in the current round of hierarchical clustering, and the initial D is C; Step S33: Select the two CLs corresponding to the two clusters with the smallest distance in the LCS according to Q i , CL j , and merge these two clusters into a new cluster CL' i ; Step S34: Remove CL from the client cluster set CLS i , CL j , and insert CL' i into CLS; Step S35: Update the set Q using the average link algorithm, as shown in formula (6); Step S36: Repeat S33 to S35 until the number of client clusters converges to the target number of clusters K, and obtain the client cluster set CLS = {CL1, CL2, …, CL K}; Step S37: Server A calculates the average feature importance within each client cluster CL in the client cluster set CLS k of the client cluster Step S38: According to in descending order to obtain the in-cluster feature importance ranking Then take the top S features as the result of feature selection and apply them to all clients within the cluster; Step S39: Assign a server to each client cluster, and initialize each global model, and perform independent federated learning in parallel for different client clusters, so as to alleviate the impact brought by the non-independent and identically distributed data among clients.

4. A method for detecting abnormal traffic in a federated clustering based on feature importance according to claim 1, characterized in that, The specific content of Step S4 is: Step S41: For each client cluster, the server sends its respective global model to it. The client cluster trains the received global model and updates the local model, and then uploads the local model parameters to the server. r is the current round; Step S42: The server uploads the corresponding client clusters CL i , i ∈ [1, K] to update the global model parameter W using the input aggregation function, as shown in formula (8); ,r+1 ​ Step S43: Repeat Step S41 to Step S42 until all global models converge; Step S44: Finally, obtain K global model parameter sets WS = {W1, W2, …, W K}, and distribute each global model to the corresponding client cluster.

Citation Information

Cited By

  • Heterogeneous federal model adjusting method based on importance sampling

    CN120725101A

  • A heterogeneous federated model adjustment method based on importance sampling

    CN120725101B