A Feature Selection and Optimization Method for a Network Structure
By building a selection factor layer and dynamic programming algorithm in a high-speed network, the traffic characteristics that are adapted to downstream tasks are selected, and the problems of SRAM memory limitation and sharp increase in traffic are solved, adaptive feature selection and fast matching are achieved, and traffic recognition accuracy is improved.
Patent Information
- Application Number
- CN202510482704.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In high-speed network environments, it is difficult for the prior art to effectively use programmable switches to perform stream-level traffic characteristics statistics. Due to the problems of SRAM memory limitation and sharp increase in traffic, the performance requirements of downstream tasks are difficult to meet.
The task model is built and the selection factor layer is added. The features that meet the memory bit width limitation are filtered out through the dynamic programming algorithm, combined with the class balance loss function to optimize the feature selection, and the early stop mechanism is used to accelerate the feature selection process, design feature selection and optimization methods.
Under limited memory resources, it adapts to the downstream task model, reduces computing overhead, quickly matches feature, solves data imbalance problem, and improves traffic recognition accuracy.
Smart Images

Figure CN120017508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network model optimization, and particularly to a method for feature selection and optimization of a network structure. Background Art
[0002] Network telemetry systems provide continuous and real-time measurements of network status, which are crucial for network operators to detect increasingly complex events ranging from performance degradation to security attacks. Thanks to the rapid development of programmable switches, modern data plane telemetry systems increasingly utilize these devices, achieving scalability in handling high-speed traffic and enabling fast and efficient flow-level feature statistics.
[0003] In network management, programmable switches supporting flow-level telemetry technology are usually adopted to count flow-level features, and various flow-level traffic features are stored in a flow feature table in the SRAM of the data plane. The flow feature table is periodically uploaded to the control plane for guiding the judgment of downstream tasks. However, in high-speed networks, collecting flow-level traffic features by programmable switches is still challenging for the following reasons:
[0004] (1) SRAM memory limitation: The SRAM in the data plane is usually limited to dozens of megabytes. For example, the SRAM size of the Trio programmable chipset used in Juniper Networks' MX series routers is only 2 - 8 MB, which greatly limits the memory capacity of flow features;
[0005] (2) High-speed networks with increasing traffic: In modern high-speed network environments, the volume of traffic data passing through programmable switches is rising sharply. For example, in large-scale data center networks, switches usually process millions of data flows per second. As the number of rows in the flow feature table (representing the keys of network flows) grows rapidly, the available memory for storing flow feature statistics in the flow table becomes very limited;
[0006] To address these challenges, a solution that effectively performs network measurements while adhering to memory constraints must be designed. For example, for 10 MB of SRAM, measuring the features of more than 10,000 flows requires limiting the bit width of each flow's features to 512 bits. Since downstream network tasks require flow features such as traffic classification and anomaly detection, it is crucial to identify a feature set in the limited available SRAM that meets the performance requirements of downstream tasks. Summary of the Invention
[0007] In view of the above existing deficiencies, the present invention provides a method for feature selection and optimization of a network structure, which can effectively solve the problems involved in the above background art.
[0008] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0009] A method for feature selection and optimization of a network structure, including the following steps: constructing a task model and adding a selection factor layer to obtain a feature selection model;
[0010] Taking the obtained traffic data set as the input of the feature selection model, screening out K statistical features and their corresponding selection factor values according to the feature selection strategy to obtain pre-screened features;
[0011] Using the memory bit width length as a feature, screening out features that meet the memory bit width limit condition by using the dynamic programming algorithm, and combining them with the pre-screened features to obtain traffic features;
[0012] For the spatial distance between the traffic features and the set label features, obtaining the accuracy; according to the accuracy, adjusting the values of the selection factors in the selection factor layer, and issuing an instruction through the control plane to adjust the matching operation table of the data plane, so as to select more suitable traffic features; wherein, the spatial distance is calculated by the total loss function constructed by introducing a class balance term and a focal term in the cross-entropy loss function.
[0013] Furthermore, the selection factor layer is parameterized by a group of learnable parameters and each parameter corresponds to an updatable feature .
[0014] Furthermore, the feature selection strategy refers to introducing an algorithm based on a heap to identify the optimal K selection factors. The specific algorithm is as follows:
[0015] Construct a small root heap with the first K selection factors, where the heap top represents the smallest value;
[0016] Iterate through the remaining (N - K) selection factors. For each factor, if its value exceeds the value of the heap top element, replace the value of the heap top element and readjust the heap;
[0017] After iterating through all the factors, the elements in the heap are the Top-K selection factors.
[0018] Furthermore, the stability of the index set composed of the Top-K selection factors is set as the Top-K stability measure TSM, and its setting is as follows: ; where is the stability measure of the i-th epoch, Q is a hyperparameter specifying the consecutive epochs to be considered, and is the index set of the Top-K features of the (i - j)-th epoch; represent the intersection of the past Q epochs ; when = K, it indicates that has remained unchanged for Q consecutive epochs, that is, the early stopping mechanism is triggered.
[0019] Furthermore, the bit-width limit condition refers to the bit-width limit constraint of SRAM, and the dynamic programming algorithm used operates as follows:
[0020] Problem definition: W represents the total bit-width limit of SRAM, represents the bit-width of the candidate feature, represents the selection factor value of the candidate feature, represents whether the th feature is selected. The goal is to maximize the sum of the selection factor values while ensuring that the total bit-width does not exceed , and the problem is expressed as: ;
[0021] State definition: Set to represent the sum of the maximum selection factor values obtained by selecting from the first features with a total bit-width limit of ; set to store the index set of the selected features corresponding to ;
[0022] Initialization: Initialize all to 0, which means no items are selected and the total value is zero. Initialize to be empty, indicating that no indexes are recorded initially;
[0023] State transition: If the th feature is not selected:
[0024] ;
[0025] If the i-th feature is selected (j ≥ ):
[0026] ;
[0027] ;
[0028] Optimal evaluation: The optimal value is , and the index set corresponding to the selected items is .
[0029] Furthermore, the total loss function is expressed as:
[0030] ,
[0031] Among them, is the class balance term, which can adjust the weight of each class according to the sample size of the class ;
[0032] is the focus term, which can pay more attention to the sample classes that are difficult to predict.
[0033] In a second aspect, a computer program, when executed, implements each functional system or module included in the feature selection and optimization method of the above-mentioned network structure.
[0034] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0035] Adaptive flow statistical feature selection for downstream task models;
[0036] Feature selection strategy for selecting the best under limited memory resources;
[0037] Reduce computational overhead and quickly perform feature matching;
[0038] Solve the data imbalance problem for downstream tasks.
[0039] Other features and advantages of the present application will be described in the subsequent description, and some will become obvious from the description, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0041] Figure 1 Schematic flow diagram of a feature selection and optimization method for a network structure provided by the present invention;
[0042] Figure 2 Schematic architecture diagram of a programmable switch provided by the present invention;
[0043] Figure 3 Schematic diagram of a task-aware traffic telemetry framework based on a programmable switch provided by the present invention;
[0044] Figure 4 Schematic diagram of a task-sensitive feature selection module provided by the present invention;
[0045] Figure 5 Schematic diagram of the feature selection module under bit width limitation provided by the present invention;
[0046] Figure 6 Schematic diagram of the change of selection factors of the Internet of Things device identification dataset provided by the present invention;
[0047] Figure 7 Schematic diagram of the feature index provided by the present invention;
[0048] Figure 8 Schematic diagram of feature selection under bit width limitation provided by the present invention;
[0049] Figure 9 Schematic diagram of the recognition accuracy of different numbers of statistical features of the UNSW-IoT dataset under the NetMamba model;
[0050] Figure 10 Schematic diagram of the recognition accuracy of different numbers of statistical features of the UNSW-NB dataset under the NetMamba model;
[0051] Figure 11 Schematic diagram of the recognition accuracy of different numbers of statistical features of the NSL-KDD dataset under the NetMamba model. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0053] The present application provides a method for feature selection and optimization of a network structure, which adopts an innovative design. Through a feature pre-selection strategy, under the limitation of memory resources, the feature selection of the network structure is optimized, and an early stopping training mechanism is proposed to efficiently solve the training process.
[0054] The design concept of the embodiments of the present application will be briefly introduced below.
[0055] The basic idea of the present invention is to provide a feature selection method for a network structure. First, collect flow-level statistical features through the data plane of a programmable switch. After uploading them to the control plane, a selection factor layer is added between the complete flow-level statistical features and the downstream task model, and it is trained together with the neural network to screen out K key candidate statistical features. According to the bit width occupied by each statistical feature and the space that can be accommodated in the flow table of the programmable switch, the SRAM-constrained feature selection problem is transformed into a knapsack problem, and the dynamic programming method is used to solve it. Since the data imbalance problem widely exists in the traffic recognition task, based on the softmax loss function, the present invention designs a class-balanced loss function to strengthen the model's attention to minority-class samples, thereby alleviating the impact of data imbalance and improving the traffic recognition accuracy. Finally, two evaluation indicators are used to judge the accuracy of traffic recognition, evaluate the selected set of statistical features, and feedback the traffic recognition accuracy under the feature selection of this method to the control plane to better adjust the issuance of measurement instructions and determine the statistical features to be measured in the data plane.
[0056] The following will Figures 1 to 5 describe the SRAM-constrained data plane flow-level telemetry task-aware feature selection method provided by the embodiments of the present invention.
[0057] As Figure 1 shown, the present invention provides a feature selection and optimization method for a network structure, mainly including feature pre-selection sensitive to downstream tasks, feature selection under bit width constraints, and feedback and update of the selection factor based on the task recognition accuracy of the network structure to select more suitable traffic features.
[0058] In the specific implementation, first, the programmable switch collects flow-level statistical feature data. Figure 2 shows the architecture of a programmable switch that supports flow-level traffic telemetry technology. The incoming data packets are parsed and processed through multiple matching operations (MA), which can calculate measurement statistics. These operations generate various flow-level traffic features stored in the flow feature table in the SRAM of the data plane. These flow features are regularly uploaded to the control plane. By utilizing the flow-level features, Internet service providers (ISPs) can perform network management, such as traffic classification and anomaly detection.
[0059] In the specific implementation, Figure 3Shows a schematic diagram of the task-aware traffic telemetry framework based on a programmable switch. This framework integrates operations on the control plane and the data plane to achieve efficient and adaptive network traffic measurement. The role of the control plane is to configure the data plane and utilize the measured traffic characteristics for downstream tasks. Specifically, it identifies the necessary traffic characteristics according to the specific requirements of downstream tasks and generates corresponding Match-Action (MA) operations to calculate these characteristics. During task execution, the control plane periodically collects traffic characteristics from the data plane and applies them to tasks such as traffic classification and anomaly detection, realizing intelligent network management. While executing the measurement instructions issued by the control plane, the data plane processes and forwards data packets. These instructions are converted into a series of MA operations to complete packet parsing and pipeline processing in the programmable switch. When data packets pass through a pipeline composed of multiple MA operations, the data plane aggregates and calculates flow feature statistics according to a unique flow key (e.g., five-tuple or destination IP address). The calculated traffic characteristics are stored in a flow table in SRAM, and each row corresponds to a unique flow.
[0060] The proposed task-aware data plane measurement framework mainly consists of three steps:
[0061] 1) Key feature selection. In the control plane, an adaptive and efficient feature selection algorithm can identify the best set of traffic features suitable for the specific requirements of downstream tasks while complying with the SRAM bitwidth constraint;
[0062] 2) Deployment of measurement instructions. According to the selected features, the control plane generates and deploys the corresponding MA operations as measurement instructions to the data plane;
[0063] 3) Reporting and utilization of traffic functions. The data plane executes the configured MA operations in the programmable packet parsing and processing pipeline. When data packets pass through the pipeline, the data plane aggregates and calculates the required flow feature statistics using a unique flow key (e.g., five-tuple or destination IP address). These statistics are stored in SRAM and reported to the control plane regularly to support downstream tasks such as traffic classification and anomaly detection.
[0064] This framework has high adaptability, dynamically selects features optimized for specific downstream tasks, and at the same time follows strict SRAM bitwidth constraints. By combining task-aware feature selection with programmable telemetry technology, this framework solves two key challenges in data plane measurement: limited SRAM memory capacity and the need for adaptive traffic measurement in a high-speed network environment.
[0065] In specific implementation, as Figure 4 shown, given a task model for a network management task , the goal is to design a feature selection algorithm that identifies key flow features from the complete set of flow features while ensuring the performance of downstream tasks and meeting the memory constraints of SRAM. Here, represents the number of selected features, where and and . In addition, the selected features must satisfy the bit-width constraint , which limits the total memory required to store the flow features in the flow table.
[0066] By introducing a new neural network layer, an end-to-end feature selection framework called the selection factor layer is designed, which is trained together with the downstream task model. The framework consists of two main modules: a downstream task-sensitive feature pre-selection module and a feature selection module under bit-width constraints.
[0067] In specific implementation, in order to make the selected features adaptable to downstream tasks, the present invention proposes a brand-new neural network layer - the selection factor layer. This layer is located between the complete traffic features and the downstream task model and is co-trained with the task model. By assigning importance scores to the features, it can identify the features that are most important for task performance, thereby optimizing the model performance.
[0068] In specific implementation, as Figure 5 shown, the selection factor layer is parameterized by a set of learnable parameters , where each parameter corresponds to a feature . To ensure that the selected traffic features adapt to different downstream task models, is jointly trained with the task model. norm regularization is incorporated into the loss function of the task model to strengthen sparsity. The combined loss function is defined as:
[0069] ;
[0070] where, is the loss function of the downstream task model, represents the model parameters, is the regularization hyperparameter, is the complete set of traffic features, represents norm.
[0071] During the training process, norm regularization will force many values in to approach zero, thereby enhancing its sparsity. reflects the relative importance of the features:
[0072] • For 0: It means that the corresponding feature is redundant and makes no contribution to the downstream task;
[0073] • Small : It means that the corresponding feature has a negligible impact on the downstream task;
[0074] • Large : It means that the corresponding feature is crucial for the task and has a significant impact on the model performance.
[0075] Based on the sparsity, the present invention adopts the following two feature selection strategies: 1. Delete the features with the selection factor value ( ) being 0 or close to 0; 2. Sort the remaining features according to the value, and select the top features with the largest value. After applying this strategy, the obtained important candidate features are:
[0076] ;
[0077] Among them, represents the selected features, is the index set of the top features with the largest value.
[0078] Although the selection factor layer can effectively identify the importance of each feature for the downstream task, waiting for the neural network model to fully converge often requires a large amount of training time. Since what needs to be concerned about is not the exact numerical value, but the index set of the top features sorted according to , therefore, can be identified in advance before the model fully converges, thus accelerating the feature selection process.
[0079] The index set tends to be stable after only a few training iterations. In a specific implementation, as Figure 6 shows the change process of the selection factors of the largest and smallest 5 selection factors during the training of the Internet of Things device recognition dataset. It can be found that in the first few training cycles, the selection factors corresponding to the selected features are significantly higher than those of the unselected features, and the selection factors of the unselected features gradually converge to zero. This phenomenon indicates that can be determined in the early stage of training before γ fully converges,
[0080] The present invention designs an early stopping mechanism to determine whether it has stabilized, thereby accelerating the feature selection process.
[0081] In a specific implementation, the Top-K Stability Metric (TSM) is introduced to accelerate the feature selection process. If it remains unchanged over multiple consecutive training epochs, it can be considered stable and the training can be terminated early. The definition of TSM is as follows:
[0082] ;
[0083] where is the stability metric for the i-th epoch, Q is a hyperparameter specifying the number of consecutive epochs to consider, and is the index set of the Top-K features for the (i - j)-th epoch. represents taking the intersection of over the past Q epochs. For example, Figure 7 shows the feature indices for 6 epochs when K = 5 and Q = 4. Each row represents the indices of γ after sorting the selection factors for each epoch, where the Top-K indices are shown in red. The calculation methods for the TSM values of the 5th and 6th epochs are:
[0084] ;
[0085] When = K, it indicates that has remained unchanged for Q consecutive epochs, i.e., triggering the early stopping mechanism.
[0086] Through task-sensitive metric selection, key candidate features can be determined. However, due to the limited SRAM capacity of programmable switches, all candidate features may not be able to be stored. Therefore, further screening is required to maximize the utility of the selected features while satisfying the SRAM bit-width constraint. For this purpose, the feature selection problem is modeled as a 0-1 knapsack problem, and an efficient dynamic programming algorithm is designed to find the optimal feature subset.
[0087] In a specific implementation, as Figure 8 shown, the present invention models the feature selection problem under the SRAM bit-width constraint as a 0-1 knapsack problem. To effectively solve this knapsack problem, a dynamic programming algorithm is adopted, which specifically includes the following steps:
[0088] (1) Problem definition: Let W represent the total bit-width limit of the SRAM, Represents the bit width of the candidate feature, Represents the selection factor value of the candidate feature, Indicates whether the \(i\)-th feature is selected. The goal is to maximize the sum of the selection factor values while ensuring that the total bit width does not exceed \(W\). Specifically, it is expressed as:
[0089] ;
[0090] (2)State definition: Let Represent the maximum sum of the selection factor values obtained by selecting from the previous features with a total bit width limit of . Let Store the index set of the selected features corresponding to ;
[0091] (3)Initialization: Initialize all to 0, which means no items are selected and the total value is zero. Initialize to be empty, indicating that no indices are recorded initially;
[0092] (4)State transition: If the \(j\)-th feature is not selected:
[0093] ;
[0094] If the \(i\)-th feature is selected (\(j\geq \)):
[0095] ;
[0096] ;
[0097] (5)Optimal evaluation: The optimal value is , and the index set corresponding to the selected items is .
[0098] By using the dynamic programming algorithm, we obtain the maximum sum of the selection factor values ( ) under the bit width limit and the corresponding index set of the selected features ( ). The time complexity of this algorithm is , where \(K\) is the number of candidate features and \(W\) is the bit width limit.
[0099] In the traffic recognition task, data imbalance is a common problem, which not only reduces the overall task performance but also causes bias in the feature selection process. This is because the accuracy of the downstream task is the key metric guiding feature selection. In specific implementation, in the traffic recognition task of Internet of Things (IoT) devices, the traffic data generated by mobile phones and laptops is significantly more than that generated by printers or cameras. In the commonly used UNSWIoT dataset for IoT device recognition, there is a serious data imbalance problem in the traffic distribution. This dataset contains traffic from 30 IoT devices, where the top 4 devices account for more than 82% of the total traffic, while the total contribution of the remaining 26 devices is less than 18%. Specifically, more than 60% of the IoT devices generate less than 1% of the total traffic. This imbalance causes the model to overfit to devices with abundant traffic and underfit to devices with sparse traffic, reducing the overall accuracy of the model.
[0100] In the downstream task model, the softmax cross-entropy loss function is widely used in multi-classification tasks. Given the input data as the flow statistic sample X, the output of the model M(X) is usually the logits (unnormalized scores) for each class. These logits are combined into a vector z, and there is . z is transformed into a probability distribution through the softmax function as follows:
[0101] ;
[0102] where, is the logit for the i-th class, is the predicted probability for the i-th class, and t is the number of classes.
[0103] The goal of the softmax cross-entropy loss function is to minimize the difference between the predicted probability and the actual label. Given the actual label as y, which is a one-hot vector, the cross-entropy loss function is defined as:
[0104] ;
[0105] However, since the softmax cross-entropy loss function treats each sample equally, any data imbalance in the dataset will cause the classes with more training data to overfit and the classes with less training data to underfit, thus significantly reducing the performance of the model.
[0106] In the downstream task model, the softmax cross-entropy loss function is widely used in multi-classification tasks. Given the input data as the flow statistic sample X, the output of the model M(X) is usually the logits (unnormalized scores) for each class. These logits are combined into a vector z, and there is . z is transformed into a probability distribution through the softmax function as follows:
[0107] ;
[0108] Among them, is the logits of the i-th class, is the predicted probability of the i-th class, and t is the number of classes.
[0109] The goal of the softmax cross-entropy loss function is to minimize the difference between the predicted probability and the actual label. Given the actual label as y, which is a one-hot vector, the cross-entropy loss function is defined as:
[0110] ;
[0111] However, since the softmax cross-entropy loss function treats each sample equally, any data imbalance in the dataset will cause the classes with more training data to be overfitted and the classes with less training data to be underfitted, thus significantly reducing the performance of the model.
[0112] To solve the problem of data imbalance, the present invention designs a class-balanced loss function, which, based on the softmax cross-entropy loss function, assigns more weights to the minority-class samples, thus encouraging the model to pay more attention to them. That is, a class-balanced term and a focal term are introduced into the loss function. The definition of the total loss function is:
[0113] ;
[0114] Among them, is the class-balanced term, which can adjust the weight of each class according to the sample size of the class. is the hyperparameter that controls the influence degree of the class-balanced term, is the focal term, which focuses on the sample classes that are difficult to predict, is the hyperparameter that controls the influence degree of the focal term.
[0115] In specific implementation, β is set between 0.999 and 0.9999, is set to 2.
[0116] According to the said accuracy, adjust the value of the selection factor in the selection factor layer, and issue an instruction through the control plane to adjust the matching operation table of the data plane, so as to select more suitable traffic features.
[0117] In specific implementation, first calculate the predicted label of the downstream task model, and compare it with the true label to obtain the recognition accuracy of the downstream task model. By adjusting the value of the selection factor, make the downstream task model reach a higher recognition accuracy to optimize the traffic recognition ability of the statistical feature selection set.
[0118] In specific implementation, two common metrics are used to evaluate the accuracy of the traffic classification task: Macro-F1 and Micro-F1. The calculation of these two metrics depends on various metrics derived from the confusion matrix, which is a table used to describe the performance of a classification model on a set of test data with known true labels. It is usually organized into four quadrants:
[0119] • True Positive (TP): The number of samples correctly predicted for the current class;
[0120] • True Negative (TN): The number of negative instances correctly predicted as negative;
[0121] • False Positive (FP): The number of negative instances incorrectly predicted as positive;
[0122] • False Negative (FN): The number of positive instances incorrectly predicted as negative.
[0123] The calculation formulas for Macro-F1 and Micro-F1 are as follows:
[0124] ;
[0125] where, is the recall rate of the k-th class, is the precision of the k-th class, and t is the total number of class types. The applicable ranges of both Macro-F1 and Micro-F1 are [0, 1]. The larger Macro-F1 and Micro-F1 are, the higher the classification accuracy of the downstream task model.
[0126] The purpose of calculating the feedback accuracy of the downstream task model above is to timely adjust the change of the selection factor, and the evolution of the selection factor is adapted to different downstream tasks. After determining the most suitable features under the limited bit width, the control plane will issue measurement instructions to adjust the matching operation (MA) table of the data plane, and then collect statistical features that are more adapted to the downstream tasks.
[0127] A computer program, when executed, implements each functional system or module included in the feature selection and optimization method of the above network structure.
[0128] Experimental Results
[0129] The present invention is deployed on tensorflow and uses the Adam optimizer with an initial learning rate of 0.0001. During training, the batch size is set to 128, and the weight of L1 regularization is set to 0.001. The class-balanced focal loss is used to handle class imbalance and has a modulation term based on the prediction probability. The initial scale factor is set to a vector of all 1s.
[0130] The present invention uses three real-world datasets: UNSW-IoT, UNSW-NB, and NSL-KDD, and evaluates them on two traffic identification tasks, namely, Internet of Things (IoT) device traffic identification and anomaly detection. The descriptions of the datasets are as follows:
[0131] UNSW-IoT is a public dataset for IoT device identification. It contains approximately 430,000 traffic samples distributed across 29 categories, including Insteon cameras, Triby speakers, and HP printers, etc. We extracted 72 statistical features from the pcap files, including packet arrival time intervals, traffic duration, and congestion window sizes of upstream and downstream traffic.
[0132] UNSW-NB is a public dataset for anomaly detection, containing approximately 200,000 traffic samples divided into normal and abnormal categories. There are 42 statistical features such as protocols, services, and TTL.
[0133] NSL-KDD is a public dataset for anomaly detection, containing approximately 150,000 traffic samples divided into normal and abnormal categories. There are mainly 41 statistical features such as error rate, duration, and service time.
[0134] In the experiment, five commonly used neural network models were adopted as the backbone models. All models were implemented using the TensorFlow [1] framework and trained on a single NVIDIA GeForce RTX 3090 GPU. A brief introduction to the five models is as follows:
[0135] VGG: A simple and effective convolutional neural network, consisting of a series of convolutional layers and max-pooling layers, followed by fully connected layers for output;
[0136] ResNet: A classic residual network that alleviates the degradation problem in deep neural networks by introducing skip (residual) connections;
[0137] ET-BERT (Encrypted Traffic Identification BERT): A transformer-based encrypted traffic classification method that utilizes a pre-trained transformer model and a multi-layer attention mechanism;
[0138] NetMamba: An efficient network traffic classification model that uses a pre-trained unidirectional Mamba architecture to address challenges in model efficiency and traffic representation;
[0139] AE (Autoencoder): A neural network designed for unsupervised learning that detects anomalies by reconstructing input data.
[0140] In specific implementation, such as Figure 9For the recognition accuracy of extracting different numbers (K) of statistical features from the UNSW-IoT dataset through the present invention under the NetMamba model, it can be seen that when 24 statistical features are selected, the model reaches the highest accuracy, which is 0.13% higher than the accuracy of the complete 72 statistical features; Figure 10 For the recognition accuracy of extracting different numbers of statistical features from the UNSW-NB dataset through the present invention under the NetMamba model, it can be seen that when 32 statistical features are selected, the model reaches the highest accuracy, which is 1.80% higher than the accuracy of the complete 42 statistical features; Figure 11 For the recognition accuracy of extracting different numbers of statistical features from the NSL-KDD dataset through the present invention under the NetMamba model, it can be seen that when 4 statistical features are selected, the model reaches the highest accuracy, which is 3.71% higher than the accuracy of the complete 41 statistical features.
[0141] When the computer program is executed, it implements each functional system or module included in the feature selection and optimization method of the above network structure.
[0142] Specifically, the experimental results are shown in Table 1. The effectiveness of this feature selection method for the accuracy of downstream tasks under different memory limitations is compared through experiments. First, we use a neural network model to evaluate all tasks with all features, which is set to "all". Secondly, the pre-selection module of the present invention is applied to select the most important K features, and the task is evaluated again, denoted as "K". Finally, we test the selected features under SRAM memory constraints of 512bit and 256bit respectively. For comparison, two classic machine learning methods for traffic recognition tasks, AppScanner and KNN, are also compared in the experiment. AppScanner is a traffic recognition method that detects and identifies application programs by analyzing network traffic features. KNN is a widely used classification method that calculates the distance between the sample to be classified and the training samples, selects the K samples with the closest distance, and then determines the category of the sample to be classified according to the categories of these K samples. It can be found from the experimental results in Table 1 that:
[0143] (1) The feature pre-selection module sensitive to downstream tasks can identify the key K features, enabling the model trained on these features to achieve comparable or higher accuracy. As shown in the results of the UNSW-IoT dataset in Table 1, by selecting the important K features, most models obtained higher accuracy than the models using the complete feature set, including ResNet, VGG, AE, and NetMamba models. At the same time, the ET-BERT model obtained accuracy comparable to that of the model with all features;
[0144] (2)The feature pre-selection module sensitive to downstream tasks identifies and eliminates feature redundancy, significantly reducing computational and memory costs. As shown in Table 1, in the traffic anomaly detection task on the NSL-KDD dataset, the selected feature subset accounts for 10%-78% of the total 41 features;
[0145] (3)Even under memory limitations of 256bit and 512bit, the feature selection optimization method of the present invention maintains high accuracy, demonstrating the effectiveness of the method in SRAM memory-constrained environments. For example, for the IoT device identification task of the UNSW-IoT dataset shown in Table 1, under 512-bit and 256-bit width constraints, most models achieve accuracy comparable to that of the model using full features, and the NetMamba model even outperforms the model using full features.
[0146] Table 1 Comparison of two metric accuracies of different traffic recognition models in 3 datasets:
[0147]
[0148] As shown in Table 2, the impact of the early stopping mechanism on feature selection was evaluated through experiments. Specifically, the impact of early stopping on training time and model accuracy was analyzed through experiments. Five backbone neural networks were used for experiments in the UNSW-IoT dataset, where No indicates not using the early stopping mechanism, Yes indicates using the early stopping mechanism, and the bold indices in the selection factor Top-K index set are the differences between the index set without the early stopping mechanism and the index set with the early stopping mechanism after using the early stopping mechanism. The experimental summary is as follows:
[0149] (1)After adding the early stopping mechanism, the training time of the NetMamba model decreased most significantly, reducing by 83.33% (100 / 120 epochs) compared to the case without the early stopping mechanism. In contrast, the reduction of ET-BERT was the smallest, only reducing by 40% (4 / 10 epochs). Therefore, after adding the early stopping mechanism, compared to the case without the early stopping mechanism, the training time of the model decreased by 40% to 83.33%, demonstrating the effectiveness of this mechanism in accelerating feature selection;
[0150] (2)The accuracy of the model trained with early stopping is comparable to that of the model trained to convergence. This indicates that the features most influential on the accuracy of downstream tasks are determined early in the training process;
[0151] (3)The selected feature sets have only slight differences between full training and early stopping, which confirms that important candidate features can remain stable during the training process.
[0152] Table 2 Impact of the early stopping mechanism in the UNSW-IoT dataset:
[0153]
[0154] As shown in Table 3, the impact of evaluating the class balance (CB) loss function on four downstream task models (ResNet, VGG, ET-BERT, and NetMamba) in the UNSW-IoT dataset is presented. AE is excluded because the autoencoder-based neural network model does not support the CB loss function. The experiments show that using the class balance loss function improves the Micro-F1 and Macro-F1 scores in all models, confirming the effectiveness of CB loss in addressing class imbalance and improving model performance.
[0155] Table 3 Comparison of accuracies with and without the class balance (CB) loss function configured:
[0156]
[0157] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0158] (1) Adaptive flow statistic feature selection for downstream task models. To achieve adaptive selection of statistical features for different task models, the present invention designs a selection factor layer located between the task model and the statistical features. The selection factor evolves continuously during the training process of the task model and gradually stabilizes. After the selection factor stabilizes, its magnitude can be used as an indicator of feature importance. Based on this indicator, the K most critical candidate features of the task model can be identified, and redundant features can be removed;
[0159] (2) Feature selection under bit-width constraints. To ensure the accuracy of downstream tasks and select the most critical features without exceeding the SRAM memory limit, the present invention transforms the original feature selection problem into a feature knapsack problem according to the value of the selection factor and its corresponding bit-width requirements. In addition, a dynamic programming-based feature selection algorithm is designed to solve this feature knapsack problem;
[0160] (3) Fast feature selection algorithm. To reduce the computational overhead and accelerate the feature selection process, the present invention first conducts a large number of experiments and finds that the largest K selection factors can be quickly determined before the model converges. Based on this, the present invention designs an early stopping metric to monitor the convergence of the selection factors and proposes an early stopping mechanism. This mechanism can quickly complete the feature selection after the selection factors converge, rather than waiting for the model training to complete, thereby accelerating the feature selection process;
[0161] (4)Data imbalance processing. Since the data plane measurement framework proposed in the present invention is task-driven, two downstream tasks are implemented on the control plane, including Internet of Things device identification and anomaly detection. In these tasks, there is a common problem (i.e., data imbalance). However, during the feature selection process, the accuracy of the downstream task model is usually required as the measurement standard. Therefore, it is crucial to ensure the fairness of the flow statistical feature selection process in solving the problem of network flow data imbalance. To solve the problem of traffic data imbalance, the present invention designs a class balance loss function. By assigning different weighted losses to different network flow data, it effectively prevents the downstream task model from overfitting to the network flow data class with a larger number of samples, thereby reducing the impact of data imbalance on the model performance.
[0162] The above content is only an example and illustration of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, as long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should fall within the protection scope of the present invention.
Claims
1. A feature selection and optimization method for a network structure, characterized in that including the following methods: Construct a task model and add a selection factor layer to obtain a feature selection model; Use the obtained traffic dataset as the input of the feature selection model, and screen out K statistical features and their corresponding selection factor values according to the feature selection strategy to obtain pre-screened features; Use the memory bit width length as a feature and adopt a dynamic programming algorithm to screen out features that meet the memory bit width limit condition, and combine them with the pre-screened features to obtain traffic features; Obtain the accuracy for the spatial distance between the traffic features and the set label features; according to the accuracy, adjust the value of the selection factor in the selection factor layer, and issue an instruction through the control plane to adjust the matching operation table of the data plane, so as to select more suitable traffic features; wherein, the spatial distance is calculated by a total loss function constructed by introducing a class balance term and a focal term in the cross-entropy loss function; the total loss function is expressed as: ; Among them, is the class balance term, which can adjust the weight of each class according to the sample size of the class ; is the focus item, which can pay more attention to difficult-to-predict sample classes; is a hyperparameter that controls the influence degree of the class balance item, is a hyperparameter that controls the influence degree of the focus item, and the actual label is y, is the predicted probability of the i-th class.
2. The feature selection and optimization method for the network structure according to claim 1, characterized in that The described selection factor layer consists of a set of learnable parameters parameterized, and each of these parameters corresponds to an updatable feature .
3. The feature selection and optimization method for the network structure according to claim 1, characterized in that, The feature selection strategy mentioned above refers to introducing a heap-based algorithm to identify the optimal K selection factors. The specific algorithm is as follows: Construct a min heap with the first K selection factors, where the heap top represents the minimum value; Iterate the remaining (N-K) selection factors; for each factor, if its value exceeds the value of the heap top element, replace the value of the top element and readjust the heap; After iterating all the factors, the elements in the heap are the Top-K selection factors.
4. The feature selection and optimization method for the network structure according to claim 3, characterized in that, The index set composed of the Top-K selection factors is set to the Top-K stability metric TSM, which is set as follows: ; where is the stability metric for the 𝑖-th epoch, Q is a hyperparameter specifying the number of consecutive epochs to consider, and is the index set of the Top-K features for the -th epoch; represents taking the intersection of the of the past Q epochs; when = 𝐾, it indicates that has remained unchanged for Q consecutive epochs, i.e., triggering the early stopping mechanism.
5. The feature selection and optimization method for the network structure according to claim 1, characterized in that The bit width limit condition mentioned above refers to the bit width limit constraint of SRAM, and the dynamic programming algorithm adopted operates as follows: Problem Definition: Represents the total bit width limit of the SRAM, Represents the bit width of the candidate feature, Represents the selection factor value of the candidate feature, Represents whether the i-th feature is selected. The goal is to maximize the sum of the selection factor values while ensuring that the total bit width does not exceed , and the problem is expressed as: ; Status definition: Set Indicates selecting from the previous features, with the total bit width limited to the sum of the maximum selection factor values obtained, set to store the index set of the selected features corresponding to ; Initialization: Set all to 0, indicating unassigned and with a total value of zero; set to empty, indicating no indexes are initially recorded; State transition: If the i-th feature is not selected: ; If the i-th feature is selected (j ≥ ): ; ; Optimized evaluation: The optimal value is , and the index set corresponding to the selected option is .
6. A computer program, characterized in that, When the computer program is executed, it realizes each functional system or module included in the feature selection and optimization method of the network structure according to any one of claims 1 to 5.
Citation Information
Patent Citations
Traffic feature extraction method and system, storage medium and electronic equipment
CN114024758A
Neural architecture searching method for chip design
CN118821863A