Feature selection and optimization method of network structure
By introducing a selection factor layer and dynamic programming algorithm into the network structure, combining a class-balanced loss function, and optimizing feature selection, the memory limitation and data imbalance problems of programmable switches collecting stream-level traffic characteristics in high-speed networks are solved, and efficient feature selection and traffic recognition are achieved.
Patent Information
- Application Number
- CN202510482704.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In high-speed network environments, programmable switches face SRAM memory limitations and data imbalances when collecting stream-level traffic characteristics, making it difficult to identify feature sets that meet the performance requirements of downstream tasks.
A feature selection and optimization method for network structure is adopted. By constructing a task model and adding a selection factor layer, K statistical features and their corresponding selection factor values are selected, and combined with dynamic programming algorithms and class balance loss functions, feature selection is optimized to adapt to downstream tasks and memory bit width limitations.
It realizes adaptive selection of flow statistical features suitable for downstream tasks under finite memory resources, reduces calculation overhead, improves traffic recognition accuracy, and alleviates data imbalance problem.
Smart Images

Figure CN120017508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network model optimization, and in particular to a method for feature selection and optimization of a network structure. Background Art
[0002] Network telemetry systems provide continuous and real-time measurements of network status, which is critical for network operators to detect increasingly complex events ranging from performance degradation to security attacks. Thanks to the rapid development of programmable switches, modern data plane telemetry systems increasingly leverage these devices, achieving scalability in handling high-speed traffic and achieving fast and efficient flow-level feature statistics.
[0003] In network management, programmable switches that support flow-level telemetry technology are usually used to count flow-level features and store various flow-level traffic features in the flow feature table in the SRAM of the data plane. The flow feature table is regularly uploaded to the control plane to guide the judgment of downstream tasks. However, in high-speed networks, it is still challenging for programmable switches to collect flow-level traffic features for the following reasons: (1) SRAM memory limitation: The SRAM in the data plane is usually limited to tens of megabytes [4]. For example, the SRAM size of the Trio programmable chipset used in Juniper Networks' MX series routers is only 2-8MB, which greatly limits the memory capacity of flow features; (2) High-speed networks with increasing traffic: In modern high-speed network environments, the amount of traffic data passing through programmable switches is increasing dramatically. For example, in large-scale data center networks, switches typically process millions of data flows per second. As the number of rows in the flow feature statistics table (the key representing the network flow) grows rapidly, the available memory in the flow table for storing flow feature statistics becomes very limited.
[0004] To address these challenges, a solution must be designed to efficiently perform network measurements while adhering to memory constraints. For example, with 10 MB of SRAM, measuring features for more than 10,000 flows requires limiting the bit width of each flow feature to 512 bits. Since flow features are required for downstream network tasks such as traffic classification and anomaly detection, it is critical to identify a feature set that meets the performance requirements of downstream tasks within the limited available SRAM. Summary of the invention
[0005] In view of the above-mentioned shortcomings currently existing, the present invention provides a feature selection and optimization method of a network structure, which can effectively solve the problems involved in the above-mentioned background technology.
[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions: A method for feature selection and optimization of a network structure, comprising the following methods: constructing a task model and adding a selection factor layer to obtain a feature selection model; The acquired traffic data set is used as the input of the feature selection model, and K statistical features and their corresponding selection factor values are screened out according to the feature selection strategy to obtain the pre-screening features; Taking the memory bit width length as a feature, a dynamic programming algorithm is used to screen out features that meet the memory bit width restriction condition, and the features are combined with the pre-screened features to obtain traffic features; The accuracy is obtained for the spatial distance between the traffic feature and the set label feature; according to the accuracy, the value of the selection factor in the selection factor layer is adjusted, and an instruction is issued through the control plane to adjust the matching operation table of the data plane, so as to select a more suitable traffic feature; wherein the spatial distance is calculated by the total loss function constructed by introducing the class balance term and the focus term in the cross entropy loss function.
[0007] Furthermore, the selection factor layer consists of a set of learnable parameters parameterized, where each parameter Corresponding to an updateable feature .
[0008] Furthermore, the feature selection strategy refers to introducing a heap-based algorithm to identify the optimal K selection factors. The specific algorithm is as follows: Use the first K selection factors to construct a small root heap, where the top of the heap represents the smallest value; Iterate the remaining (NK) selection factors. For each factor, if its value exceeds the value of the top element of the heap, replace the value of the top element and re-adjust the heap. After iterating over all factors, the elements in the heap are the Top-K selection factors.
[0009] Furthermore, the index set composed of the Top-K selection factors The stability of is set as the Top-K stability metric TSM, which is set as follows: ;in, is the stability measure for the i-th epoch, Q is a hyperparameter that specifies the number of consecutive epochs to consider, and is the index set of Top-K features of the (ij)th epoch; Represents the past Q epochs The intersection of When The value has remained unchanged for Q consecutive epochs, triggering the early stopping mechanism.
[0010] Further, the bit width restriction condition refers to the bit width restriction constraint of the SRAM, and the dynamic programming algorithm adopted is operated as follows: Problem definition: Indicates the total bit width limit of SRAM, represents the bit width of the candidate feature, represents the selection factor value of the candidate feature, Indicates whether the The goal is to maximize the sum of the selected factor values while ensuring that the total bit width does not exceed , the problem is expressed as: ; State definition: set Indicates that in the past feature selection, the total bit width is limited to The sum of the maximum selectivity factor values obtained; set Storage and The corresponding index set of the selected features; Initialization: Set all Initialized to 0, which means no items are selected and the total value is zero. Initialized to empty, indicating that no index is recorded initially; State transition: If you do not select Features: , If we select the i-th feature (j≥ ): ; ; Optimization evaluation: The optimal value is , the index set corresponding to the selected option is .
[0011] Further, the total loss function is expressed as: ; in, It is a class balancing item that can be used to balance the sample size of each class. Adjust the weight of each class; It is a focus item that can pay more attention to the sample classes that are difficult to predict.
[0012] In a second aspect, a computer program is provided, which, when executed, implements the various functional systems or modules included in the above-mentioned network structure feature selection and optimization method.
[0013] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: Stream statistical feature selection that is adaptive to downstream task models; The optimal feature selection strategy under limited memory resources; Reduce computational overhead and quickly perform feature matching; Solve the data imbalance problem for downstream tasks.
[0014] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0016] Figure 1 A schematic diagram of a flow chart of a feature selection and optimization method for a network structure provided by the present invention; Figure 2 A schematic diagram of the architecture of a programmable switch provided by the present invention; Figure 3 A schematic diagram of a task-aware traffic telemetry framework based on a programmable switch provided by the present invention; Figure 4 A schematic diagram of a task-sensitive feature selection module provided by the present invention; Figure 5 A schematic diagram of a feature selection module under bit width restriction provided by the present invention; Figure 6 A schematic diagram of changes in selection factors of an IoT device identification data set provided by the present invention; Figure 7 A schematic diagram of a feature index provided by the present invention; Figure 8 A schematic diagram of feature selection under bit width restriction provided by the present invention; Fig. 9 Schematic diagram of recognition accuracy of different quantitative statistical features of UNSW-IoT dataset under NetMamba model; Fig.10 Schematic diagram of recognition accuracy of different quantitative statistical features of UNSW-NB dataset under NetMamba model; Fig.11Schematic diagram of the recognition accuracy of different quantitative statistical features of the NSL-KDD dataset under the NetMamba model. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The present application provides a feature selection and optimization method for a network structure, adopts an innovative design, optimizes the feature selection of the network structure under the limitation of memory resources through a feature pre-screening strategy, and proposes an early stopping training mechanism to efficiently solve the training process.
[0019] The following is a brief introduction to the design concept of the embodiments of the present application.
[0020] The basic idea of the present invention is to provide a feature selection method for a network structure. First, flow-level statistical features are collected through the data plane of a programmable switch, and after being uploaded to the control plane, a selection factor layer is added between the complete flow-level statistical features and the downstream task model, so that it is trained together with the neural network, and then K key candidate statistical features are screened out; according to the bit width occupied by each statistical feature and the space that can be accommodated in the programmable switch flow table, the SRAM-limited feature selection problem is converted into a knapsack problem, and a dynamic programming method is used to solve it; because the data imbalance problem is widely present in the traffic identification task, the present invention designs a class balance loss function based on the softmax loss function, strengthens the model's attention to minority class samples, and then alleviates the impact of data imbalance and improves the accuracy of traffic identification. Finally, two evaluation indicators are used to judge the accuracy of traffic identification, and the selected statistical feature selection set is evaluated. The traffic identification accuracy under the feature selection of this method is fed back to the control plane to better adjust the issuance of measurement instructions and determine the statistical features to be measured in the data plane.
[0021] Combine the following Figures 1 to 5 The present invention describes a method for selecting stream-level telemetry task-aware features in a SRAM-constrained data plane provided in an embodiment of the present invention.
[0022] like Figure 1As shown, the present invention provides a feature selection and optimization method for a network structure, which mainly includes pre-selection of features sensitive to downstream tasks, feature selection under bit width constraints, and updating of selection factors based on feedback of task recognition accuracy of the network structure to select more suitable traffic features.
[0023] In the specific implementation, the programmable switch first collects flow-level statistical feature data. Figure 2 The architecture of a programmable switch that supports flow-level traffic telemetry is presented. Incoming packets are parsed and processed through multiple matching operations (MAs) that can calculate measurement statistics. These operations generate various flow-level traffic features stored in the flow feature table in the SRAM of the data plane. These flow features are uploaded to the control plane periodically. By leveraging flow-level features, Internet Service Providers (ISPs) can perform network management such as traffic classification and anomaly detection.
[0024] In the specific implementation, Figure 3 The present invention shows a schematic diagram of a task-aware traffic telemetry framework based on a programmable switch. The framework integrates operations on the control plane and the data plane to achieve efficient and adaptive network traffic measurement. The role of the control plane is to configure the data plane and use the measured traffic features for downstream tasks. Specifically, it identifies the necessary traffic features according to the specific requirements of the downstream tasks and generates corresponding match action (MA) operations to calculate these features. During task execution, the control plane regularly collects traffic features from the data plane and applies them to tasks such as traffic classification and anomaly detection, thereby realizing intelligent network management. The data plane processes and forwards data packets while executing measurement instructions issued by the control plane. These instructions are converted into a series of MA operations to complete packet parsing and pipeline processing in the programmable switch. When a data packet passes through a pipeline consisting of multiple MA operations, the data plane aggregates and calculates flow feature statistics based on a unique flow key (e.g., a five-tuple or a destination IP address). The calculated traffic features are stored in a flow table in SRAM, with each row corresponding to a unique flow.
[0025] The proposed task-aware data plane measurement framework is mainly divided into three steps: 1) Key feature selection. In the control plane, an adaptive and efficient feature selection algorithm can identify the best set of traffic features that suits the specific requirements of downstream tasks while complying with SRAM bit width constraints; 2) Deploy measurement instructions. Based on the selected features, the control plane generates and deploys the corresponding MA operations as measurement instructions on the data plane; 3) Reporting and exploiting flow features. The data plane performs configured MA operations in a programmable packet parsing and processing pipeline. As packets traverse the pipeline, the data plane aggregates and computes the required flow feature statistics using unique flow keys (e.g., quintuples or destination IP addresses). These statistics are stored in SRAM and periodically reported to the control plane to support downstream tasks such as traffic classification and anomaly detection.
[0026] The framework is highly adaptive and dynamically selects features optimized for specific downstream tasks while adhering to strict SRAM bit width constraints. By combining task-aware feature selection with programmable telemetry, the framework addresses two key challenges in data plane measurement: limited SRAM memory capacity and the need for adaptive flow measurement in high-speed network environments.
[0027] In specific implementation, Figure 4 As shown, given a network management task Task Model , the goal is to design a feature selection algorithm that extracts Identify key flow features in , while ensuring the performance of downstream tasks and meeting the memory constraints of SRAM. Here, represents the number of selected features, where and In addition, the selected features must satisfy the bit width constraint , which limits the total memory required to store flow features in the flow table.
[0028] An end-to-end feature selection framework is designed by introducing a new neural network layer, called the selection factor layer, which is trained together with the downstream task model. The framework consists of two main modules: a downstream task-sensitive feature pre-selection module and a feature selection module under bitwidth constraints.
[0029] In the specific implementation, in order to make the selected features adaptive to the downstream tasks, the present invention proposes a new neural network layer - the selection factor layer. This layer is located between the complete traffic features and the downstream task model, and is trained in conjunction with the task model. By assigning importance scores to features, it can identify the features that are most important to task performance, thereby optimizing model performance.
[0030] In specific implementation, Figure 5 As shown, the selection factor layer consists of a set of learnable parameters Parameterized, where each parameter Corresponding to a feature In order to ensure that the selected traffic features are suitable for different downstream task models, Joint training with the task model. Norm regularization is incorporated into the loss function of the task model to strengthen The sparsity of the combined loss function is defined as: ; in, is the loss function of the downstream task model, represents the model parameters, is the regularization hyperparameter, is the complete set of traffic features, express Norm.
[0031] During the training process, Norm regularization forces Many values in are close to zero, thus enhancing its sparsity. Reflects the relative importance of features: • 0: indicates the corresponding feature It is redundant and does not contribute to downstream tasks; • Small : Indicates the corresponding features The impact on downstream tasks is negligible; •Big : Indicates the corresponding features It is crucial to the task and has a significant impact on the performance of the model.
[0032] based on To address the sparsity of feature selection, the present invention adopts the following two feature selection strategies: 1. Delete the selection factor value ( ) is 0 or close to 0; 2. For the remaining features, Sort the values and select the one with the largest value. After applying this strategy, we get The important candidate features are: ; in, represents the selected features, yes The largest value The index set of features.
[0033] Although the selection factor layer can effectively identify the importance of each feature to the downstream task, waiting for the neural network model to fully converge often requires a lot of training time. The exact value of The largest after sorting The index set of features , so it is possible to identify in advance before the model fully converges , thereby accelerating the feature selection process.
[0034] Index Set After only a few training iterations, it becomes stable. Figure 6 The change process of the selection factors of the IoT device recognition dataset during the training process is shown. It can be found that in the first few training cycles, the selection factors corresponding to the selected features are significantly higher than the selection factors of the unselected features, and the selection factors of the unselected features gradually converge to zero. This phenomenon shows that before γ fully converges, can be determined early in training, thus speeding up the feature selection process.
[0035] The present invention designs an early stopping mechanism to judge Is it stable, thus speeding up the feature selection process.
[0036] In the specific implementation, the Top-K Stability Metric (TSM) is introduced to speed up the feature selection process. If it remains unchanged in multiple consecutive training cycles, it can be considered to be stable and the training is terminated early. The definition of TSM is as follows: ; in, is the stability measure of the i-th epoch, Q is a hyperparameter that specifies the number of consecutive epochs to consider, and It is the index set of Top-K features of the (ij)th epoch. Represents the past Q epochs For example, Figure 7 The feature index of 6 epochs is shown when K=5 and Q=4. Each row represents the index of γ after sorting the selection factor of each epoch, where the Top-K index Shown in red. The calculation method of TSM value of the 5th epoch and the 6th epoch is: ; ; when When The value has remained unchanged for Q consecutive epochs, triggering the early stopping mechanism.
[0037] By selecting task-sensitive indicators, it is possible to determine However, due to the limited SRAM capacity of programmable switches, It may be impossible to store all candidate features. Therefore, further screening is required to maximize the utility of the selected features while satisfying the SRAM bit width constraint. To this end, the feature selection problem is modeled as a 0-1 knapsack problem, and an efficient dynamic programming algorithm is designed to find the optimal feature subset.
[0038] In specific implementation, Figure 8 As shown, the present invention models the feature selection problem under the SRAM bit width constraint as a 0-1 knapsack problem. In order to effectively solve this knapsack problem, a dynamic programming algorithm is used, which specifically includes the following steps: (1) Problem definition: Let W represent the total bit width limit of SRAM. represents the bit width of the candidate feature, represents the selection factor value of the candidate feature, Indicates whether the i-th feature is selected. The goal is to maximize the sum of the selection factor values while ensuring that the total bit width does not exceed W. Specifically expressed as: ; (2) State definition: Indicates that in the past feature selection, the total bit width is limited to The sum of the maximum selectivity factor values obtained. Let Storage and The corresponding index set of the selected features; (3) Initialization: Set all Initialized to 0, which means no items are selected and the total value is zero. Initialized to empty, indicating that no index is recorded initially; (4) State transition: If you do not select Features: ; If we select the i-th feature (j≥ ): ; ; (5) Optimization evaluation: The optimal value is , the index set corresponding to the selected option is .
[0039] By using a dynamic programming algorithm, we obtain the maximum sum of selectivity factor values under the bit width constraint ( ) and the corresponding index set of selected features ( ). The time complexity of this algorithm is , where K is the number of candidate features and W is the bit width limit.
[0040] In traffic identification tasks, data imbalance is a common problem that not only degrades the overall task performance but also biases the feature selection process. This is because the accuracy of downstream tasks is a key metric to guide feature selection. In specific implementations, in the traffic identification task of IoT devices, the traffic data generated by mobile phones and laptops is significantly more than that generated by printers or cameras. In the commonly used UNSWIoT dataset for IoT device identification, there is a serious data imbalance problem in traffic distribution. The dataset contains traffic from 30 IoT devices, of which the top 4 devices account for more than 82% of the total traffic, while the remaining 26 devices contribute less than 18% in total. Specifically, more than 60% of IoT devices generate less than 1% of the total traffic. This imbalance causes the model to overfit to devices with abundant traffic and overfit to devices with sparse traffic, reducing the overall accuracy of the model.
[0041] In downstream task models, the softmax cross entropy loss function is widely used in multi-classification tasks. Given the input data as a stream statistic sample X, the output of the model M(X) is usually the logits (unnormalized scores) for each category. These logits are combined into a vector z, . z is converted into a probability distribution through the softmax function as follows: ; in, is the logits of the i-th class, is the predicted probability of class i, and t is the number of classes.
[0042] The goal of the softmax cross entropy loss function is to minimize the difference between the predicted probability and the actual label. Given the actual label y, which is a one-hot vector, the cross entropy loss function is defined as: ; However, since the softmax cross entropy loss function treats each sample equally, any data imbalance in the dataset will cause the categories with more training data to overfit and the categories with less training data to underfit, thus significantly reducing the performance of the model.
[0043] In downstream task models, the softmax cross entropy loss function is widely used in multi-classification tasks. Given the input data as a stream statistic sample X, the output of the model M(X) is usually the logits (unnormalized scores) for each category. These logits are combined into a vector z, . z is converted into a probability distribution through the softmax function as follows: ; in, is the logits of the i-th category, is the predicted probability of class i, and t is the number of classes.
[0044] The goal of the softmax cross entropy loss function is to minimize the difference between the predicted probability and the actual label. Given the actual label y, which is a one-hot vector, the cross entropy loss function is defined as: ; However, since the softmax cross entropy loss function treats each sample equally, any data imbalance in the dataset will cause the categories with more training data to overfit and the categories with less training data to underfit, thus significantly reducing the performance of the model.
[0045] In order to solve the problem of data imbalance, the present invention designs a class-balanced loss function, which assigns more weights to minority class samples based on the softmax cross entropy loss function, thereby encouraging the model to pay more attention to them. That is, a class-balanced term and a focus term are introduced into the loss function. The total loss function is defined as: ; in, It is a class balancing term that can be used to balance the sample size of each class. Adjust the weight of each class. is a hyperparameter that controls the degree of influence of the class balance term, is the focus item, focusing on the sample class that is difficult to predict, is a hyperparameter that controls the influence of the focal term.
[0046] In the specific implementation, β is set between 0.999 and 0.9999. Set to 2.
[0047] According to the accuracy, the value of the selection factor in the selection factor layer is adjusted, and instructions are sent through the control plane to adjust the matching operation table of the data plane, so as to select more suitable traffic characteristics.
[0048] In the specific implementation, the predicted label of the downstream task model is first calculated and compared with the true label to obtain the recognition accuracy of the downstream task model. By adjusting the value of the selection factor, the downstream task model can achieve a higher recognition accuracy to optimize the traffic recognition capability of the statistical feature selection set.
[0049] In the specific implementation, two commonly used indicators are used to evaluate the accuracy of the traffic classification task: Macro-F1 and Micro-F1. The calculation of these two indicators relies on various indicators derived from the confusion matrix. The confusion matrix is a table used to describe the performance of a classification model on a set of test data with known true labels. It is usually organized into four quadrants: •True Positives (TP): the number of samples that correctly predict the current category; • True Negatives (TN): the number of negative instances correctly predicted as negative; • False Positives (FP): the number of negative instances that were incorrectly predicted as positive; • False Negatives (FN): the number of positive instances that are incorrectly predicted as negative; The calculation formulas for Macro-F1 and Micro-F1 are as follows: ; in, is the recall rate of the kth class, is the accuracy of the kth class, and t is the total number of class types. The applicable range of Macro-F1 and Micro-F1 is [0,1]. The larger the Macro-F1 and Micro-F1 are, the higher the classification accuracy of the downstream task model.
[0050] The purpose of calculating the feedback accuracy of the downstream task model is to timely adjust the changes in the selection factors so that the evolution of the selection factors is suitable for different downstream tasks. After determining the most suitable features under the limited bit width, the control plane will issue measurement instructions to adjust the matching operation (MA) table of the data plane, and then collect statistical features that are more suitable for downstream tasks.
[0051] A computer program, when executed, realizes each functional system or module included in the above-mentioned network structure feature selection and optimization method.
[0052] Experimental Results The invention is deployed on tensorflow and uses the Adam optimizer with an initial learning rate of 0.0001. During training, the batch size is set to 128 and the weight of L1 regularization is set to 0.001. The class-balanced focal loss is used to handle class imbalance with a modulation term based on the predicted probability. The initial scaling factor is set to an all-one vector.
[0053] This paper uses three real datasets: UNSW-IoT, UNSW-NB and NSL-KDD, and evaluates them on two traffic identification tasks (i.e., IoT device traffic identification and anomaly detection). The description of the datasets is as follows: UNSW-IoT is a public dataset for IoT device identification. It contains about 430,000 traffic samples distributed in 29 categories, including Insteon cameras, Triby speakers, and HP printers. We extracted 72 statistical features from the pcap files, including packet arrival time, traffic duration, and congestion window size of upstream and downstream traffic.
[0054] UNSW-NB is a public dataset for anomaly detection, containing about 200,000 flow samples, divided into normal and abnormal categories, including 42 statistical features such as protocol, service, and TTL.
[0055] NSL-KDD is a public dataset for anomaly detection, containing about 150,000 flow samples, divided into two categories: normal and abnormal. It mainly includes 41 statistical features such as error rate, duration, and service time.
[0056] Five commonly used neural network models were used as backbone models in the experiment. All models were implemented using the Tensorflow [1] framework and trained on a single NVIDIA GeForce RTX 3090 GPU. A brief introduction to the five models is as follows: VGG: A simple yet effective convolutional neural network consisting of a series of convolutional and max pooling layers followed by a fully connected layer at the output.
[0057] ResNet: A classic residual network that alleviates the degradation problem in deep neural networks by introducing skip (residual) connections.
[0058] ET-BERT (Encrypted Traffic Identification BERT): A Transformer-based approach for encrypted traffic classification that leverages a pre-trained Transformer model with a multi-layer attention mechanism.
[0059] NetMamba: An efficient network traffic classification model that uses a pre-trained unidirectional Mamba architecture and addresses challenges in model efficiency and traffic characterization.
[0060] AE (Autoencoder): A neural network designed for unsupervised learning that detects anomalies by reconstructing input data.
[0061] In specific implementation, Fig. 9 To illustrate the recognition accuracy of different numbers (K) of statistical features extracted from the UNSW-IoT dataset by the present invention under the NetMamba model, it can be seen that when 24 statistical features are selected, the model achieves the highest accuracy, which is 0.13% higher than the accuracy of the complete 72 statistical features; Fig.10The recognition accuracy of different numbers of statistical features of the UNSW-NB dataset extracted by the present invention under the NetMamba model shows that when 32 statistical features are selected, the model achieves the highest accuracy, which is 1.80% higher than the accuracy of the complete 42 statistical features; Fig.11 The recognition accuracy of different numbers of statistical features of the NSL-KDD data set extracted by the present invention under the NetMamba model shows that when four statistical features are selected, the model achieves the highest accuracy, which is 3.71% higher than the accuracy of the complete 41 statistical features.
[0062] When the computer program is executed, each functional system or module included in the above-mentioned network structure feature selection and optimization method is implemented.
[0063] Specifically, the experimental results are shown in Table 1. The effectiveness of the feature selection method on the accuracy of downstream tasks under different memory constraints is compared through experiments. First, we use the neural network model to evaluate all tasks with all features, set to "all". Secondly, the pre-selection module of the present invention is applied to select the most important K features, and the task is evaluated again, denoted as "K". Finally, we test the selected features under the SRAM memory constraints of 512bit and 256bit, respectively. For comparison, two classic machine learning methods for traffic identification tasks are also compared in the experiment: AppScanner and KNN. AppScanner is a traffic identification method that detects and identifies applications by analyzing network traffic features. KNN is a widely used classification method that calculates the distance between the sample to be classified and the training sample, selects the K samples with the closest distance, and then determines the category of the sample to be classified based on the category of these K samples. From the experimental results in Table 1, it can be found that: (1) The downstream task-sensitive feature pre-selection module can identify the key K features, so that the models trained on these features can achieve comparable or higher accuracy. As shown in the results of the UNSW-IoT dataset in Table 1, by selecting the important K features, most models achieve higher accuracy than the models using the full feature set, including ResNet, VGG, AE, and NetMamba models. At the same time, the ET-BERT model achieves comparable accuracy to the model with all features; (2) The feature pre-selection module recognizes the downstream task sensitivity and eliminates feature redundancy, and significantly reduces computational and memory costs. As shown in Table 1, in the traffic anomaly detection task on the NSL-KDD dataset, the selected feature subset accounts for 10%-78% of the total 41 features; (3) Even under the memory constraints of 256 bits and 512 bits, the feature selection optimization method of the present invention maintains a high accuracy, proving the effectiveness of the method in SRAM memory-constrained environments. For example, for the IoT device identification task of the UNSW-IoT dataset shown in Table 1, under the constraints of 512 bits and 256 bits, most models achieve comparable accuracy to the model using the full features, and the NetMamba model even outperforms the model using the full features; Table 1 Comparison of the accuracy of two indicators of different traffic identification models in three data sets
[0064] As shown in Table 2, the impact of early stopping mechanism on feature selection was evaluated through experiments. Specifically, the impact of early stopping on training time and model accuracy was analyzed through experiments. Five backbone neural networks were used in the UNSW-IoT dataset for experiments, where No means not using early stopping mechanism, Yes means using early stopping mechanism, and the bold index in the selection factor Top-K index set is the difference between the index set without early stopping mechanism and the index set with early stopping mechanism after using early stopping mechanism. The experiment is summarized as follows: (1) After adding the early stopping mechanism, the training time of the NetMamba model was reduced most significantly, by 83.33% (100 / 120 epochs) compared to the case without the early stopping mechanism. In contrast, ET-BERT had the smallest reduction, only 40% (4 / 10 epochs). Therefore, after adding the early stopping mechanism, the training time of the model was reduced by 40% to 83.33% compared to the case without the early stopping mechanism, proving the effectiveness of the mechanism in accelerating feature selection; (2) The accuracy of the model trained with early stopping is comparable to the model trained to convergence. This suggests that the features that have the greatest impact on the accuracy of the downstream task are identified early in the training process; (3) The selected feature set differs only slightly between full training and early stopping, which confirms that important candidate features remain stable during training; Table 2 The impact of early stopping mechanism in UNSW-IoT dataset:
[0065] Table 3 shows the impact of the class-balanced (CB) loss function on four downstream task models (ResNet, VGG, ET-BERT, and NetMamba) evaluated in the UNSW-IoT dataset. AE is excluded because the autoencoder-based neural network model does not support the CB loss function. The experiments show that the use of the class-balanced loss function improves the Micro-F1 and Macro-F1 scores in all models, confirming the effectiveness of the CB loss in solving class imbalance and improving model performance; Table 3 Comparison of accuracy with and without class balance (CB) loss function:
[0066] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) Stream statistical feature selection that is adaptive to the downstream task model. In order to achieve adaptive selection of statistical features of different task models, the present invention designs a selection factor layer located between the task model and the statistical features. The selection factor evolves continuously during the training process of the task model and gradually stabilizes. After the selection factor stabilizes, its size can be used as an indicator of feature importance. Based on this indicator, the most critical K candidate features of the task model can be identified and redundant features can be removed; (2) Feature selection under bit width constraints. In order to ensure the accuracy of downstream tasks and select the most critical features without exceeding the SRAM memory limit, the present invention converts the original feature selection problem into a feature knapsack problem according to the value of the selection factor and its corresponding bit width requirement. In addition, a dynamic programming feature selection algorithm is designed to solve this feature knapsack problem. (3) Fast feature selection algorithm. In order to reduce computational overhead and speed up the feature selection process, the present invention first conducted a large number of experiments and found that the largest K selection factors can be quickly determined before the model converges. On this basis, the present invention designs an early stopping indicator to monitor the convergence of the selection factors and proposes an early stopping mechanism. This mechanism can quickly complete feature selection after the selection factors converge, rather than waiting for the model training to be completed, thereby speeding up the feature selection process; (4) Data imbalance processing. Since the data plane measurement framework proposed in the present invention is task-driven, two downstream tasks are implemented on the control plane, including IoT device identification and anomaly detection. In these tasks, there is a common problem (i.e., data imbalance). However, in the feature selection process, the accuracy of the downstream task model is usually required as the measurement standard. Therefore, in solving the problem of network flow data imbalance, it is crucial to ensure the fairness of the flow statistical feature selection process. In order to solve the problem of traffic data imbalance, the present invention designs a class-balanced loss function. By assigning different weight losses to different network flow data, the downstream task model is effectively prevented from overfitting to the network flow data class with a larger sample size, thereby reducing the impact of data imbalance on model performance.
[0067] The above contents are merely examples and explanations of the structure of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the protection scope of the present invention.
Claims
1. A method for feature selection and optimization of a network structure, characterized in that: The following methods are included: Build the task model and add the selection factor layer to obtain the feature selection model; The acquired traffic data set is used as the input of the feature selection model, and K statistical features and their corresponding selection factor values are screened out according to the feature selection strategy to obtain the pre-screening features; Taking the memory bit width length as a feature, a dynamic programming algorithm is used to screen out features that meet the memory bit width restriction condition, and the features are combined with the pre-screened features to obtain traffic features; The accuracy is obtained for the spatial distance between the traffic feature and the set label feature; according to the accuracy, the value of the selection factor in the selection factor layer is adjusted, and an instruction is issued through the control plane to adjust the matching operation table of the data plane, so as to select a more suitable traffic feature; wherein the spatial distance is calculated by the total loss function constructed by introducing the class balance term and the focus term in the cross entropy loss function.
2. The feature selection and optimization method of the network structure according to claim 1, characterized in that: The selection factor layer consists of a set of learnable parameters parameterized, where each parameter Corresponding to an updateable feature .
3. The feature selection and optimization method of the network structure according to claim 1, characterized in that: The feature selection strategy is to introduce a heap-based algorithm to identify the optimal K selection factors. The specific algorithm is as follows: Use the first K selection factors to construct a small root heap, where the top of the heap represents the smallest value; Iterate the remaining (NK) selection factors; for each factor, if its value exceeds the value of the top element of the heap, replace the value of the top element and re-adjust the heap; After iterating over all factors, the elements in the heap are the Top-K selection factors.
4. The feature selection and optimization method of the network structure as claimed in claim 3, characterized in that: The index set composed of the Top-K selection factors The stability of is set as the Top-K stability metric TSM, which is set as follows: ;in, is the stability measure for the 𝑖th epoch, Q is a hyperparameter that specifies the number of consecutive epochs to consider, and It is the ) epoch’s Top-K feature index set; Represents the past Q epochs The intersection of When The value has remained unchanged for Q consecutive epochs, triggering the early stopping mechanism.
5. The method for feature selection and optimization of a network structure according to claim 1, characterized in that: The bit width restriction condition refers to the bit width restriction constraint of the SRAM, and the dynamic programming algorithm used is operated as follows: Problem definition: Indicates the total bit width limit of SRAM, represents the bit width of the candidate feature, represents the selection factor value of the candidate feature, Indicates whether the i-th feature is selected. The goal is to maximize the sum of the selection factor values while ensuring that the total bit width does not exceed , the problem is expressed as: ; State definition: set Indicates that in the past feature selection, the total bit width is limited to The sum of the maximum selectivity factor values obtained, set Storage and The corresponding index set of the selected features; Initialization: Set all Initialized to 0, indicating no value is assigned and the total value is zero; Initialized to empty, indicating that no index is recorded initially; State transition: If the i-th feature is not selected: , If we select the i-th feature (j≥ ): ; ; Optimization evaluation: The optimal value is , the index set corresponding to the selected option is .
6. The method for feature selection and optimization of a network structure according to claim 1, characterized in that: The total loss function is expressed as: ; in, It is a class balancing item that can be used to balance the sample size of each class. Adjust the weight of each class; It is a focus item that can pay more attention to the sample classes that are difficult to predict.
7. A computer program, characterized in that When the computer program is executed, each functional system or module included in the feature selection and optimization method of the network structure as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Traffic feature extraction method and system, storage medium and electronic equipment
CN114024758A
Neural architecture searching method for chip design
CN118821863A