Network traffic classification method and related devices based on bayesian rule distillation
Patent Information
- Application Number
- CN202611184357.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-06
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-08-06
AI Technical Summary
但相关技术中,浅层决策虽然易于部署,但其分类能力有限,而深层决策树或随机森林虽然能够取得更高的分类性能,但其会产生大量叶路径规则,容易消耗过多数据面资源
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a network traffic classification method, apparatus, electronic device, and storage medium based on Bayesian rule distillation. This scheme acquires network traffic training data to obtain corresponding network traffic features in advance, and then trains a preset tree model based on these features to construct a teacher model. Next, the embodiments of this invention extract preset rule information from the leaf paths of the teacher model, including candidate rules and leaf node category data. Then, the embodiments of this invention construct empirical Bayesian prior data based on the leaf node category data, and then combine this with training sample matching data under a preset matching strategy, i.e., preset sample matching data, to update the posterior category distribution of the candidate rules, determine preset posterior probability information, and select candidate rules using a greedy algorithm based on this preset posterior probability information to generate an expected rule table. Further, the embodiments of this invention generate switch data plane deployment rules based on the expected rule table using a preset bitmap and encoding mechanism, and deploy them to the target switch. Then, by inputting the traffic data to be classified into the target switch for analysis and processing, the traffic classification result is obtained, thus achieving network traffic classification. It is readily understood that the embodiments of the present invention effectively alleviate the problem of rule proliferation by constructing a teacher model based on a preset tree model and then distilling the teacher model into a desired rule table. This effectively reduces hardware resource consumption and rule size while maintaining network traffic classification accuracy. Furthermore, the embodiments of the present invention introduce the category distribution information learned from the leaf nodes of the teacher model into the rule modeling process through empirical Bayesian priors, thereby improving network traffic classification accuracy. Moreover, by using a preset bitmap and encoding mechanism to transform the desired rule table into deployment rules for the switch data plane, the embodiments of the present invention alleviate the problem of the data plane being unable to perform complex calculations, effectively improving the deployability of the programmable switch data plane and achieving a hardware-friendly form of network traffic classification.
Smart Images

Figure CN122678985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network technology, and in particular to a network traffic classification method and related equipment based on Bayesian rule distillation. Background Technology
[0002] Network traffic classification is a fundamental capability in modern network management and security, applicable to scenarios such as application identification, intrusion detection, and malicious traffic identification. Traditional network traffic classification typically relies on switches or routers to send sampled data, aggregated statistics, or mirrored traffic to a centralized server for inference. This approach incurs additional control plane communication overhead, higher response latency, and privacy risks due to the exposure of raw traffic. With the development of programmable switches and the P4 language, traffic classification logic can be pushed down to the switch data plane, allowing data packets to complete inference directly in the forwarding path, thereby reducing data transmission between the control plane and data plane and improving response speed. Since the data plane of a programmable switch is not a general-purpose computing platform, its pipeline typically only supports simple integer operations, comparisons, bitwise operations, and a limited matching action table, while floating-point operations, matrix multiplication, loops, recursion, and complex nonlinear functions are difficult to implement directly. Decision trees, with their core focus on conditional judgment and path matching, have a natural similarity to the matching action pipeline of a switch and are therefore often used for network classification model deployment. However, while shallow decision trees are easy to deploy, their classification capabilities are limited. Deep decision trees or random forests, on the other hand, can achieve higher classification performance, but they generate a large number of leaf path rules, which can easily consume too much data surface resources.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this application is to propose a network traffic classification method and related equipment based on Bayesian rule distillation, which can effectively improve the accuracy of network traffic classification, while reducing hardware resource consumption and rule scale, and achieving network traffic classification in a hardware-friendly manner.
[0005] To achieve the above objectives, one aspect of this application proposes a network traffic classification method based on Bayesian rule distillation, the method comprising: Obtain network traffic training data, and then extract the corresponding network traffic features; The teacher model is constructed by training the preset tree model based on the network traffic characteristics. Preset rule information is extracted from the leaf paths of the teacher model; wherein, the preset rule information includes candidate rules and leaf node category data; Empirical Bayesian prior data is constructed based on the leaf node category data, and then the posterior category distribution of the candidate rules is updated by combining it with preset sample matching data to determine preset posterior probability information; wherein, the preset sample matching data includes training sample matching data under the preset matching strategy. Based on the preset posterior probability information, a greedy algorithm is used to select the candidate rules and generate an expected rule table. Based on the desired rule table, switch data plane deployment rules are generated using a preset bitmap and encoding mechanism, and then deployed to the target switch; The traffic data to be classified is input into the target switch for analysis and processing to obtain the traffic classification result.
[0006] In some embodiments, extracting preset rule information from the leaf paths of the teacher model includes: The node judgment information on each rule path of the teacher model is converted into atomic predicates; wherein, the atomic predicates include feature indexes, comparison operators, and thresholds; Traverse the teacher model, convert each rule path from the root node to the leaf node of the teacher model into the corresponding candidate rule, and combine the atomic predicates corresponding to the rule path according to the path order to construct the preset rule antecedent; The category count vectors of the model leaf nodes corresponding to the candidate rules are counted to obtain the leaf node category data.
[0007] In some embodiments, the step of constructing empirical Bayesian prior data based on the leaf node category data, and then updating the posterior category distribution of the candidate rules in conjunction with preset sample matching data to determine preset posterior probability information includes: The category count vectors of the model leaf nodes corresponding to the candidate rules are normalized to obtain the leaf node category distribution; Dirichlet prior parameters are constructed based on the smoothed baseline term and the leaf node category distribution; Based on the order of the candidate rules, each network traffic training sample is allocated according to the preset matching strategy to obtain the sample allocation result; Based on the sample allocation results, the training sample categories allocated to each candidate rule are statistically analyzed to obtain a matching count vector; Posterior parameters are constructed based on the matching count vector and the Dirichlet prior parameters, and the corresponding posterior mean is calculated using the posterior parameters; wherein the posterior mean serves as the class probability vector of the candidate rule.
[0008] In some embodiments, the step of allocating each network traffic training sample according to the order of the candidate rules and the preset matching strategy to obtain the sample allocation result includes: The coverage of each network traffic training sample is analyzed sequentially according to the order of the candidate rules to obtain sample coverage information; When it is determined from the sample coverage information that the network traffic training sample satisfies several candidate rules, the network traffic training sample is matched with the first coverage rule; or, when it is determined from the sample coverage information that the network traffic training sample does not satisfy any of the candidate rules, the network traffic training sample is matched with the default category branch; wherein, the first coverage rule includes the candidate rule that covers the network traffic training sample first in the order of arrangement.
[0009] In some embodiments, the step of selecting candidate rules based on the preset posterior probability information using a greedy algorithm to generate an expected rule table includes: Construct an ordered rule table; The preset F1 increment is evaluated according to the candidate rules to obtain preset increment data; wherein, the preset F1 increment includes the macro-average F1 increment generated on the validation set after adding the candidate rules that were not selected to the ordered rule table; Based on the preset incremental data and the preset posterior probability information, the desired selection rule is determined from each of the candidate rules, and then the desired selection rule is added to the ordered rule table to construct the desired rule table.
[0010] In some embodiments, generating switch data plane deployment rules based on the desired rule table using a preset bitmap and encoding mechanism, and deploying them to the target switch, includes: Assign a corresponding bit position to each candidate rule in the expected rule table, divide each network traffic feature in the expected rule table into regions, and then construct a feature bitmap lookup table based on the correspondence between the constructed feature regions and the bit positions. A priority parsing table is constructed according to a preset priority rule; wherein, the preset priority rule includes, when there is at least one set bit in the target fusion bitmap, taking the candidate rule corresponding to the set bit with the smallest rule number as the target matching rule; Constructing a bitmap and fusion logic; wherein, the bitmap and fusion logic includes performing a bitwise AND operation on several feature bitmaps determined by the input data packet, and then determining the target matching rule through the priority parsing table based on the bit setting information of the obtained first fused bitmap; The feature bitmap lookup table, the priority parsing table, and the bitmap and fusion logic are deployed to the target switch.
[0011] In some embodiments, the bitmap and fusion logic further includes: The corresponding block bitmap is obtained by querying the combination of feature intervals within the block based on the input data packet; wherein, the block bitmap is constructed by dividing each network traffic feature in the expected rule table into several feature blocks, and performing a bitwise AND operation on the feature bitmap of each feature interval combination within each feature block; Perform a bitwise AND operation on the block bitmap to obtain a second fused bitmap; The target matching rule is determined by the priority parsing table based on the bit setting information of the second fused bitmap.
[0012] To achieve the above objectives, another aspect of this application proposes a network traffic classification device based on Bayesian rule distillation, the device comprising: The first module is used to acquire network traffic training data and then extract the corresponding network traffic features. The second module is used to train the preset tree model based on the network traffic characteristics to construct the teacher model; The third module is used to extract preset rule information from the leaf paths of the teacher model; wherein, the preset rule information includes candidate rules and leaf node category data; The fourth module is used to construct empirical Bayesian prior data based on the leaf node category data, and then update the posterior category distribution of the candidate rules by combining it with preset sample matching data to determine preset posterior probability information; wherein, the preset sample matching data includes training sample matching data under the preset matching strategy. The fifth module is used to select candidate rules based on the preset posterior probability information using a greedy algorithm, and generate an expected rule table; The sixth module is used to generate switch data plane deployment rules based on the expected rule table through a preset bitmap and encoding mechanism, and deploy them to the target switch; The seventh module is used to input the traffic data to be classified into the target switch for analysis and processing to obtain the traffic classification result.
[0013] To achieve the above objectives, another aspect of this application provides an electronic device, the electronic device comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a network traffic classification method, apparatus, electronic device, and storage medium based on Bayesian rule distillation. This scheme acquires network traffic training data to obtain corresponding network traffic features in advance, and then trains a preset tree model based on these features to construct a teacher model. Next, the embodiments of this invention extract preset rule information from the leaf paths of the teacher model, including candidate rules and leaf node category data. Then, the embodiments of this invention construct empirical Bayesian prior data based on the leaf node category data, and then combine this with training sample matching data under a preset matching strategy, i.e., preset sample matching data, to update the posterior category distribution of the candidate rules, determine preset posterior probability information, and select candidate rules using a greedy algorithm based on this preset posterior probability information to generate an expected rule table. Further, the embodiments of this invention generate switch data plane deployment rules based on the expected rule table using a preset bitmap and encoding mechanism, and deploy them to the target switch. Then, by inputting the traffic data to be classified into the target switch for analysis and processing, the traffic classification result is obtained, thus achieving network traffic classification. It is readily understood that the embodiments of the present invention effectively alleviate the problem of rule proliferation by constructing a teacher model based on a preset tree model and then distilling the teacher model into a desired rule table. This effectively reduces hardware resource consumption and rule size while maintaining network traffic classification accuracy. Furthermore, the embodiments of the present invention introduce the category distribution information learned from the leaf nodes of the teacher model into the rule modeling process through empirical Bayesian priors, thereby improving network traffic classification accuracy. Moreover, by using a preset bitmap and encoding mechanism to transform the desired rule table into deployment rules for the switch data plane, the embodiments of the present invention alleviate the problem of the data plane being unable to perform complex calculations, effectively improving the deployability of the programmable switch data plane and achieving a hardware-friendly form of network traffic classification. Attached Figure Description
[0016] Figure 1 This is a flowchart of the network traffic classification method based on Bayesian rule distillation provided in this embodiment of the invention; Figure 2 This is a schematic diagram of a system architecture for control plane and data plane collaboration provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the Bayesian distillation process provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of bitmaps and encoding provided in an embodiment of the present invention; Figure 5 This is a block bitmap and encoding schematic diagram provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the network traffic classification device based on Bayesian rule distillation provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” or “when” as used herein may be interpreted as “when…” or “in response to determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0022] Greedy algorithm: refers to an algorithm that makes the best or optimal choice in each step of the current state, in order to achieve the best or optimal result in the whole world.
[0023] Tree models are models applied in the fields of machine learning and data structures. Their structure resembles an inverted tree, such as decision tree models and forest models.
[0024] Network traffic classification is a fundamental capability in modern network management and security, applicable to scenarios such as application identification, intrusion detection, and malicious traffic identification. Traditional network traffic classification typically relies on switches or routers to send sampled data, aggregated statistics, or mirrored traffic to a centralized server for inference. This approach incurs additional control plane communication overhead, higher response latency, and privacy risks due to the exposure of raw traffic. With the development of programmable switches and the P4 language, traffic classification logic can be pushed down to the switch data plane, allowing data packets to complete inference directly in the forwarding path, thereby reducing data transmission between the control plane and data plane and improving response speed. Since the data plane of a programmable switch is not a general-purpose computing platform, its pipeline typically only supports simple integer operations, comparisons, bitwise operations, and a limited matching action table, while floating-point operations, matrix multiplication, loops, recursion, and complex nonlinear functions are difficult to implement directly. Decision trees, with their core focus on conditional judgment and path matching, have a natural similarity to the matching action pipeline of a switch and are therefore often used for network classification model deployment. However, among related technologies, shallow decision trees are easy to deploy, but their classification ability is limited, while deep decision trees or random forests can achieve higher classification performance, but they generate a large number of leaf path rules, which can easily consume too much data surface resources.
[0025] In view of this, this application provides a network traffic classification method, apparatus, electronic device, and storage medium based on Bayesian rule distillation. This scheme acquires network traffic training data to obtain corresponding network traffic features in advance, and then trains a preset tree model based on these features to construct a teacher model. Next, this embodiment extracts preset rule information from the leaf paths of the teacher model, including candidate rules and leaf node category data. Then, this embodiment constructs empirical Bayesian prior data based on the leaf node category data, and then combines this with training sample matching data under a preset matching strategy, i.e., preset sample matching data, to update the posterior category distribution of the candidate rules, determine preset posterior probability information, and selects candidate rules using a greedy algorithm based on this preset posterior probability information to generate an expected rule table. Furthermore, in this embodiment of the invention, based on the desired rule table, a switch data plane deployment rule is generated through a preset bitmap and encoding mechanism and deployed to the target switch. Then, by inputting the traffic data to be classified into the target switch for analysis and processing, the traffic classification result is obtained, thereby realizing network traffic classification. This can effectively improve the accuracy of network traffic classification, while reducing hardware resource consumption and rule size, and achieving a hardware-friendly form of network traffic classification.
[0026] The network traffic classification method based on Bayesian rule distillation provided in this application relates to the field of computer network technology. This method can be applied to terminals, servers, or software running on either a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the network traffic classification method based on Bayesian rule distillation, but is not limited to the above forms.
[0027] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0028] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0029] Figure 1 This is an optional flowchart of a network traffic classification method based on Bayesian rule distillation provided in this application embodiment. Figure 1 The method may include, but is not limited to, steps S110 to S170: Step S110: Obtain network traffic training data, and then extract the corresponding network traffic features; Step S120: Train the preset tree model based on network traffic characteristics to construct the teacher model; Step S130: Extract preset rule information from the leaf paths of the teacher model. The preset rule information includes candidate rules and leaf node category data. Step S140: Construct empirical Bayesian prior data based on leaf node category data, and then update the posterior category distribution of candidate rules by combining it with preset sample matching data to determine preset posterior probability information. The preset sample matching data includes training sample matching data under the preset matching strategy. Step S150: Based on the preset posterior probability information, a greedy algorithm is used to select candidate rules and generate an expected rule table; Step S160: Generate switch data plane deployment rules based on the desired rule table using a preset bitmap and encoding mechanism, and deploy them to the target switch; Step S170: Input the traffic data to be classified into the target switch for analysis and processing to obtain the traffic classification result.
[0030] In the operation of this specific embodiment, the present invention first acquires network traffic training data to extract corresponding network traffic features. Specifically, the network traffic training data in this embodiment refers to data used for network traffic classification training, including multiple network flow samples, each of which includes a traffic feature vector and a category label. The network traffic features in this embodiment refer to traffic features that can be deployed on the programmable switch data plane, such as packet-level features that can be directly parsed from packet header fields, and / or flow-level features maintained through finite state registers, counters, or time windows. The network traffic features in this embodiment include one or more of the following: packet length statistics, flow duration, number of forward or backward packets, number of bytes, flag statistics, port information, and time-related statistics. Accordingly, the network traffic features in this embodiment can be used for network attack detection, malicious traffic identification, or application traffic classification. For example, in this embodiment, the programmable switch control plane first acquires labeled network traffic training data. The network traffic training data is shown in the following formula: ; Where, in the formula Indicates the first The feature vector of a network traffic sample Indicates the first Category labels for each network traffic sample This indicates the number of samples. For the C classification task, belong .
[0031] Next, after acquiring the training data, the control plane performs feature filtering and feature transformation operations. Since the programmable switch data plane is not suitable for performing floating-point operations, complex functions, cyclic statistics, or high-dimensional matrix operations, this embodiment of the invention prioritizes lightweight features that can be parsed or maintained by the switch. For example, packet header fields can be directly extracted by the parser; packet length, port number, protocol number, and TCP flags can be directly used as packet-level features; flow duration, number of flow packets, number of flow bytes, maximum packet length, and minimum packet length can be maintained through registers, counters, and time windows. Furthermore, this embodiment of the invention divides the original network traffic training data into a training set, a validation set, and a test set through the control plane. The training set is used to train the teacher model and count the matching categories of candidate rules; the validation set is used to calculate the macro-average F1 (Macro-F1) increment after adding candidate rules to the ordered rule table under rule budget constraints; and the test set is used to evaluate the classification performance of the finally generated compact ordered rule table. It should be noted that, in this embodiment of the invention, the control plane can also preprocess traffic features, including missing value imputation, outlier handling, integerization, discretization, feature range pruning, category label encoding, and feature name normalization. Simultaneously, for features that need to be deployed to the data plane, the control plane can also record their data plane acquisition method, such as direct parsing from packet header fields, maintenance by registers, maintenance by counters, or pre-configuration by the control plane. Through the above steps, this embodiment of the invention ensures that the teacher model and candidate rules obtained through subsequent training are based on features that can be obtained or maintained by the programmable switch data plane, thereby avoiding the generation of data plane rules that cannot be deployed.
[0032] Next, in this embodiment of the invention, a pre-defined tree model is trained based on network traffic characteristics to construct a teacher model. Specifically, the pre-defined tree model in this embodiment includes a decision tree model or a forest tree model. Correspondingly, the teacher model trained in this embodiment can be a C4.5 decision tree, a random forest, or a tree model composed of multiple decision trees. After training the pre-defined tree model on the programmable switch control plane based on the acquired network traffic characteristics to obtain the teacher model, this embodiment does not directly deploy it to the programmable switch data plane, but uses it as the source model for candidate rule extraction and rule distillation. It is easy to understand that although deep decision tree or forest models can express complex nonlinear classification boundaries and usually have high classification performance in network attack detection, malicious traffic identification, and application traffic classification tasks, deep decision tree or forest models contain a large number of leaf nodes, each corresponding to a path rule from the root node to the leaf node. If directly deployed to the programmable switch data plane, it will lead to an excessive number of rules, an excessively large table size, excessively deep pipeline dependencies, or excessive metadata width. Therefore, in this embodiment of the invention, a teacher model is used to learn complex classification knowledge, and then a deployable compact and ordered rule table, namely the expected rule table, is obtained through subsequent Bayesian rule distillation and budget-aware rule selection.
[0033] Furthermore, this embodiment of the invention extracts preset rule information from the leaf paths of the teacher model. Specifically, the preset rule information in this embodiment includes candidate rules and leaf node category data. When the teacher model is a single decision tree, the control surface traverses all leaf nodes of the decision tree to obtain candidate rules. When the teacher model is a random forest, the control surface traverses all leaf nodes of each tree in the forest, forming a candidate rule pool from all leaf path rules. For identical or equivalent leaf path rules in the forest model, the control surface can choose to retain multiple sources, merge statistical information, or remove duplicate rules. Simultaneously, this embodiment of the invention statistically analyzes the distribution of the number of training samples at each category after they reach the corresponding leaf nodes, obtaining leaf node category data. Then, this embodiment of the invention constructs empirical Bayesian prior data based on the leaf node category data, and then updates the posterior category distribution of the candidate rules by combining it with preset sample matching data, determining the preset posterior probability information. Specifically, in this embodiment of the invention, the control surface constructs an empirical Bayesian prior based on the category count vector of the leaf node corresponding to each candidate rule, i.e., the leaf node category data. By using the category distribution of the leaf nodes of the teacher model as the empirical prior, the candidate rules inherit the local category structure learned by the teacher model at the corresponding leaf nodes. Correspondingly, the preset sample matching data in this embodiment includes training sample matching data under a preset matching strategy. The preset matching strategy refers to the pre-set matching semantics between training samples and candidate rules. In this embodiment, the posterior category distribution of each candidate rule is updated based on the empirical Bayesian prior data combined with the training sample matching data under the preset matching strategy, thereby obtaining the corresponding preset posterior probability information.
[0034] Further, in this embodiment of the invention, candidate rules are selected using a greedy algorithm based on preset posterior probability information to generate an expected rule table. Then, based on the expected rule table, a switch data plane deployment rule is generated using a preset bitmap and encoding mechanism, and deployed to the target switch. Specifically, under preset rule budget constraints, this embodiment of the invention flexibly selects each candidate rule in the candidate rule pool based on the macro-average F1 gain of the validation set, thereby obtaining a compact and ordered rule table that satisfies hardware resource constraints, i.e., the expected rule table. Then, in this embodiment of the invention, the control plane generates deployment rules for the programmable switch data plane based on this compact and ordered rule table. That is, through a corresponding bitmap and encoding mechanism, the compact and ordered rule table is transformed into a feature interval bitmap and corresponding decision entries to improve the deployability of the programmable switch data plane. For example, this embodiment of the invention first divides the value space of each feature into multiple feature intervals based on the feature thresholds involved in the compact and ordered rule table, and then generates a rule bitmap for each feature interval. Each bit in the rule bitmap corresponds to a rule, used to indicate whether the rule satisfies the current feature interval. Then, decision entries are generated based on the rule number, prediction category, and priority. Accordingly, in the deployment process of this embodiment of the invention, a P4 program and runtime entries are first generated based on the corresponding switch data plane deployment rules. The P4 program is then compiled and deployed to a programmable switch based on the PISA architecture. The runtime entries are then distributed to the programmable switch via the runtime interface, enabling the classification logic to execute at line speed in the packet forwarding path. Finally, this embodiment of the invention inputs the traffic data to be classified into the target switch for analysis and processing to obtain the traffic classification results. Specifically, the traffic data to be classified in this embodiment refers to traffic data that requires network traffic classification analysis, such as network traffic input data to the programmable switch data plane. Accordingly, after the traffic data to be classified enters the programmable switch data plane, i.e., the target switch data plane, the data plane parser first extracts packet header fields, such as source IP address, destination IP address, source port, destination port, protocol type, packet length, and TCP flags. Simultaneously, for flow-level features that need to be maintained, the data plane can look up or update registers, counters, and time window states based on the flow identifier to obtain the statistical features corresponding to the current flow. Then, the target switch performs table lookups and bitwise AND operations according to the deployment rules based on the parsed feature information to obtain the corresponding traffic classification results. This invention addresses the problem that deep decision tree models or forest models are difficult to deploy directly on resource-constrained programmable switches by distilling the classification knowledge in the high-performance tree model of the control plane into a compact and ordered rule table that can be executed in the data plane, and achieving efficient rule matching through bitmap AND encoding.
[0035] In some embodiments of the present invention, preset rule information is extracted from the leaf paths of the teacher model, including but not limited to the following steps: The node judgment information on each rule path of the teacher model is converted into atomic predicates. These atomic predicates include feature indices, comparison operators, and thresholds. Traverse the teacher model, convert each rule path from the root node to the leaf node into a corresponding candidate rule, and combine the atomic predicates corresponding to the rule path according to the path order to construct the preset rule antecedent. The leaf node category data is obtained by counting the category count vectors of the model leaf nodes corresponding to the candidate rules.
[0036] In this specific embodiment, the present invention converts the node judgment information on each rule path of the teacher model into atomic predicates, and traverses the teacher model to convert each rule path from the root node to the leaf node into a corresponding candidate rule. Simultaneously, the atomic predicates corresponding to the rule paths are combined according to the path order to construct the corresponding preset rule antecedent. Specifically, in this embodiment, the programmable switch control plane traverses the leaf nodes in the teacher model, converting each path from the root node to the leaf node into a candidate rule. At the same time, for each internal node judgment on the path, i.e., the node judgment information, the control plane converts it into an atomic predicate. Accordingly, the atomic predicate includes a feature index, a comparison operator, and a threshold. The comparison operator includes one of less than, less than or equal to, greater than, and greater than or equal to. For example, if a path from the root node to the leaf node in the teacher model sequentially passes through three judgment conditions: feature... Less than or equal to the threshold ;feature Greater than the threshold ;feature Less than or equal to the threshold The path can then be converted into a candidate rule, whose presupposition is the ordered conjunction of the three atomic predicates mentioned above. Furthermore, a candidate rule is considered to cover an input sample only if the input sample simultaneously satisfies all atomic predicates. Simultaneously, this embodiment of the invention statistically analyzes the class count vectors of the model leaf nodes corresponding to the candidate rules to obtain leaf node class data. Specifically, in this embodiment, each candidate rule is associated with the class count vector of its source leaf node, and the class count vector represents the distribution of the number of training samples arriving at that leaf node across different classes. In this embodiment, the class count vectors of the training samples in the leaf nodes are used as statistical information for the candidate rules. Correspondingly, each candidate rule in this embodiment... This can be represented as a pair, as shown in the following equation: ; Where, in the formula Represents an ordered conjunctive set consisting of multiple atomic predicates. Candidate rules Source leaf node The category count vector. This is used to represent the distribution of the number of training samples in each category after reaching this leaf node. For example, for Classification tasks, It can be expressed as the following formula: ; Where, in the formula Indicates reaching the leaf node And the number of training samples of category c.
[0037] It should be noted that candidate rules are also constructed in some embodiments of the present invention. For input samples The covering function. When the input sample satisfy When all atomic predicates in the formula are used, it indicates a candidate rule. Covering input samples Otherwise, it indicates a candidate rule. Do not cover input samples The covering function is shown in the following equation: ; The coverage function represents the candidate rule. Covering input samples Conversely, if it is not covered, it is represented as: .
[0038] It is readily understood that, through the aforementioned candidate rule extraction process, this embodiment of the invention converts the tree path in the teacher model into corresponding conditional rules. Each rule consists of several feature conditions and carries leaf node category count information. This category count information reflects the category structure learned by the teacher model in local regions and can serve as an important basis for subsequent empirical Bayesian modeling.
[0039] In some embodiments of the present invention, empirical Bayesian prior data is constructed based on leaf node category data, and then the posterior category distribution of candidate rules is updated by combining it with preset sample matching data to determine preset posterior probability information, including but not limited to the following steps: The category count vectors of the model leaf nodes corresponding to the candidate rules are normalized to obtain the leaf node category distribution.
[0040] Dirichlet prior parameters are constructed based on the smoothed baseline term and the leaf node category distribution.
[0041] Based on the order of the candidate rules, the training samples of each network traffic are allocated according to the preset matching strategy to obtain the sample allocation results.
[0042] Based on the sample allocation results, the training sample categories assigned to each candidate rule are statistically analyzed to obtain the matching count vector.
[0043] The posterior parameters are constructed based on the matching count vector and the Dirichlet prior parameters, and the corresponding posterior mean is calculated using these parameters. The posterior mean serves as the class probability vector for the candidate rule.
[0044] In this specific embodiment, the present invention first normalizes the class count vectors of the model leaf nodes corresponding to the candidate rules, and then constructs Dirichlet prior parameters by combining the obtained leaf node class distribution and smoothing baseline term. Specifically, in this embodiment, the control surface constructs an empirical Bayesian prior based on the class count vectors of the leaf nodes from which each candidate rule originates. Unlike a completely uniform manual prior, this embodiment uses the class distribution of the teacher model leaf nodes as an empirical prior, enabling the candidate rules to inherit the local class structure learned by the teacher model at the corresponding leaf nodes. Accordingly, for candidate rules... The control surface first counts the class vectors of its source leaf nodes. After normalization, the leaf node category distribution is obtained, as shown in the following formula: ; Where, in the formula Candidate rules Normalized class distribution of source leaf nodes.
[0045] Next, in this embodiment of the invention, the control surface adds the smoothed baseline term to the weighted term of the leaf node category distribution to obtain the candidate rule. The Dirichlet prior parameters are shown in the following equation: ; Where, in the formula Candidate rules Prior parameters, Represents the smoothing coefficient. express A dimensional vector of all 1s This represents the prior weights of the teacher model. Candidate rules The normalized category distribution of the source leaf nodes, i.e., the leaf node category distribution. It should be noted that, in this embodiment of the invention, As a smoothing baseline term, it prevents the probability of a class not appearing in a leaf node from degenerating to zero. For example, if a minority class sample is absent from a leaf node, without a smoothing term, the prior probability of that class might be zero, leading to overly aggressive subsequent posterior estimates. Simultaneously, it... As a prior term of the teacher model, it is used to control the strength of candidate rules inheriting the class structure of the teacher model. Wherein, when When the size is large, the candidate rules depend more on the class distribution of the leaf nodes of the teacher model; conversely, when... When the information level is low, the candidate rules are closer to the uninformed prior. Accordingly, this embodiment of the invention constructs Dirichlet prior parameters based on the leaf node category distribution and combined with the smoothed baseline term, which can transform the leaf node statistics in the teacher model into the probabilistic prior of the candidate rules, giving the rules a quantifiable initial category tendency.
[0046] Furthermore, in this embodiment of the invention, the network traffic training samples are allocated according to the order of the candidate rules and a preset matching strategy to obtain the sample allocation result. The categories of training samples allocated to each candidate rule are then statistically analyzed based on the sample allocation result to obtain a matching count vector. Specifically, in this embodiment of the invention, the control plane sequentially determines whether each training sample is covered by a rule according to the current order of the candidate rules, and then determines the matching status of the training samples, i.e., the sample allocation result, according to the preset matching semantics. Then, the control plane counts the categories of training samples allocated to each candidate rule to obtain a matching count vector, as shown in the following formula: ; Where, in the formula This indicates that candidate rules are assigned according to a preset matching strategy. And the category is The number of training samples.
[0047] Furthermore, in this embodiment of the invention, posterior parameters are constructed based on the matching count vector and Dirichlet prior parameters, and the corresponding posterior mean is calculated using the posterior parameters. Specifically, in this embodiment of the invention, the class probability vector of the candidate rule is modeled as a Dirichlet-Categorical conjugate model. Since the Dirichlet distribution is a commonly used prior for multi-class probability vectors and has a conjugate relationship with the class distribution or multinomial distribution, posterior updates can be completed by adding the prior parameters to the observed class counts, without requiring complex iterative optimization, sampling, or backpropagation training. Accordingly, in this embodiment of the invention, the candidate rule... The posterior parameters are shown in the following formula: ; Where, in the formula Candidate rules The posterior parameters, Represents the weights of the matching counts for training samples. This indicates that candidate rules are assigned according to a preset matching strategy. The training sample class count vector.
[0048] Accordingly, based on posterior parameters The control surface can compute candidate rules. The posterior mean is used as the class probability vector of the candidate rule, as shown in the following formula: ; Where, in the formula Candidate rules Predicted as category The posterior probability; Meanwhile, candidate rules in the embodiments of the present invention The predicted category can be determined by the category with the highest posterior probability, as shown in the following formula: ; In addition, candidate rules in the embodiments of the present invention The confidence level can be defined as the maximum posterior probability, as shown in the following formula: ; It is readily understood that, through the above steps, the embodiments of the present invention can provide each candidate rule with a posterior class probability, predicted class, and confidence level, enabling the rule selection process to not only rely on rule coverage but also consider the statistical reliability of rule prediction. Furthermore, in some embodiments of the present invention, the control surface can also calculate the rule confidence level based on the concentration of the posterior distribution, the class probability interval, the number of matched samples, or the posterior entropy. For example, when a rule has a high maximum posterior probability and a large number of matched samples, the rule can be considered to have high reliability; when the posterior probabilities of different classes are similar, the rule can be considered to have high classification uncertainty.
[0049] In some embodiments of the present invention, each network traffic training sample is allocated according to the order of candidate rules and a preset matching strategy to obtain the sample allocation result, including but not limited to the following steps: The coverage of each network traffic training sample is analyzed sequentially according to the order of the candidate rules to obtain sample coverage information.
[0050] When the network traffic training sample is determined to satisfy several candidate rules based on the sample coverage information, it is matched with the first coverage rule. Alternatively, when the network traffic training sample is determined not to satisfy any candidate rule based on the sample coverage information, it is matched with the default category branch. The first coverage rule includes the candidate rules that cover the network traffic training sample first, in the order listed.
[0051] In this specific embodiment, the present invention first determines whether each network traffic training sample is covered by a rule according to the order of the candidate rules, that is, analyzes the coverage of each network traffic training sample, and then determines the candidate rules to match the network traffic training sample based on the coverage. Specifically, when it is determined through sample coverage information that a network traffic training sample satisfies several candidate rules, the present invention matches the network traffic training sample with the first coverage rule. Accordingly, the first coverage rule in the present invention includes the candidate rule that covers the network traffic training sample first in the order of arrangement. In addition, when it is determined that a network traffic training sample does not satisfy any candidate rule, the present invention matches the network traffic training sample with the default category branch. Exemplarily, in the present invention, the control plane determines whether each training sample is covered by a rule according to the current order of the candidate rules. Wherein, for any training sample If a sample satisfies multiple candidate rules, the control surface assigns it only to the first candidate rule that covers the training sample; otherwise, if it does not satisfy any candidate rule, it is assigned to the default category branch. Accordingly, this process can be expressed by the assignment function as follows: ; At the same time, when there are no covered training samples When choosing candidate rules, It can be set as the default branch identifier.
[0052] In some embodiments of the present invention, candidate rules are selected using a greedy algorithm based on preset posterior probability information to generate an expected rule table, including but not limited to the following steps: Construct an ordered rule table.
[0053] The preset F1 increment is evaluated based on the candidate rules to obtain the preset increment data. The preset F1 increment includes the macro-average F1 increment generated on the validation set after adding the candidate rules that were not selected to the ordered rule table.
[0054] Based on preset incremental data and preset posterior probability information, the desired selection rule is determined from each candidate rule, and then the desired selection rule is added to the ordered rule table to construct the desired rule table.
[0055] In this specific embodiment, the present invention first constructs an ordered rule table and evaluates the preset F1 increment based on candidate rules to obtain preset increment data. Specifically, within a given preset rule budget, the control plane of the present invention selects a portion of rules from the candidate rule pool to form a compact ordered rule table, i.e., the desired rule table. The preset rule budget in this embodiment can be set according to the TCAM, SRAM, number of pipeline stages, metadata width, bitmap width, and runtime table entry capacity of the programmable switch data plane. Furthermore, the present invention does not directly deploy all candidate rules, but instead performs a greedy selection on the validation set with incremental Macro-F1 as the target. The Macro-F1 is the unweighted average of the F1 values for each category, which reduces the impact of category imbalance on rule selection, ensuring that minority attack traffic, abnormal traffic, or low-frequency application categories receive the same importance as majority traffic during rule selection. Accordingly, the present invention first initializes the ordered rule table to be empty and sets a default category. The default category can be the category with the most samples in the training set, the most common category among unmatched samples, or a category specified by the user based on the application scenario. For example, in an intrusion detection scenario, the default category can be set to normal traffic or to an unknown category that needs to be further submitted to the control plane. In this embodiment of the invention, the macro-average F1 increment generated on the validation set after adding unselected candidate rules to the ordered rule table is used as the preset F1 increment to determine the corresponding preset increment data. For example, in each round of selection, the control plane sequentially evaluates the candidate rules that have not yet been selected in the candidate rule pool. Then, for each candidate rule... The control plane temporarily appends it to the current ordered rule table M to obtain a temporary rule table M'. Further, the control plane performs classification on the validation set according to a preset matching strategy and calculates the Macro-F1 increment of the temporary rule table M' relative to the current rule table M, i.e., the preset F1 increment, to obtain the preset increment data. Accordingly, in this embodiment of the invention, the macro-average F1 (Macro-F1) can be expressed as follows: ; Where, in the formula Indicates category The corresponding F1 value.
[0056] Further, in this embodiment of the invention, the desired selection rule is determined from each candidate rule based on preset incremental data and preset posterior probability information, and then the corresponding desired selection rule is added to the ordered rule table to construct the desired rule table. Specifically, this embodiment of the invention selects the desired selection rule from each candidate rule based on the determined preset incremental data and adds it to the ordered rule table to construct the desired rule table. For example, the control surface selects the candidate rule that generates the largest positive increment as the desired selection rule based on the preset incremental data and appends it to the current ordered rule table M. Correspondingly, if a given preset rule budget is reached, or if there is no candidate rule that can generate a positive increment, the selection stops and the current ordered rule table is output as a compact ordered rule table. In addition, this embodiment of the invention can also combine preset posterior probability information to determine the desired selection rule. For example, when multiple candidate rules generate the same or approximately the same Macro-F1 increment, the control surface performs a parallel rule processing step. In this invention, the control surface can prioritize candidate rules with higher posterior confidence. If the confidence levels are still the same or close, it prioritizes candidate rules covering more validation samples. If the number of covered samples is still the same or close, it prioritizes candidate rules with fewer rule predicates. If the number of rule predicates is still the same or close, it prioritizes candidate rules with lower estimated data surface entry costs. Furthermore, in this embodiment, the control surface can dynamically update the default category during the greedy selection process. For example, when some samples are not covered by the current ordered rule table, the control surface can set the default category based on the majority category of these uncovered samples, thereby improving overall classification performance. It is easy to understand that by combining preset posterior probability information and macro-average F1 increments for greedy selection to construct the expected rule table, this embodiment can filter out a small number of rules that significantly contribute to the Macro-F1 validation set from a large pool of candidate rules, significantly reducing the number of rules and data surface deployment costs while maintaining classification performance.
[0057] In some embodiments of the present invention, switch data plane deployment rules are generated based on a desired rule table using a preset bitmap and encoding mechanism, and then deployed to the target switch, including but not limited to the following steps: Each candidate rule in the expected rule table is assigned a corresponding bit position, and each network traffic feature in the expected rule table is divided into regions. Then, a feature bitmap lookup table is constructed based on the correspondence between the constructed feature regions and bit positions.
[0058] A priority parsing table is constructed based on preset priority rules. These preset priority rules include, when at least one set bit exists in the target fusion bitmap, selecting the candidate rule corresponding to the set bit with the smallest rule number as the target matching rule.
[0059] The bitmap and fusion logic are constructed. This includes performing a bitwise AND operation on several feature bitmaps determined by the input data packet, and then determining the target matching rule based on the bit setting information of the obtained first fused bitmap using a priority parsing table.
[0060] Deploy the feature bitmap lookup table, priority parsing table, and bitmap and fusion logic to the target switch.
[0061] In this specific embodiment, the present invention first assigns a corresponding bit position to each candidate rule in the desired rule table, and divides each network traffic feature in the desired rule table into regions, so as to construct a feature bitmap lookup table based on the correspondence between the constructed feature regions and bit positions. Specifically, the present invention first assigns a unique bit position to each rule in the compact ordered rule table (i.e., the desired rule table) through the programmable switch control plane. If the compact ordered rule table contains K rules, then each rule corresponds to a bit in a bitmap of length K. The smaller the rule number, the higher the priority of the rule in the ordered rule table. Next, for each traffic feature, the control plane collects all thresholds related to that traffic feature in the compact ordered rule table and sorts them. Then, based on the sorted thresholds, the control plane divides the value range of that traffic feature into multiple non-overlapping intervals. For example, if a certain feature f involves a threshold... , and ,and Then the range of values for this feature can be divided into less than or equal to , arrive between, arrive Between and greater than Multiple intervals are considered. Further, the control plane generates a bitmap of length K equal to the number of rules for each feature interval. For a given interval, if it does not violate the feasible region of a rule on that traffic feature, the corresponding bit of that rule is set to 1; if it violates the feasible region of that rule on that traffic feature, the corresponding bit of that rule is set to 0. Additionally, for features not covered by a rule, none of the intervals of that feature will violate the rule, so the corresponding bit of that rule can be set to 1. Then, the control plane converts the correspondence between feature intervals and bitmaps into matching action entries in the programmable switch, constructing a feature bitmap lookup table. Each entry can include interval matching conditions and the corresponding bitmap output action. Accordingly, the bitmap can be stored in action parameters, metadata fields, registers, or other data plane accessible resources.
[0062] Next, this embodiment of the invention constructs a priority parsing table according to preset priority rules. Specifically, the preset priority rules in this embodiment include, when there is at least one set bit in the target fused bitmap, selecting the candidate rule corresponding to the set bit with the smallest rule number as the target matching rule. Accordingly, this embodiment of the invention generates a priority parsing table through the control plane to parse the fused bitmap into a first matching rule number. For example, if there is at least one set bit in the fused bitmap, the rule corresponding to the set bit with the smallest rule number is selected as the first matching rule, i.e., the target matching rule. Accordingly, the control plane can pre-map common fused bitmap patterns to rule numbers, or the first set bit parsing can be implemented in the data plane through a hierarchical parsing table, a priority encoding table, or multiple matching action tables.
[0063] Furthermore, embodiments of the present invention construct bitmaps and fusion logic. Specifically, the bitmaps and fusion logic constructed in embodiments of the present invention includes performing a bitwise AND operation on several feature bitmaps determined from the input data packet, and then determining the target matching rule through a priority parsing table based on the set information of the obtained first fused bitmap. Here, set information refers to data information in the fused bitmap that contains set bits. For example, embodiments of the present invention obtain multiple feature bitmaps by performing interval searches on multiple traffic features of the input data packet in the data plane. Then, a bitwise AND operation is performed on the multiple feature bitmaps to obtain the first fused bitmap. Accordingly, if at least one set bit exists in the first fused bitmap, the rule corresponding to the set bit with the smallest rule sequence number is selected as the first matching rule, and the predicted category of that rule is output; otherwise, if the first fused bitmap is all zeros, the default category is output.
[0064] Finally, this embodiment of the invention deploys the constructed feature bitmap lookup table, priority parsing table, and bitmap and fusion logic to the target switch. Specifically, this embodiment of the invention generates a P4 program and runtime entries through the programmable switch control plane based on the feature bitmap lookup table, bitmap AND fusion logic (bitmap and fusion logic), and priority parsing table. The P4 program is compiled and deployed to the programmable switch based on the PISA architecture, and runtime entries are issued to the programmable switch through the runtime interface, enabling the classification logic to execute at line speed in the packet forwarding path. It is easy to understand that this embodiment of the invention solves the problem that deep decision trees or forests are difficult to deploy directly on resource-constrained programmable switches by distilling the classification knowledge in the high-performance tree model of the control plane into a compact and ordered rule table executable by the data plane, and achieving efficient rule matching through bitmap AND encoding.
[0065] It should be noted that during the analysis and processing of traffic data to be classified by the target switch, the programmable switch data plane performs feature lookup, bitmap AND fusion, and first-match parsing on the input data packet (i.e., the traffic data to be classified), and outputs the traffic category corresponding to the input data packet. For example, when the input data packet enters the programmable switch data plane, the data plane parser first extracts the packet header fields, such as source IP address, destination IP address, source port, destination port, protocol type, packet length, and TCP flags. Additionally, for flow-level features that need to be maintained, the data plane can look up or update registers, counters, and time window states based on the flow identifier to obtain the statistical characteristics corresponding to the current flow. Then, the data plane performs feature interval lookup based on multiple traffic features of the input data packet or data flow to obtain multiple feature bitmaps. Each feature bitmap represents a set of rules that may still be matched under that feature condition. Then, the data plane performs a bitwise AND operation on the multiple feature bitmaps to obtain a fused bitmap. This fused bitmap represents a set of candidate rules that simultaneously satisfy multiple feature conditions. Accordingly, if at least one set bit exists in the fused bitmap, the data plane selects the rule corresponding to the set bit with the smallest rule sequence number as the first matching rule, i.e., the target matching rule, through the priority parsing table. Subsequently, the data plane queries the decision table to obtain the predicted category corresponding to the rule and outputs the predicted category as the classification result of the input data packet or data stream. If the fused bitmap is all zeros, it indicates that the input data packet or data stream does not satisfy any rule in the compact ordered rule table, and the data plane outputs the default category. Furthermore, after the data plane outputs the classification result, this embodiment of the invention can also perform further actions based on the classification result. For example, for normal traffic, it can continue to forward according to the original forwarding logic; for attack traffic, it can perform dropping, rate limiting, mirroring, reporting to the control plane, or redirection; for unknown traffic, it can trigger further analysis by the control plane; and for application classification results, it can be used for quality of service control, path selection, or traffic statistics.
[0066] In some embodiments of the present invention, the bitmap and fusion logic further include: The corresponding block bitmap is obtained by querying the combination of feature intervals within the block based on the input data packet. The block bitmap is constructed by dividing each network traffic feature in the expected rule table into several feature blocks and performing a bitwise AND operation on the feature bitmap of each feature interval combination within each feature block.
[0067] Perform a bitwise AND operation on the block bitmap to obtain the second merged bitmap.
[0068] The target matching rule is determined by using the priority parsing table based on the bit setting information of the second fused bitmap.
[0069] In this specific embodiment, the present invention first uses the input data packet to generate the corresponding block bitmap by combining feature intervals within the block, and then performs a bitwise AND operation on the block bitmap to obtain a second fused bitmap. Then, based on the bit information of the second fused bitmap, the target matching rule is determined through a priority parsing table. Specifically, the present invention divides each network traffic feature in the expected rule table into several feature blocks, and performs a bitwise AND operation on the feature bitmap of each feature interval combination within each feature block to construct the block bitmap. The present invention uses block bitmap AND encoding to reduce the dependency depth in the data plane pipeline. Correspondingly, the control plane divides multiple traffic features into several feature blocks, and pre-calculates the block bitmap corresponding to multiple feature interval combinations within each feature block. This allows the data plane to only need to query the block bitmap commonly corresponding to multiple features in each feature block during data plane operation, and perform a bitwise AND operation on the outputs of different feature blocks, thereby reducing the pipeline dependency depth caused by continuous AND operations per feature.
[0070] For example, if a compact rule table involves 8 features, the control plane can divide it into 4 feature blocks, each containing 2 features. For each feature within a block, the control plane first divides the feature intervals according to the rule thresholds and generates a feature bitmap for each interval. The block bitmap is obtained by pre-operating a bitwise AND operation on multiple feature bitmaps within the same feature block, representing the set of rules that may still be hit after simultaneously satisfying all feature conditions within the feature block. Then, during data plane runtime, it is not necessary to obtain each feature bitmap within the block separately and perform AND operations step by step; instead, the corresponding block bitmap is queried directly based on the combination of feature intervals within the block. Subsequently, in this embodiment of the invention, only the block bitmaps output by different feature blocks are bitwise ANDed to obtain the final fused bitmap. Thus, block bitmap encoding advances some AND calculations from the data plane runtime to the control plane pre-computation stage, reducing pipeline dependency depth. At the same time, this embodiment of the invention can also make trade-offs between the number of entries, pipeline stages, metadata bit width, and runtime computational complexity by adjusting the feature block size.
[0071] The following section provides a detailed introduction and explanation of the solution in this embodiment of the invention, using a specific scenario of network traffic classification based on Bayesian rule distillation: For example, such as Figure 2 As shown in the embodiments of the present invention, the control plane and data plane of the programmable switch operate collaboratively. Specifically, the embodiments of the present invention first train a teacher model on the control plane based on training data and feature configuration, then construct a compact ordered rule table through Bayesian rule distillation and a greedy algorithm, and then convert the compact ordered rule table into a feature interval bitmap, a block-level bitmap, and decision table entries, and distribute them to the data plane, thereby performing feature interval lookup, bitmap bitwise AND operation, and classification decision on the programmable switch data plane to output network traffic classification results. (Refer to...) Figure 3The Bayesian rule distillation process in this embodiment includes steps such as candidate rule input, leaf node category distribution extraction, empirical Bayesian prior construction, first matching count statistics, Dirichlet posterior update, rule prediction category and confidence calculation, and rule output. Specifically, this embodiment inputs a candidate rule pool into the control surface. Each candidate rule in the pool originates from a leaf path of the teacher model and carries a category count vector of the source leaf node. This category count vector reflects the statistical experience of the teacher model in the local feature space. Next, the control surface normalizes the leaf node category count vector to a leaf node category distribution and constructs Dirichlet prior parameters based on a smoothing baseline term and the teacher model's prior weights. This embodiment enables candidate rules to make stable predictions using the category structure of the teacher model's leaf nodes even when there are no matching samples or few matching samples. Further, the control surface counts the matching of training samples according to the first matching semantics, i.e., the preset matching strategy. Unlike ordinary rule coverage statistics, the first matching statistics ensure that each training sample contributes to at most one rule, thus avoiding duplicate counting of multiple overlapping rules and ensuring that the training semantics of the subsequent ordered rule table is consistent with the priority matching semantics of the data surface. Then, the control surface weighted sums the prior parameters with the training sample matching counts to obtain the rule posterior parameters. Due to the use of the Dirichlet-Categorical conjugate structure, this posterior update process has a closed form, is computationally simple, and is suitable for handling a large number of candidate rules in the control surface. Finally, the control surface calculates the rule category probability vector, predicted category, and confidence score based on the posterior parameters. The rule category probability vector represents the rule's prediction tendency for each category, the predicted category is used in the decision table deployed to the data surface, and the confidence score is used for handling parallel rules in the greedy rule selection process and can also be used for subsequent rule auditing and manual maintenance.
[0072] It should be noted that the Bayesian rule distillation process provided in this embodiment of the invention can utilize the leaf node category distribution of the teacher model as an empirical prior, improving the stability of rule estimation in sparse regions of the samples. It can also provide a probabilistic interpretation for each rule, rather than just a hard label output, and can reduce the complexity of control surface rule modeling through closed posterior updates. Furthermore, by combining it with the first matching semantics, consistency between rule training and data surface deployment can be ensured.
[0073] Reference Figure 4Regarding the bitmap and encoding in this embodiment of the invention, when the compact ordered rule table contains K rules, the control plane assigns a bit position to each rule. Therefore, any bitmap can be represented as a binary vector of length K. If the k-th bit in the bitmap is 1, it indicates that the k-th rule may still be hit under the current conditions; if the k-th bit is 0, it indicates that the k-th rule cannot be hit under the current conditions. Specifically, for each traffic feature, the control plane divides disjoint intervals according to the threshold values appearing in the compact ordered rule table and generates a corresponding bitmap for each interval. For example, if a rule requires feature f to be greater than 10 and less than or equal to 20, then for the interval (10, 20] of feature f, the corresponding bit of the rule is 1; for intervals less than or equal to 10 or greater than 20, the corresponding bit of the rule is 0. If a rule does not contain a restriction on feature f, then none of the intervals of feature f will exclude the rule, so the corresponding bit of the rule is 1. Accordingly, during data plane operation, for each feature value of the input data packet or data stream, the data plane retrieves a bitmap from the corresponding feature bitmap lookup table. If a rule contains multiple feature conditions, then only when all relevant feature conditions are satisfied simultaneously will the corresponding bit of the rule be 1 in all feature bitmaps. This embodiment of the invention obtains the final fused bitmap by performing a bitwise AND operation on multiple feature bitmaps, as shown in the following formula: ; Where, in the formula Indicates the number of features involved in the classification. Indicates the first The rule bitmap found by each feature. Correspondingly, if... The presence of multiple set bits indicates that the input data packet or data stream simultaneously satisfies multiple rules. In this case, this embodiment of the invention selects the rule corresponding to the set bit with the smallest rule sequence number according to the first matching semantics. If If no set bit is specified, the default category will be output.
[0074] In addition, to reduce the pipeline dependency depth caused by data plane continuity and operations, embodiments of the present invention may also employ block bitmaps and encoding, such as... Figure 5As shown. Specifically, the control plane divides multiple features into several feature blocks and first generates corresponding feature bitmaps for each interval of each feature within a block. Then, a bitwise AND operation is pre-executed on multiple feature bitmaps corresponding to a certain set of feature intervals within the same feature block to obtain the block bitmap corresponding to that interval combination, and the mapping between the feature interval combination and the block bitmap is written into the data plane block bitmap lookup table. Correspondingly, when the data plane runs, instead of outputting the bitmaps of each feature within the block separately and performing AND operations step by step, a block bitmap is directly obtained by querying based on the feature interval combination within the block. Then, only the block bitmaps output from different feature blocks are subjected to bitwise AND operations to obtain the final fused bitmap. This embodiment of the invention advances some AND calculations from the data plane runtime to the control plane pre-computation stage, thereby reducing the number of runtime AND operations and pipeline dependency depth. Therefore, this embodiment of the invention can select an appropriate block granularity based on the number of pipeline stages, table capacity, PHV metadata bit width, and bitmap length of the programmable switch. It is easy to understand that, through bitmaps and encoding, the embodiments of the present invention convert multi-condition rule matching into programmable switch-friendly interval lookup and integer bitwise operations, avoiding complex tree traversal, loops, recursion, floating-point operations, matrix calculations or nonlinear functions in the data plane, thereby improving the hardware adaptability of rule deployment.
[0075] It should be noted that, by distilling high-precision decision trees or forest models into compact ordered rule tables, this embodiment of the invention alleviates the problem of rule expansion caused by directly deploying deep tree models or large-scale forest models, significantly reducing the size of rules and storage overhead while maintaining high classification accuracy. Simultaneously, by introducing the class distribution information learned from the leaf nodes of the teacher model into the rule modeling process through empirical Bayesian priors, candidate rules not only possess deterministic matching conditions but also quantifiable posterior class distributions and prediction confidence, thereby improving the reliability of rule selection and classification decisions. Furthermore, this embodiment employs a Dirichlet-class distribution conjugate modeling approach, enabling posterior updates of candidate rules to be completed by adding prior parameters to class counts, eliminating the need for complex iterative optimization, sampling, or deep neural network training processes, making it suitable for efficiently generating deployable rules on the control plane. Finally, this embodiment employs a greedy rule selection method based on macro-average F1 gain, ensuring that minority attack traffic, abnormal traffic, or low-frequency application categories are fully considered during rule selection, mitigating the problem of majority class traffic dominating classification results. Meanwhile, by employing a first matching semantic to construct and evaluate an ordered rule table, this embodiment of the invention ensures that the classification behavior in the control plane rule selection stage is consistent with the priority matching behavior in the programmable switch data plane, reducing semantic deviations between model training and hardware deployment. Furthermore, by using bitmaps and encoding mechanisms, multi-condition rule matching is converted into feature interval lookup and bitwise AND operations, avoiding complex tree traversal, matrix multiplication, nonlinear functions, or floating-point operations on the data plane, thus conforming to the hardware execution characteristics of the programmable switch's matching-action pipeline. Additionally, this embodiment of the invention achieves rule matching compression through feature interval bitmaps, block-level bitmaps, and bitwise AND operations, reducing the table entry explosion problem in traditional combined lookup methods and achieving flexible trade-offs between pipeline stages, SRAM, TCAM, and metadata width. Moreover, each rule generated in this embodiment originates from the leaf path of the teacher tree model and can be represented as a conjunct of multiple feature conditions, enabling network operators to intuitively understand the classification criteria and facilitating rule auditing, model debugging, and security policy maintenance. Furthermore, by deploying a compact rule table onto the data plane of the programmable switch, embodiments of the present invention enable online classification of data packets or data flows along the forwarding path. This reduces the need to upload raw traffic, packet content, or a large number of features to the control plane, thereby reducing communication overhead, classification response latency, and potential privacy exposure risks. Additionally, the network traffic classification method based on Bayesian rule distillation provided by embodiments of the present invention is applicable to various network traffic classification tasks such as intrusion detection, malicious traffic identification, application classification, and abnormal traffic detection. The number of rules can be configured according to the hardware resource budget of different programmable switches, exhibiting good cross-dataset and cross-scenario adaptability.
[0076] Please see Figure 6This application also provides a network traffic classification device based on Bayesian rule distillation, which can implement the above method. The device includes: The first module 210 is used to acquire network traffic training data and then extract the corresponding network traffic features; The second module 220 is used to train the preset tree model based on network traffic characteristics to construct the teacher model; The third module 230 is used to extract preset rule information from the leaf paths of the teacher model; wherein, the preset rule information includes candidate rules and leaf node category data; The fourth module 240 is used to construct empirical Bayesian prior data based on leaf node category data, and then update the posterior category distribution of candidate rules by combining it with preset sample matching data to determine preset posterior probability information. The preset sample matching data includes training sample matching data under a preset matching strategy. The fifth module 250 is used to select candidate rules based on preset posterior probability information using a greedy algorithm, and generate an expected rule table; The sixth module 260 is used to generate switch data plane deployment rules based on the desired rule table through a preset bitmap and encoding mechanism, and deploy them to the target switch; Module 7, 270, is used to input the traffic data to be classified into the target switch for analysis and processing to obtain the traffic classification results.
[0077] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0078] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0079] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0080] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 310 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 320 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 320 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 320 and is called and executed by the processor 310 using the methods described in the embodiments of this application. Input / output interface 330 is used to realize information input and output; The communication interface 340 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 350 transmits information between various components of the device (e.g., processor 310, memory 320, input / output interface 330, and communication interface 340); The processor 310, memory 320, input / output interface 330 and communication interface 340 are connected to each other within the device via bus 350.
[0081] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0082] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0083] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0084] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0085] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0087] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0088] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0089] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0090] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0091] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0092] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0093] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0094] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A network traffic classification method based on Bayesian rule distillation, characterized in that, The method includes the following steps: Obtain network traffic training data, and then extract the corresponding network traffic features; The teacher model is constructed by training the preset tree model based on the network traffic characteristics. Preset rule information is extracted from the leaf paths of the teacher model; wherein, the preset rule information includes candidate rules and leaf node category data; Based on the leaf node category data, empirical Bayesian prior data is constructed, and then combined with preset sample matching data to update the posterior category distribution of the candidate rules, thereby determining preset posterior probability information; wherein, the preset sample matching data includes training sample matching data under the preset matching strategy. Based on the preset posterior probability information, a greedy algorithm is used to select the candidate rules and generate an expected rule table. Based on the desired rule table, switch data plane deployment rules are generated using a preset bitmap and encoding mechanism, and then deployed to the target switch; The traffic data to be classified is input into the target switch for analysis and processing to obtain the traffic classification result; The step of generating switch data plane deployment rules based on the desired rule table using a preset bitmap and encoding mechanism, and deploying them to the target switch, includes: Assign a corresponding bit position to each candidate rule in the expected rule table, divide each network traffic feature in the expected rule table into regions, and then construct a feature bitmap lookup table based on the correspondence between the constructed feature regions and the bit positions. A priority parsing table is constructed according to a preset priority rule; wherein, the preset priority rule includes, when there is at least one set bit in the target fusion bitmap, taking the candidate rule corresponding to the set bit with the smallest rule number as the target matching rule; Constructing a bitmap and fusion logic; wherein, the bitmap and fusion logic includes performing a bitwise AND operation on several feature bitmaps determined by the input data packet, and then determining the target matching rule through the priority parsing table based on the bit setting information of the obtained first fused bitmap; The feature bitmap lookup table, the priority parsing table, and the bitmap and fusion logic are deployed to the target switch.
2. The method according to claim 1, characterized in that, The step of extracting preset rule information from the leaf path of the teacher model includes: The node judgment information on each rule path of the teacher model is converted into atomic predicates; wherein, the atomic predicates include feature indexes, comparison operators, and thresholds; Traverse the teacher model, convert each rule path from the root node to the leaf node of the teacher model into the corresponding candidate rule, and combine the atomic predicates corresponding to the rule path according to the path order to construct the preset rule antecedent; The category count vectors of the model leaf nodes corresponding to the candidate rules are counted to obtain the leaf node category data.
3. The method according to claim 1, characterized in that, The step of constructing empirical Bayesian prior data based on the leaf node category data, and then updating the posterior category distribution of the candidate rules by combining it with preset sample matching data to determine preset posterior probability information includes: The category count vectors of the model leaf nodes corresponding to the candidate rules are normalized to obtain the leaf node category distribution; Dirichlet prior parameters are constructed based on the smoothed baseline term and the leaf node category distribution; Based on the order of the candidate rules, each network traffic training sample is allocated according to the preset matching strategy to obtain the sample allocation result; Based on the sample allocation results, the training sample categories allocated to each candidate rule are statistically analyzed to obtain a matching count vector; Posterior parameters are constructed based on the matching count vector and the Dirichlet prior parameters, and the corresponding posterior mean is calculated using the posterior parameters; wherein the posterior mean serves as the class probability vector of the candidate rule.
4. The method according to claim 3, characterized in that, The step of allocating each network traffic training sample according to the order of the candidate rules and the preset matching strategy to obtain the sample allocation result includes: The coverage of each network traffic training sample is analyzed sequentially according to the order of the candidate rules to obtain sample coverage information; When it is determined from the sample coverage information that the network traffic training sample satisfies several candidate rules, the network traffic training sample is matched with the first coverage rule; or, when it is determined from the sample coverage information that the network traffic training sample does not satisfy any of the candidate rules, the network traffic training sample is matched with the default category branch; wherein, the first coverage rule includes the candidate rule that covers the network traffic training sample first in the order of arrangement.
5. The method according to claim 1, characterized in that, The step of selecting candidate rules based on the preset posterior probability information using a greedy algorithm to generate an expected rule table includes: Construct an ordered rule table; The preset F1 increment is evaluated according to the candidate rules to obtain preset increment data; wherein, the preset F1 increment includes the macro-average F1 increment generated on the validation set after adding the candidate rules that were not selected to the ordered rule table; Based on the preset incremental data and the preset posterior probability information, the desired selection rule is determined from each of the candidate rules, and then the desired selection rule is added to the ordered rule table to construct the desired rule table.
6. The method according to claim 1, characterized in that, The bitmap and fusion logic also include: The corresponding block bitmap is obtained by querying the combination of feature intervals within the block based on the input data packet; wherein, the block bitmap is constructed by dividing each network traffic feature in the expected rule table into several feature blocks, and performing a bitwise AND operation on the feature bitmap of each feature interval combination within each feature block; Perform a bitwise AND operation on the block bitmap to obtain a second fused bitmap; The target matching rule is determined by the priority parsing table based on the bit setting information of the second fused bitmap.
7. A network traffic classification device based on Bayesian rule distillation, characterized in that, The device includes: The first module is used to acquire network traffic training data and then extract the corresponding network traffic features. The second module is used to train the preset tree model based on the network traffic characteristics to construct the teacher model; The third module is used to extract preset rule information from the leaf paths of the teacher model; wherein, the preset rule information includes candidate rules and leaf node category data; The fourth module is used to construct empirical Bayesian prior data based on the leaf node category data, and then update the posterior category distribution of the candidate rules by combining it with preset sample matching data to determine preset posterior probability information; wherein, the preset sample matching data includes training sample matching data under the preset matching strategy. The fifth module is used to select candidate rules based on the preset posterior probability information using a greedy algorithm, and generate an expected rule table; The sixth module is used to generate switch data plane deployment rules based on the expected rule table through a preset bitmap and encoding mechanism, and deploy them to the target switch; The step of generating switch data plane deployment rules based on the desired rule table using a preset bitmap and encoding mechanism, and deploying them to the target switch, includes: Assign a corresponding bit position to each candidate rule in the expected rule table, divide each network traffic feature in the expected rule table into regions, and then construct a feature bitmap lookup table based on the correspondence between the constructed feature regions and the bit positions. A priority parsing table is constructed according to a preset priority rule; wherein, the preset priority rule includes, when there is at least one set bit in the target fusion bitmap, taking the candidate rule corresponding to the set bit with the smallest rule number as the target matching rule; Constructing a bitmap and fusion logic; wherein, the bitmap and fusion logic includes performing a bitwise AND operation on several feature bitmaps determined by the input data packet, and then determining the target matching rule through the priority parsing table based on the bit setting information of the obtained first fused bitmap; The feature bitmap lookup table, the priority parsing table, and the bitmap and fusion logic are deployed to the target switch; The seventh module is used to input the traffic data to be classified into the target switch for analysis and processing to obtain the traffic classification result.
8. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
24 solar term health guidance system based on artificial intelligence
CN121565395A
Tree model based network traffic classification data distillation and online model updating method
CN122508355A