A method for improving the packet classification rate based on decision trees

By converting the original policy set into a new policy set with no intersection and creating a decision tree based on the new policy set, the problem of too many decision tree layers and low search rate in the existing technology is solved, and more efficient packet classification is achieved.

CN115412423BActive Publication Date: 2025-06-17BEIJING ZUOJIANG TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210979061.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2025-06-17
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

The existing data packet classification algorithm based on decision tree has problems with strategy replication and too many layers, resulting in too low search rate.

Method used

By converting the original policy set to a new policy set, make sure that the policies in the new policy set have no intersection, and create a decision tree based on this new policy set, avoiding policy replication and reducing the number of layers of the decision tree.

Benefits of technology

It effectively reduces the average number of layers and maximum number of layers of the decision tree, improves the average search rate and worst search rate of data packets, and increases the rate by 20% and 10%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115412423B_ABST
    Figure CN115412423B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for improving the packet classification rate based on a decision tree, belonging to the field of communication technologies. The present invention converts the original policy set into a new policy set, and there is no intersection among the policies in the new policy set. During the process of creating a decision tree using the new policy set, since there is no intersection among the policies, the policies can be evenly divided from the parent node to two child nodes, reducing the average number of levels and the maximum number of levels of the decision tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to a method for improving the classification rate of data packets based on a decision tree. Background Art

[0002] The data packet classification algorithm is a commonly used algorithm in a network. Its purpose is to classify data packets into different categories according to the header information carried by the packets, generally including five dimensions: source IP, destination IP, source port, destination port, and protocol number, so that different processing can be performed on them. Common applications include: access control lists, firewalls, and flow-based charging statistics. It has different applications in both core network devices and edge network devices.

[0003] To perform data packet classification, a set of rule sets needs to be defined first, and then searched in the rule sets according to the header information of the data packets. Figure 1 Taking it as a rule example, each field represents in turn: source IP address value / source IP mask, destination IP address value / destination IP mask, source port number range, destination port number range, and protocol number value / protocol number mask value. Each rule represents a range, and the rule with a smaller mask represents a larger range.

[0004] Among many data packet classification algorithms, the hypersplit decision tree is a commonly used one.

[0005] However, the hypersplit decision tree has two obvious problems:

[0006] 1) There are intersections in the ranges represented by the policies in the original user policy set. Therefore, when dividing the policy of the parent node into two child nodes during the decision tree creation process, policy replication will occur, which will further lead to an excessive number of layers of the decision tree.

[0007] 2) When creating a decision tree using the original user policy set, when dividing the policy of the parent node into two child nodes, two factors, namely the uniformity of the division and the small degree of policy replication, need to be considered simultaneously. If only the uniformity is considered, the policy of the parent node will be evenly distributed to the two child nodes, and the number of layers of the decision tree is log2 N, where N is the total number of policies in the policy set. When additionally considering a small degree of replication, the policy of the parent node can no longer be evenly distributed to the two child nodes, which will lead to an excessive maximum number of layers of the decision tree, and further lead to a too low worst-case search rate for data packet classification based on the decision tree.

[0008] 3) The creation of the decision tree is expected to approximate the global optimum with the local optimum layer by layer. The local optimum layer by layer means that the policy division of each node is planned in the best way according to the information of the current node, and this best division method is ensured by the information entropy. The global optimum means that the average number of layers and the maximum number of layers of the entire decision tree are the smallest among all tree-building methods. After the tree building is completed, multiple policies of the leaf nodes can be merged into one policy based on the priority. This merging scheme reduces the storage requirement of the entire decision tree and improves the query performance. However, due to the change in the number of rules contained in each branch of the decision tree caused by the merging, the original balance is destroyed, and there is a possibility of further optimization of the decision tree. Summary of the Invention

[0009] (1) Technical problems to be solved

[0010] The technical problem to be solved by the present invention is: how to design a method for improving the data packet classification rate based on a decision tree.

[0011] (2) Technical solutions

[0012] To solve the above technical problems, the present invention provides a method for improving the data packet classification rate based on a decision tree, including the following steps:

[0013] 1) Create the first decision tree according to the original policy set input by the user;

[0014] 2) Traverse the first decision tree and output the policies in each leaf node of the decision tree to the policy file;

[0015] 3) Release the first decision tree;

[0016] 4) Read the policy file to obtain a new policy set, and there is no intersection among the policies in the new policy set;

[0017] 5) Create the second decision tree according to the new policy set and perform data packet lookup in the second decision tree.

[0018] Preferably, the structures of the first decision tree and the second decision tree are the same. In the decision tree, the basis for becoming a leaf node is: the node has only one policy; or, the node has multiple policies, but the policies cannot be further divided among the child nodes, that is, each child node has the same policy as the parent node.

[0019] Preferably, in the decision tree, the root node has all the policies in the policy array, and the same policy can be divided among multiple child nodes.

[0020] Preferably, the creation processes of the first decision tree and the second decision tree are the same, and the creation process is as follows:

[0021] 11) Create a root node and assign all the policies in the policy set to the root node;

[0022] 12) Calculate the maximum information entropy of the current node;

[0023] 13) If the maximum information entropy is 0, mark the current node as a leaf node; otherwise, execute step 14);

[0024] 14) Create left and right child nodes, and use the edge value corresponding to the maximum information entropy as the division edge value; use the division edge value to divide the policies of the parent node into the left and right child nodes;

[0025] 15) Respectively use the left and right child nodes as the new parent nodes;

[0026] 16) Repeat steps 12) to 15) until the ends of all paths of the hypersplit decision tree are leaf nodes.

[0027] Preferably, in step 14, the process of selecting the division edge value for dividing the policies of the parent node into two child nodes by the maximum information entropy is as follows:

[0028] Step 141: Assume that the parent node contains several policies, and the number of policies is denoted as cur. These policies are to be divided into the left and right child nodes; count the different edge values of the policies of the parent node in different dimensions. Assume that the number of different edge values is num_edge; define num_edge pairs of counters, and each edge value corresponds to a pair of counters {lc, rc}, where lc represents the number of policies of the parent node divided into the left child node using the specified edge value, and rc represents the number of policies of the parent node divided into the right child node using the specified edge value; initialize the num_edge pairs of counters to 0, and execute step 142;

[0029] Step 142: Obtain a policy filt in the parent node and execute step 143;

[0030] Step 143: Select an edge value cut_value in step 141 and execute step 144;

[0031] Step 144: Assume that dim represents the dimension to which cut_value belongs; filt is a variable that stores the low edge value and high edge value of all dimensions of the policy. filt->value[dim][0] represents the low edge value of the policy filt in dimension dim, and filt->value[dim][1] represents the high edge value of the policy filt in dimension dim;

[0032] If the low edge value of the policy in dimension dim is less than or equal to the candidate partition edge value cut_value, it indicates that the policy filt will be partitioned to the left child node, then step 146 is executed; otherwise, step 145 is executed;

[0033] Step 145: If the high edge value of the policy in dimension dim is greater than the candidate partition edge value cut_value, it indicates that the policy filt will be partitioned to the right child node, then step 147 is executed; otherwise, step 146 and step 147 are respectively executed, indicating that the policy filt will be partitioned to both the left child node and the right child node;

[0034] Step 146: The left counter lcn is incremented by one, and step 148 is executed;

[0035] Step 147: The right counter rcn is incremented by one, and step 148 is executed;

[0036] Step 148: Repeat steps 143 to 147 until the counting statistics of the num_edge edge values for the specified filt are completed, and then step 149 is executed;

[0037] Step 149: Repeat steps 142 to 148 until the counting statistics of all the policies of the parent node for the num_edge edge values are completed, and then step 1410 is executed;

[0038] Step 1410: Obtain a total of num_edge pairs of counters, and then step 1411 is executed;

[0039] Step 1411: Calculate num_edge information entropies based on the num_edge pairs of counters, and the edge value corresponding to the maximum information entropy is the partition edge value of the parent node policy.

[0040] Preferably, the calculation formula of the information entropy is as follows:

[0041]

[0042] Where:

[0043] Represents the percentage of the parent node policy partitioned to the left child node;

[0044] Represents the percentage of the parent node policy partitioned to the right child node;

[0045] 0 < a <= 1, 0 < b <= 1, a + b >= 1.

[0046] Preferably, in the calculation formula of the information entropy The part represents the degree of uniformity of the division of the parent node's policy into two child nodes. The larger this value, the more uniform the policy division. The log2(a + b) part in the formula represents the degree of replication when the parent node's policy is divided into two child nodes. The larger this value, the higher the degree of replication.

[0047] Preferably, when a policy of the parent node is simultaneously divided into two child nodes, it is said that the policy has undergone one replication.

[0048] The present invention also provides a system for implementing the method to improve the data packet classification rate based on a decision tree.

[0049] The present invention also provides an application of the method in the field of communication technology.

[0050] (III) Beneficial effects

[0051] 1. Through information entropy, the optimal division edge value can be calculated, and the policy of the parent node can be divided as evenly as possible into two child nodes.

[0052] 2. When the information entropy of a certain node is zero, the node is marked as a leaf node. The information entropy being zero can ensure that multiple rules in the leaf node can be merged into one.

[0053] 3. Create a decision tree according to the original user policy set, output the policies in the leaf node to the policy file, and generate a new policy set. The policies in the new policy set have no intersection. By creating the decision tree for the first time, the original policy set with intersections is converted into a new policy set without intersections.

[0054] 4. When creating a decision tree according to the new policy set, since the policies in the new policy set have no intersection, when the policy of the parent node is divided into child nodes, no policy replication will occur. The policies in the parent node are almost evenly divided between the two child nodes, reducing the average number of levels and the maximum number of levels of the decision tree, and improving the average search rate and the worst-case search rate of the data packet. The average number of levels and the maximum number of levels of the decision tree created according to the policy set without intersections are reduced by 20% and 10% respectively compared to the average number of levels and the maximum number of levels of the decision tree created according to the original policy set. Therefore, the present invention can increase the average rate and the worst rate of data packet classification based on a decision tree by 20% and 10% respectively. Description of the drawings

[0055] Figure 1 It is an example of a five-tuple rule;

[0056] Figure 2 It is the overall processing flow chart of the method of the present invention;

[0057] Figure 3 It is the overall structural schematic diagram of the decision tree designed by the present invention;

[0058] Figure 4 Overall flowchart created for the decision tree designed for the present invention;

[0059] Figure 5 Flowchart for selecting the optimal division edge value by maximum information entropy designed for the present invention;

[0060] Figure 6 Schematic diagram of the changes after the policy is divided into child nodes. Detailed implementation manners

[0061] To make the objectives, content and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments.

[0062] A method for improving the data packet classification rate based on a decision tree provided by the present invention converts the original policy set into a new policy set, and there is no intersection among the policies in the new policy set. During the process of creating a decision tree using the new policy set, since there is no intersection among the policies, the policies can be evenly divided from the parent node to two child nodes, reducing the average number of levels and the maximum number of levels of the decision tree. The present invention achieves the purpose of reducing the average number of levels and the maximum number of levels of the decision tree by creating the decision tree twice. The overall processing flow is as Figure 2 shown and includes the following steps:

[0063] 1) Create the first decision tree according to the original policy set input by the user;

[0064] 2) Traverse the first decision tree and output the policies in each leaf node of the decision tree to a policy file;

[0065] 3) At this time, the purpose of creating the first decision tree is achieved, so the first decision tree is released;

[0066] 4) Read the policy file to obtain a new policy set. By creating the first decision tree, the original policy set is converted into a new policy set. There is no intersection among the policies in the new policy set, which is determined by the creation principle of the decision tree. The creation process of the decision tree is briefly described below.

[0067] 5) Create the second (hypersplit) decision tree according to the new policy set; The search for data packets is performed in this second decision tree.

[0068] The following uses an example diagram to illustrate the overall structure of the decision tree. The structures of the first and second decision trees are the same.

[0069] In Figure 3 , it is assumed that there are a total of 10 policies in the policy set, and the above decision tree is created based on these 10 policies. The basic information of the decision tree in Figure 3 is described as follows:

[0070] 1) A circle represents a node in the decision tree; the number above the circle represents the index of the policy array, that is, it means the node has these policies;

[0071] 2) The "leaf" node represents a leaf node, and the "cld" node represents a child node;

[0072] 3) The basis for becoming a leaf node:

[0073] 3.1) The node has only one policy; for example, the first node in the third layer has only one policy, which is policy 1, so it is marked as a leaf node;

[0074] 3.2) The node has multiple policies, but the policies cannot be further divided among the child nodes, that is, the child nodes all have the same policies as the parent node; for example, the third node in the third layer has policies 1 and 3;

[0075] 4) The root node has all the policies in the policy array;

[0076] 5) Policies can be copied: the same policy can be assigned to multiple child nodes; for example, policy 1 in the first node of the first layer is assigned to its two child nodes.

[0077] 6) To reduce the tree height and thus improve the search rate of data packets, it is expected that the policies of the parent node can be evenly divided among the child nodes.

[0078] As Figure 4 shown, the creation process of the hypersplit decision tree is described as follows (the creation processes of the first decision tree and the second decision tree are the same):

[0079] 11) Create a root node and assign all the policies in the policy set to the root node;

[0080] 12) Calculate the maximum information entropy of the current node;

[0081] 13) If the maximum information entropy is 0, mark the current node as a leaf node; otherwise, execute step 14);

[0082] 14) Create left and right child nodes, and use the edge value corresponding to the maximum information entropy as the division edge value; use the division edge value to divide the policies of the parent node into the left and right child nodes;

[0083] 15) Process the left and right child nodes in a depth-first recursive method, and use the left and right child nodes as new parent nodes respectively;

[0084] 16) Repeat steps 12) to 15) until the ends of all paths of the hypersplit decision tree are leaf nodes.

[0085] As Figure 5 shown, in step 14, the process of selecting the division edge value of the parent node to two child nodes by the maximum information entropy strategy is as follows:

[0086] Step 141: Assume that the parent node contains several strategies, and the number of strategies is denoted as cur. These strategies are to be divided into left and right child nodes. Count the different edge values of the parent node's strategies in different dimensions. Assume the number of different edge values is num_edge. Define num_edge pairs of counters, and each edge value corresponds to a pair of counters {lc, rc}. lc represents the number of parent node strategies divided into the left child node using the specified edge value, and rc represents the number of parent node strategies divided into the right child node using the specified edge value. Initialize the num_edge pairs of counters to 0, and then execute step 142;

[0087] Step 142: Obtain a strategy filt in the parent node, and then execute step 143;

[0088] Step 143: Select an edge value cut_value in step 141, and then execute step 144;

[0089] Step 144: Let dim represent the dimension to which cut_value belongs; filt is a variable that stores the low edge value and high edge value of all dimensions of the strategy. filt->value[dim][0] represents the low edge value of the strategy filt in dimension dim, and filt->value[dim][1] represents the high edge value of the strategy filt in dimension dim;

[0090] If the low edge value of the strategy in dimension dim is less than or equal to the candidate division edge value cut_value, it means that the strategy filt will be divided into the left child node, and then execute step 146; otherwise, execute step 145;

[0091] Step 145: If the high edge value of the strategy in dimension dim is greater than the candidate division edge value cut_value, it means that the strategy filt will be divided into the right child node, and then execute step 147; otherwise, execute step 146 and step 147 respectively, which means that the strategy filt will be divided into both the left child node and the right child node;

[0092] Step 146: Increment the left counter (denoted as: lcn) by one, and then execute step 148;

[0093] Step 147: Increment the right counter (denoted as: rcn) by one, and then execute step 148;

[0094] Step 148: Repeat steps 143 to 147 until the counting statistics of the num_edge edge values for the specified filt are completed, and then execute step 149;

[0095] Step 149: Repeat steps 142 to 148 until the counting statistics of all the policies of the parent node for the num_edge edge values are completed, and then execute step 1410;

[0096] Step 1410: Obtain a total of num_edge pairs of counters, and then execute step 1411;

[0097] Step 1411: Calculate num_edge information entropies based on the num_edge pairs of counters. The edge value corresponding to the maximum information entropy is the division edge value of the parent node's policy.

[0098] The calculation formula for the information entropy is as follows:

[0099]

[0100] Among them:

[0101] 1) cur represents the number of policies of the parent node, lc represents the number of policies of the left child node, and a represents the percentage of the parent node's policy divided to the left child node;

[0102] 2) cur represents the number of policies of the parent node, rc represents the number of policies of the right child node, and b represents the percentage of the parent node's policy divided to the right child node;

[0103] 3) 0 < a <= 1, 0 < b <= 1, a + b >= 1;

[0104] 4) In the formula, The part represents the degree of uniformity of the parent node's policy divided into two child nodes. The larger this value is, the more uniform the policy division is; the log2(a + b) part in the formula represents the degree of replication when the parent node's policy is divided into two child nodes (when a policy of the parent node is divided into two child nodes at the same time, it is called that the policy has undergone one replication). The larger this value is, the higher the replication degree is. The complete information entropy formula expects that the selected division bit can divide the parent node's policy evenly and with a low replication degree to the left and right child nodes.

[0105] The following describes the changes after the policy is divided from the parent node to the child node. As Figure 6 shown, for the sake of simplifying the description process, the policies in the following examples only contain one dimension. The division of policies with multiple dimensions is similar.

[0106] In Figure 6Among them, the parent node contains two policies. The range of Policy 1 is from 0 to 100, and the range of Policy 2 is from 1 to 150. By using the method of calculating the maximum information entropy to select the division edge value introduced above, it can be known that dividing the two policies with the right edge value 100 of Policy 1 is the best. The division process is as follows:

[0107] 1) The right edge value 100 of Policy 1 is less than or equal to the division edge value 100, and Policy 1 is divided into the left child node;

[0108] 2) The range represented by Policy 2 contains the division edge value 100. Therefore, Policy 2 will be divided into two parts: 0 to 100 is divided into the left child node, and 101 to 150 is divided into the right child node;

[0109] The ranges represented by the two policies in the left child node are the same, and only the policy with the highest priority is retained. After that, the left child node is marked as a leaf node.

[0110] The right node contains only one policy and is directly marked as a leaf node.

[0111] Output the policies in the two leaf nodes to the policy file to generate a new policy set. The original policy set contains two policies: 0 to 100 and 0 - 150, and there is an intersection between these two policies. The policies in the new policy set are 0 - 100 and 101 to 150, and there is no intersection between these two policies.

[0112] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A method for improving the data packet classification rate based on a decision tree, characterized in that, It includes the following steps: 1) Create the first decision tree according to the original policy set input by the user; 2) Traverse the first decision tree and output the policies in each leaf node of the decision tree to the policy file; 3) Release the first decision tree; 4) Read the policy file to obtain a new policy set, and there is no intersection among the policies in the new policy set; 5) Create the second decision tree according to the new policy set and perform packet lookup in the second decision tree; The creation processes of the first decision tree and the second decision tree are the same, and the creation process is as follows: 11) Create the root node and assign all the policies in the policy set to the root node; 12) Calculate the maximum information entropy of the current node; 13) If the maximum information entropy is 0, mark the current node as a leaf node; Otherwise, execute step 14); 14) Create the left and right child nodes, and use the edge value corresponding to the maximum information entropy as the division edge value; use the division edge value to divide the policies of the parent node into the left and right child nodes; 15) Respectively use the left and right child nodes as the new parent nodes; 16) Repeat steps 12) to 15) until the ends of all paths of the hypersplit decision tree are leaf nodes; In step 14, the process of selecting the division edge value for dividing the policies of the parent node into two child nodes through the maximum information entropy is as follows: Step 141: Assume that the parent node contains several policies, and the number of policies is denoted as cur. These policies are to be divided into the left and right child nodes; count the different edge values of the policies of the parent node in different dimensions. Assume that the number of different edge values is num_edge; define num_edge pairs of counters, and each edge value corresponds to a pair of counters {lc, rc}. lc represents the number of policies of the parent node divided into the left child node using the specified edge value, and rc represents the number of policies of the parent node divided into the right child node using the specified edge value; initialize the num_edge pairs of counters to 0 and execute step 142; Step 142: Obtain a policy filt in the parent node and execute step 143; Step 143: Select an edge value cut_value in step 141 and execute step 144; Step 144: Let dim represent the dimension to which cut_value belongs; filt is a variable that stores the low edge value and high edge value of all dimensions of the policy. filt->value[dim][0] represents the low edge value of the policy filt in dimension dim, and filt->value[dim][1] represents the high edge value of the policy filt in dimension dim; If the low edge value of the policy in dimension dim is less than or equal to the to-be-selected division edge value cut_value, it means that the policy filt will be divided into the left child node, and then execute step 146; Otherwise, execute step 145; Step 145: If the high edge value of the policy in dimension dim is greater than the to-be-selected division edge value cut_value, it means that the policy filt will be divided into the right child node, and then execute step 147; Otherwise, execute Step 146 and Step 147 respectively, indicating that policy filt will be divided into the left child node and the right child node simultaneously; Step 146: Increment the left counter lcn by one and execute Step 148; Step 147: Increment the right counter rcn by one and execute Step 148; Step 148: Repeat Step 143 to Step 147 until the counting statistics of the num_edge edge values for the specified filt are completed, and then execute Step 149; Step 149: Repeat Step 142 to Step 148 until the counting statistics of all the policies of the parent node for the num_edge edge values are completed, and then execute Step 1410; Step 1410: Obtain a total of num_edge pairs of counters and execute Step 1411; Step 1411: Calculate num_edge information entropies based on the num_edge pairs of counters, and the edge value corresponding to the maximum information entropy is the division edge value of the parent node's policy.

2. The method according to claim 1, characterized in that, The structures of the first decision tree and the second decision tree are the same. In a decision tree, the basis for becoming a leaf node is as follows: a node has only one policy; or, a node has multiple policies, but the policies cannot be further divided among the child nodes, that is, each child node has the same policies as the parent node.

3. The method according to claim 2, characterized in that, In a decision tree, the root node has all the policies in the policy array, and the same policy can be divided among multiple child nodes.

4. The method according to claim 1, characterized in that, The calculation formula for information entropy is as follows: Where: Indicates the percentage of the parent node policy divided to the left child node; Indicates the percentage of the parent node policy partitioned to the right child node; 0 < a <= 1, 0 < b <= 1, a + b > 1.

5. The method according to claim 4, characterized in that The part in the formula for calculating information entropy represents the degree of uniformity of the parent node strategy divided into two child nodes. The larger this value is, the more uniform the strategy division is. The part log2(a + b) in the formula represents the degree of replication when the parent node strategy is divided into two child nodes. The larger this value is, the higher the degree of replication is.

6. The method according to claim 5, characterized in that When a policy of the parent node is divided into two child nodes simultaneously, it is said that the policy has undergone a replication.

7. A system for implementing the method according to any one of claims 1 to 6, which improves the data packet classification rate based on a decision tree.

Citation Information

Patent Citations

  • Data packet classification method based on information entropy

    CN114637773A