Message classification method and device, electronic device, and readable medium
By using a multi-bit dwarf tree set for packet classification, the existing decision tree algorithm has solved the problem of tree shape and density control, and efficient and accurate packet classification is achieved, and dynamic rule updates are supported.
Patent Information
- Application Number
- CN202010348580.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-04-27
AI Technical Summary
Existing decision tree algorithms have difficulty controlling the shape and density of the tree in message classification, resulting in limited performance and cannot support dynamic download rule lists.
Multi-bit dwarf tree set is used to classify packets, find matching message classification rules sets by searching key values, and use the shorter tree shape of multi-bit dwarf tree to reduce search delay.
It improves the efficiency and accuracy of message classification, supports dynamic update rules, does not need to obtain all rules in advance, and has good scalability.
Smart Images

Figure CN113642594B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a message classification method and device, an electronic device, and a readable medium. Background Art
[0002] With the gradual increase in the bandwidth of physical transmission media and the explosive growth in the number of network terminals, more and more rules need to be followed in communication networks. At the same time, the evolution of communication network protocols has also brought more and more different classification requirements, and the classification problem has gradually become a bottleneck restricting the transmission rate of communication networks. At present, the classification of Internet Protocol (IP) packets can be achieved through the following methods, including: hardware parallel comparison method based on ternary content addressable memory (TCAM), hash-based algorithm, space decomposition-based algorithm, and decision tree-based algorithm. Among them, the core idea of the decision tree-based algorithm is to continuously divide the rule set by selecting appropriate decision content until all rules are distinguished.
[0003] However, most existing decision tree algorithms are implemented in software, which is inferior to hardware solutions in terms of packet classification efficiency and stability. In addition, the few feasible decision tree-based packet classification methods that are implemented in hardware cannot support dynamic download of rule lists, and can only obtain all rule sets in advance and perform a lot of preprocessing. This makes it impossible to control the shape and density of the decision tree, resulting in limited performance of the decision tree-based algorithm. Summary of the invention
[0004] The embodiments of the present application provide a message classification method and device, an electronic device, and a readable medium, which are used to overcome the problem in the prior art that the shape and density of a decision tree are difficult to control, resulting in limited performance of an algorithm based on a decision tree.
[0005] In a first aspect, an embodiment of the present application provides a message classification method, the method comprising: searching a first multi-bit dwarf tree set based on a search key value of a message to be classified, to obtain a second message classification rule set that matches the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, which is a decision tree constructed based on the rules in the first message classification rule set; classifying the message to be classified based on the rules in the second message classification rule set, and determining the message type of the message to be classified.
[0006] In a second aspect, an embodiment of the present application provides a message classification device, comprising: a search module, used to search a first multi-bit dwarf tree set based on a search key value of a message to be classified, and obtain a second message classification rule set that matches the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, which is a decision tree constructed based on the rules in the first message classification rule set; a classification module, used to classify the message to be classified based on the rules in the second message classification rule set, and determine the message type of the message to be classified.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0009] The method provided in the embodiment of the present application obtains a second message classification rule set that matches the message to be classified by searching a first multi-bit dwarf tree set based on the search key value of the message to be classified, wherein the first multi-bit dwarf tree set includes multiple multi-bit dwarf trees, and the multi-bit dwarf tree is a decision tree constructed based on the first message classification rule set. By utilizing the shorter tree shape of the multi-bit dwarf tree, the hardware structure using the message classification method has a shorter pipeline stage, thereby reducing the search delay. The message to be classified is classified according to the rules in the second message classification rule set to determine the message type of the message to be classified, thereby ensuring the correctness and accuracy of the classification of the message to be classified. When targeting different message classification rules, there is no need to change the hardware resource settings, and it has good scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings, in which:
[0011] Figure 1 A flowchart of a message classification method according to an embodiment of the present application.
[0012] Figure 2 This is a schematic diagram of the structure of a multi-bit dwarf tree in this application.
[0013] Figure 3This is a flow chart of a method for inserting message classification rules to be updated into a multi-bit dwarf tree in the present application.
[0014] Figure 4 A flowchart of a message classification method according to another embodiment of the present application.
[0015] Figure 5 A block diagram of the composition of a message classification device provided in one embodiment of the present application.
[0016] Figure 6 This is a block diagram of the search module implemented in the form of a multi-bit dwarf tree in this application.
[0017] Figure 7 A block diagram of the composition of a message classification device provided in another embodiment of the present application.
[0018] Figure 8 This is a structural diagram of an exemplary hardware architecture of an electronic device according to the message classification method and device in the embodiments of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solution of the present application, the method and device, electronic device, and computer-readable medium provided by the present application are described in detail below in conjunction with the accompanying drawings.
[0020] Example embodiments will be described more fully below with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, the purpose of providing these embodiments is to make this application thorough and complete and to enable those skilled in the art to fully understand the scope of this application.
[0021] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in a sequence other than the content illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning unless explicitly defined herein.
[0023] Figure 1 FIG. 1 is a flow chart of a message classification method according to an embodiment of the present application. The message classification method can be applied to a message classification device. Figure 1 As shown, the message classification method includes the following steps.
[0024] Step 110: search the first multi-bit dwarf tree set according to the search key value of the message to be classified, and obtain a second message classification rule set that matches the message to be classified.
[0025] The message to be classified includes a search key value, which is a number of fields extracted from the complete message to be classified, such as certain identifiers in the header information of the message to be classified. The first multi-bit dwarf tree set includes a multi-bit dwarf tree, which is a decision tree constructed according to the rules in the first message classification rule set.
[0026] Compared with the traditional single-bit tree (i.e., binary tree), each node in the multi-bit dwarf tree can have multiple subtree branches. In addition, the multi-bit dwarf tree has fewer layers of tree nodes, for example, only 3 layers of tree nodes, or 4 layers of tree nodes. By processing the rules in the first message classification rule set, the decision bit corresponding to the root node of the multi-bit dwarf tree is obtained; based on the number of decision bits, the number of layers of the tree nodes of the multi-bit dwarf tree is specifically determined. The number of decision bits is an integer greater than or equal to 1.
[0027] Figure 2 This is a schematic diagram of the structure of a multi-bit dwarf tree in specific implementation. Figure 2 As shown, the multi-bit dwarf tree includes a three-layer tree structure, specifically including a 0th layer node (i.e., a root node 200), a 1st layer node, and a 2nd layer node. Among them, the 1st layer node includes multiple first layer tree nodes 210; the 2nd layer node includes multiple leaf nodes 220. The leaf nodes 220 of the multi-bit dwarf tree are all located in the last layer, and the multi-bit dwarf tree is a full tree. The advantage of building a multi-bit dwarf tree into a full tree is that only simple decision logic is required to directly store the rules in the first message classification rule set into a random access memory (Random Access Memory, RAM), without considering the problem of uncertainty in the leaf level of the original decision tree-based algorithm, so that the shape and density of the decision tree are controlled, thereby improving the performance of the algorithm.
[0028] In some specific implementations, step 110 can be implemented in the following manner: based on the search key value, the multi-bit dwarf trees in the first multi-bit dwarf tree set are searched in parallel to obtain a primary message classification rule set; the search key value is compared with the rules in the primary message classification rule set to filter and obtain a second message classification rule set, wherein the rules in the second message classification rule set are rules that match the messages to be classified.
[0029] Specifically, with the search key value as the index, the tree nodes at all levels and each leaf node of each multi-bit dwarf tree in the first multi-bit dwarf tree set are searched in parallel until the message classification rule corresponding to the search key value is found, and these message classification rules constitute the primary message classification rule set. Then, the search key value is finely compared with each message classification rule in the primary message classification rule set (for example, a bit-level comparison is performed), and the second message classification rule set is obtained by screening. The rules in the second message classification rule set can be more suitable for classifying the messages to be classified.
[0030] Step 120: classify the message to be classified according to the rules in the second message classification rule set, and determine the message type of the message to be classified.
[0031] In some specific implementations, step 120 can be implemented in the following manner: obtain the priority level corresponding to the rules in the second message classification rule set; based on the priority level, perform priority arbitration on the rules in the second message classification rule set to generate the optimal message classification rule; classify the messages to be classified based on the optimal message classification rule to determine the message type of the messages to be classified.
[0032] It should be noted that the priority levels corresponding to the rules in the second message classification rule set can be set manually or determined based on the degree of matching between different rules and the messages to be classified. The priority levels here are only examples. Other ways of setting priority levels are also within the scope of protection of this application and will not be repeated here.
[0033] Specifically, a message classification method with a small number of escape buckets or a TCAM message classification method can also be used to classify messages to be classified, and generate a small number of escape bucket classification rules and TCAM classification rules. Then, the optimal message classification rule, a small number of escape bucket classification rules and TCAM classification rules are prioritized again to generate a final message classification rule; the final message classification rule is used to classify the messages to be classified to obtain a more accurate message type of the messages to be classified. It should be noted that the above description of the message classification method is only an example, and can be specifically set according to specific needs. Other message classification methods that are not illustrated can also be used in combination with the message classification method in this application, and the generated multiple message classification rules are prioritized to obtain a more accurate message classification method, which will not be repeated here.
[0034] In this embodiment, by obtaining a message to be classified, searching a first multi-bit dwarf tree set according to the search key value of the message to be classified, a second message classification rule set matching the message to be classified is obtained, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, and the multi-bit dwarf tree is a decision tree constructed according to the first message classification rule set. By utilizing the shorter tree shape of the multi-bit dwarf tree, the hardware structure using the message classification method has a shorter pipeline stage, thereby reducing the search delay. The message to be classified is classified according to the rules in the second message classification rule set to determine the message type of the message to be classified, thereby improving the classification correctness and accuracy of the message to be classified. When targeting different message classification rules, there is no need to change the hardware resource settings, and it has good scalability.
[0035] In one embodiment, before step 120, the following may also be included: step 130, inserting the message classification rule to be updated into the multi-bit dwarf tree; or, step 140, deleting the message classification rule to be deleted stored in the multi-bit dwarf tree.
[0036] In specific implementation, the constructed multi-bit dwarf trees can be associated in a linked list manner to form a multi-bit dwarf tree set, and the message classification rules to be updated can be inserted into a tree in the multi-bit dwarf tree set. Other collection forms can also be used to implement the insertion of message classification rules. The insertion of rules may trigger the reconstruction of one or some tree nodes of a single multi-bit dwarf tree, or may trigger the partial reconstruction of multiple multi-bit dwarf trees. The reconstruction brings the need to update multiple RAM contents. The bubble method can be used to perform the insertion one by one, and follow the update principle of inserting first and then deleting, to ensure that the system still has real-time search capabilities while updating.
[0037] The above implementation methods of insertion are only examples and can be set according to actual conditions. Other implementation methods of insertion not illustrated are also within the protection scope of this application and will not be described in detail here.
[0038] In this embodiment, by inserting or deleting the multi-bit tree, the message classification rules stored in the multi-bit tree are updated, ensuring that the changes in the message classification rules can be dynamically supported. When the multi-bit tree is applied to the hardware structure, it is not necessary to pre-extract all the message classification rules and perform a large amount of preprocessing, so that by inserting or deleting the multi-bit tree, the message classification rules can be dynamically updated, avoiding the tedious preprocessing and improving the processing efficiency.
[0039] In one embodiment, Figure 3 is a flow chart of a method for inserting the message classification rules to be updated into a multi-bit dwarf tree. Figure 3 As shown, inserting the message classification rules to be updated into the multi-bit dwarf tree in step 130 includes steps 131 to 133.
[0040] Step 131, traverse the first multi-bit tree set, and determine whether a multi-bit tree in the first multi-bit tree set has space to store the message classification rule to be updated.
[0041] It should be noted that if it is determined that the multi-bit tree in the first multi-bit tree set has space to store the message classification rule to be updated, step 132 is executed, otherwise step 133 is executed.
[0042] Step 132, construct a new multi-bit dwarf tree according to the message classification rule to be updated.
[0043] Specifically, a bit decision method can be used to process the message classification rules to be updated and determine the storage address information of the tree nodes at each layer of the new multi-bit dwarf tree; finally, the message classification rules to be updated are filled into the new multi-bit dwarf tree to complete the construction of a new multi-bit dwarf tree.
[0044] Step 133, rebuilding the message classification rule to be updated and the multi-bit shrubs in the first multi-bit shrub set according to the filling degree of the multi-bit shrubs and a preset filling degree threshold.
[0045] It should be noted that the reconstruction may be a partial reconstruction of the multi-bit shrub in the first multi-bit shrub set, i.e., inserting the message classification rules to be updated into the first multi-bit shrub set, or reselecting the decision bits of the message classification rules to be updated and the rules in the multi-bit shrub in the first multi-bit shrub set to establish a new multi-bit shrub set.
[0046] In a specific implementation, step 133 can be implemented in the following manner: a second multi-bit short tree set is selected from the first multi-bit short tree set, wherein the filling degree of the multi-bit short trees in the second multi-bit short tree set is less than a preset filling degree threshold; the message classification rule to be updated and the second multi-bit short tree set are rebuilt to obtain a third multi-bit short tree set; it is determined whether the number of multi-bit short trees in the third multi-bit short tree set is less than or equal to the number of multi-bit short trees in the first multi-bit short tree set; if so, it is determined that the reconstruction is successful, and the multi-bit short trees in the second multi-bit short tree set are locally updated according to the message classification rule to be updated.
[0047] In specific implementation, a linked list method can be used to achieve the reconstruction of the message classification rules to be updated and the multi-bit short trees in the first multi-bit short tree set. For example, the first multi-bit short tree set is set to be a fat tree (Fat Tree, FT) linked list, recorded as FT_list; the message classification rule to be updated is marked as a pointer p_rule; the temporary linked list of the first multi-bit short tree set is currently traversed, marked as a pointer p_FT.
[0048] First, traverse the FT_list and try to insert p_rule into the current p_FT. If the insertion is successful, return to insert the message classification rule to be updated into the first multi-bit dwarf tree successfully; otherwise, continue to traverse the FT_list. If p_rule cannot be inserted after traversing all FT_lists, the reconstruction of the message classification rule to be updated and the multi-bit dwarf tree in the first multi-bit dwarf tree set is triggered.
[0049] The specific reconstruction process includes: comparing whether the filling degree of each multi-bit short tree in the first multi-bit short tree set FT_list is less than a preset filling degree threshold, screening the multi-bit short trees whose filling degree is less than the preset filling degree threshold, and obtaining a second multi-bit short tree set; then putting the second multi-bit short tree set and the message classification rule to be updated (i.e., p_rule) into a temporary rule set, and marking the temporary rule set as Rtemp. Then, the multi-bit short trees in the temporary rule set Rtemp are reconstructed to generate a third multi-bit short tree set.
[0050] Determine whether the number of multi-bit trees in the third multi-bit tree set is less than or equal to the number of multi-bit trees in the first multi-bit tree set FT_list; if so, determine that the reconstruction is successful, and replace all the multi-bit trees in the second multi-bit tree set with the multi-bit trees in the temporary rule set (i.e., Rtemp); otherwise, keep the first multi-bit tree set FT_list unchanged, apply for a new FT resource, and insert the message classification rule p_rule to be updated, that is, construct a new multi-bit tree, which is a decision tree generated based on the message classification rule to be updated.
[0051] In this embodiment, the message classification rules to be updated and the multi-bit trees in the first multi-bit tree set are reconstructed according to the filling degree of the multi-bit tree and the preset filling degree threshold, so that the multi-bit trees in the first multi-bit tree set can be updated in time, dynamically support the changes in the message classification rules, avoid the tedious preprocessing, and improve the processing efficiency.
[0052] In one embodiment, the deletion of the message classification rules to be deleted stored in the multi-bit dwarf tree in step 140 can be implemented in the following manner: based on the message classification rules to be deleted, a matching search is performed on each branch node of the multi-bit dwarf tree in the first multi-bit dwarf tree set to determine the branch node to be deleted; and the message classification rules to be deleted stored in the branch node to be deleted are deleted.
[0053] Specifically, according to the search key value of the message classification rule to be deleted, each branch node of the multi-bit dwarf tree in the first multi-bit dwarf tree set can be searched to obtain the storage address information corresponding to the branch node to be deleted, and then according to the storage address information corresponding to the branch node to be deleted, the message classification rule to be deleted can be obtained and deleted.
[0054] In this embodiment, by deleting the message classification rules to be deleted stored in the branch nodes to be deleted, the useless message classification rules in the first multi-bit dwarf tree set can be deleted. It is worth noting that only the message classification rules are deleted, and the deleted branch nodes are still retained, so that when the message classification rules are updated next time, the storage address corresponding to the branch nodes to be deleted can be used to save other message classification rules.
[0055] In one embodiment, before step 120, the method may further include: modifying attribute information of branch nodes in the multi-bit dwarf tree, wherein the attribute information includes priority levels corresponding to rules in the first message classification rule set.
[0056] The attribute information may also include the size of the storage space occupied by the branch node, that is, whether the current branch node is saturated, etc. The above attribute information is only an example, which can be set according to actual needs. Other unspecified attribute information is also within the scope of protection of this application and will not be repeated here.
[0057] In this embodiment, by modifying the attribute information of the branch node, each branch node can timely update the priority level of the message classification rules stored therein, so that the message classification rules can be updated in a timely manner. When classifying the messages to be classified, the message classification rule with the highest priority can be screened out to ensure the accuracy of the classification of the messages to be classified.
[0058] Figure 4 is a flow chart of a message classification method in another embodiment of the present application, such as Figure 4 As shown, the specific steps include the following steps.
[0059] Step 410, using a bit decision method, processes the rules in the input first message classification rule set to determine the storage address information of each layer of nodes in the multi-bit dwarf tree.
[0060] The bit decision method is to observe the distribution of each bit of the rules in the first message classification rule set, screen and obtain bits that match each message classification rule, and then use the matching bits to determine the storage address information of each layer of nodes in the multi-bit dwarf tree, so that the multi-bit dwarf tree can match the rules in the first message classification rule set, so as to facilitate the subsequent search of the multi-bit dwarf tree.
[0061] Step 420, fill the rules in the first message classification rule set into the multi-bit dwarf tree according to the storage address information of each layer of nodes in the multi-bit dwarf tree.
[0062] In some specific implementations, step 420 is implemented in the following manner: grouping the rules in the first message classification rule set to generate N groups, where N is an integer greater than or equal to 1; determining whether the number of rules in each group is greater than a preset rule number threshold, and obtaining a first determination result; extracting groups for which the first determination result is yes, and generating a fill set; and filling the M message classification rules in each group in the fill set into the multi-bit dwarf tree based on the storage address information of each layer of nodes in the multi-bit dwarf tree, where M is less than or equal to the preset rule number threshold.
[0063] It should be noted that, since the rules in the first message classification rule set are not evenly distributed when grouping, the number of message classification rules included in each group can be 1, 2, 5, etc. When the message classification rules in each group are advanced, it is necessary to first determine whether the number of rules in each group is greater than the rule preset number threshold, where the rule preset number threshold is an integer greater than or equal to 1.
[0064] For example, when the threshold value of the preset number of rules is 3, the rules in the first message classification rule set are divided into 3 groups: the first group includes 4 message classification rules, the second group includes 5 message classification rules, and the third group includes 6 message classification rules. A filling set is generated by extracting 1 message classification rule in the first group, 2 message classification rules in the second group, and 3 message classification rules in the third group respectively; so that each constructed multi-bit dwarf tree can store as many rules as possible, and the storage resources are used to the greatest extent. In particular, if the message classification rules in each group are relatively small, the rules in the first message classification rule set can be filled into the multi-bit dwarf tree at one time, and there is no need to construct a new multi-bit dwarf tree.
[0065] In some specific implementations, after the step of filling the M message classification rules in each group in the filling set into the multi-bit dwarf tree according to the storage address information of each layer of nodes in the multi-bit dwarf tree, it also includes: recycling the remaining message classification rules other than the M message classification rules in each group; wherein the remaining message classification rules are used to fill the next multi-bit dwarf tree.
[0066] For example, when the threshold value of the preset number of rules is 3, the rules in the first message classification rule set are divided into 3 groups: the first group includes 4 message classification rules, the second group includes 5 message classification rules, and the third group includes 6 message classification rules. After the first multi-bit dwarf tree is generated by extracting the message classification rules in each group, it is necessary to recycle the remaining message classification rules other than the M message classification rules in each group, that is, the remaining message classification rules include: 1 message classification rule in the first group, 3 message classification rules in the second group, and 3 message classification rules in the third group. Then, the above method is used again to screen and obtain the filling set, and the rules in the filling set are filled into the next multi-bit dwarf tree.
[0067] Step 430: search the first multi-bit dwarf tree set according to the search key value of the message to be classified, and obtain a second message classification rule set that matches the message to be classified.
[0068] Step 440: classify the message to be classified according to the rules in the second message classification rule set, and determine the message type of the message to be classified.
[0069] It should be noted that steps 430 to 440 in this embodiment are consistent with steps 110 to 120 in the previous embodiment, and will not be described in detail here.
[0070] In this embodiment, the rules in the input first message classification rule set are processed by using a bit decision method to determine the storage address information of each layer of nodes in the multi-bit dwarf tree, so that a multi-bit dwarf tree matching the rules in the first message classification rule set can be obtained, and then the rules in the first message classification rule set are filled into the multi-bit dwarf tree, so that the desired message classification rules can be obtained by searching the multi-bit dwarf tree later. The shorter tree shape of the multi-bit dwarf tree is used to make the hardware structure using the message classification method have a shorter pipeline stage, thereby reducing the search delay.
[0071] In one embodiment, step 410 is implemented in the following manner: using a bit decision method, processing the rules in the first message classification rule set, determining K decision bits of the root node of the multi-bit dwarf tree, K is an integer greater than or equal to 1; determining the storage address information of the first layer tree node of the multi-bit dwarf tree based on the K decision bits of the root node; dividing the rules in the first message classification rule set into two k A first-level message classification rule subset is prepared; the first-level message classification rule subset is used as a new first message classification rule set, and the rules in the first-level message classification rule subset are processed in a bit decision manner to determine K decision bits of the first-level tree nodes of the multi-bit dwarf tree; based on the storage address information of the first-level tree nodes and the K decision bits of the first-level tree nodes, the storage address information of the second-level tree nodes is determined.
[0072] It should be noted that each message classification rule in the first message classification rule set includes A fields, where A is an integer greater than or equal to 1. The A fields are A fields extracted from the complete data message, and the A fields include B bits in total. Since 1 bit can represent 2 different message classification rules (for example, 00 and 01), B is an integer greater than or equal to 1.
[0073] For example, Table 1 is a rule format table of message classification rules. Among them, A is equal to 3, B is equal to 12, that is, the message classification rule includes 3 fields (i.e., field 1, field 2, and field 3), each field has 4 bits (for example, field 1 in rule A is represented by 1100, field 2 in rule A is represented by 0000, field 3 in rule A is represented by 10**, etc.), and these 3 fields have a total of 12 bits. Among them, * indicates that the bit can be any value between 0 and 1.
[0074] Table 1 Format of message classification rules
[0075]
[0076]
[0077] By observing Table 1, at least one feasible combination of decision bits can be obtained. For example, the second bit of field 1 and the first bit of field 2 are used as decision bits. These two bits can correspond to four values: {00, 01, 10, 11}, and different values correspond to different message classification rules, that is, 00 corresponds to message classification rule B, 01 corresponds to message classification rule C, 10 corresponds to message classification rule A, and 11 corresponds to message classification rule D.
[0078] Then, based on these two decision bits, the storage address information of the tree node in the multi-bit dwarf tree is determined. If the storage address information is represented by 10 bits, the remaining bits of the storage address information except the decision bit are processed by filling 0. That is, the storage addresses corresponding to the message classification rules {A, B, C, D} are {address A, address B, address C, address D} in sequence. Among them, address A = 10 000 00000, address B = 00 000 00000, address C = 01 000 00000, address D = 11 000 00000.
[0079] Furthermore, when the multi-bit tree is implemented in a hardware structure, each layer of tree nodes of the multi-bit tree corresponds to a first-level pipeline in the hardware structure. Since the current multi-bit tree only stores four rules, A, B, C, and D, the multi-bit tree only includes a root node and two layers of tree nodes. Therefore, when the multi-bit tree is applied to the hardware structure, the number of pipeline levels in the corresponding hardware structure is 3. When searching the multi-bit tree according to the search key value of the message to be classified, it is necessary to perform two bit decisions and one parallel comparison of multiple message classification rules in the leaf node.
[0080] For example, the search key values of the message to be classified include two: key1 = 0011 1110 1100 and key2 = 00111011 0000. When searching the multi-bit dwarf tree with key1 and key2 as indexes, it is necessary to find address C (i.e. 01000 00000) to obtain the message classification rule C stored at address C. Then, the message classification rule C is accurately compared with key1 and key2 at the bit level to obtain the second message classification rule set, that is, key1 corresponds to message classification rule C, and key2 has no corresponding matching message classification rule. Finally, the message to be classified is classified using message classification rule C to obtain the message type of the message to be classified.
[0081] The multi-bit dwarf tree generated in the above manner has fewer layers of tree nodes, so that the hardware structure corresponding to the multi-bit dwarf tree has a shorter pipeline stage, thereby reducing the search delay and improving the processing speed when searching for message classification rules in the multi-bit dwarf tree.
[0082] In a specific implementation, a bit decision method is adopted to process the rules in the first message classification rule set to determine K decision bits of the root node of the multi-bit dwarf tree, including: respectively calculating the bit discrimination between each bit of the rules in the first message classification rule set to obtain a bit discrimination set; determining K decision bits based on the position of each bit, the bit discrimination threshold and the bit discrimination of each bit in the bit discrimination set.
[0083] It should be noted that the bit discrimination can be calculated using a cost function, which is a function that maps the value of a random event or its related random variables into a non-negative real number to represent the "risk" or "loss" of the random event.
[0084] Specifically, it is necessary to calculate the bit discrimination of each bit of the rule in the first message classification rule set, and then compare the bit discrimination of each bit with the bit discrimination threshold (for example, 0.8), sort them, extract the bits with bit discrimination greater than 0.8, and then select K decision bits from the set of bits with bit discrimination greater than 0.8 according to the position of each bit. And, each time the screening is performed, the K bits with the highest discrimination are arbitrarily selected from the A domains of the rules in the current message classification rule set.
[0085] In some scenarios, since different message classification rules care about different bit positions and bit numbers, when the number of bits that different message classification rules care about is insufficient to effectively group the rules in the first message classification rule set, K bits that are of concern to most rules in the first message classification rule set and have the highest bit discrimination can be selected as decision bits. For the remaining message classification rules in the first message classification rule set except the message classification rules corresponding to the decision bits, the remaining message classification rules can be saved in multiple groups at the same time by rule replication, or the remaining message classification rules can be eliminated and wait for the construction of the next multi-bit dwarf tree.
[0086] In this embodiment, the rules in the first message classification rule set are processed by bit decision, K decision bits of the root node of the multi-bit dwarf tree are determined, and the storage address information of the first layer tree node of the multi-bit dwarf tree is determined based on the K decision bits; then the rules in the first message classification rule set are divided into 2 kThe first-level message classification rule subset is used as a new first message classification rule set, and the bit decision method is continued to be used for processing, so that the storage address information of the tree nodes of each layer of the multi-bit dwarf tree can be obtained to complete the construction of the multi-bit dwarf tree. The multi-bit dwarf tree has fewer layers of tree nodes, and the corresponding hardware structure has a shorter pipeline stage, so that when searching for message classification rules in the multi-bit dwarf tree, the search delay can be reduced and the processing speed can be improved.
[0087] Figure 5 The following is a schematic diagram of the structure of a message classification device provided in an embodiment of the present application. The specific implementation of the device can refer to the relevant description of embodiment 1, and the repeated parts will not be repeated. It is worth noting that the specific implementation of the device in this embodiment is not limited to the above embodiment, and other undescribed embodiments are also within the protection scope of the device.
[0088] like Figure 5 As shown, the message classification device specifically includes: a search module 510 is used to search a first multi-bit dwarf tree set according to a search key value of the message to be classified, and obtain a second message classification rule set that matches the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, and the multi-bit dwarf tree is a decision tree constructed according to the rules in the first message classification rule set; a classification module 520 is used to classify the message to be classified according to the rules in the second message classification rule set, and determine the message type of the message to be classified.
[0089] In one specific implementation, the search module 510 is implemented in the form of a multi-bit dwarf tree. Figure 6 As shown, the search module 510 includes: a root node register 511, a first-level tree node address generator 512, a first-level tree node random access memory 513, a leaf node address generator 514 and a leaf node random access memory 515.
[0090] The root node register 511 is used to store the decision content of the root node of the multi-bit dwarf tree (for example, the position information of 5 decision bits, etc.).
[0091] The first-layer tree node address generator 512 is used to detect the search key value of the input message to be classified, and compare the search key value with the decision content stored in the root node register 511 to obtain a comparison result, and output the comparison result to the first-layer tree node random access memory 513.
[0092] The first-layer tree node random access memory 513 is used to store the decision content of the first-layer tree nodes of the multi-bit dwarf tree.
[0093] The leaf node address generator 514 is used to obtain the 2-bit address information input by the first-layer tree node address generator 512 (or, obtain the decision content of the first-layer tree node from the first-layer tree node random access memory 513, that is, 5 bits of address information), and use the 5-bit address information to accurately address each leaf node; at the same time, obtain the remaining bits of the position information of each leaf node, and determine the position information of the leaf node that matches the search key value based on the remaining bits and the 5-bit address information, and then search the leaf node random access memory 515 based on the position information.
[0094] The leaf node random access memory 515 is used to retrieve all message classification rules corresponding to the search key value of the message to be classified, generate a message classification rule set corresponding to the search key value, i.e., a second message classification rule set, and then output the rules in the second message classification rule set to the classification module 520.
[0095] In this embodiment, the message to be classified is acquired through the acquisition module, and the search module is used to search the first multi-bit dwarf tree set according to the search key value of the message to be classified to obtain the second message classification rule set that matches the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, and the multi-bit dwarf tree is a decision tree constructed according to the first message classification rule set. By utilizing the shorter tree shape of the multi-bit dwarf tree, the hardware structure using the message classification method has a shorter pipeline stage, thereby reducing the search delay. Then, the classification module is used to classify the message to be classified according to the rules in the second message classification rule set to determine the message type of the message to be classified, thereby improving the classification correctness and accuracy of the message to be classified. When targeting different message classification rules, there is no need to change the hardware resource settings, and it has good scalability.
[0096] Figure 7 This is a block diagram of a message classification device provided in another embodiment of the present application. Figure 7 As shown, the message classification device specifically includes: a key value generation module 710, multiple search modules (for example, search module 720-1, search module 720-2, search module 720-3, ..., search module 720-X, etc., where X is an integer greater than or equal to 1), multiple parallel comparison modules (for example, parallel comparison module 730-1, parallel comparison module 730-2, parallel comparison module 730-3, ..., parallel comparison module 730-X, etc.), and a priority arbitration module 740.
[0097] It should be noted that each search module searches for its corresponding multi-bit shrub, multiple multi-bit shrubs constitute a first multi-bit shrub set, and the multi-bit shrub is a decision tree constructed according to the rules in the first message classification rule set.
[0098] Among them, the key value generation module 710 is used to process the header information of the message to be classified and generate a search key value, which represents the core bit of the key field of the message to be classified. The key value generation module 710 distributes the generated search key value to each search module, for example, to the search module 720-1, the search module 720-2, the search module 720-3, ..., the search module 720-X, etc.
[0099] Each search module is used to search the multi-bit tree step by step according to the search key value, and each search module can output the message classification rules stored in a multi-bit tree, which may include 0, 1 or at most M message classification rules, where M is an integer greater than or equal to 1. Each search module outputs the at most M message classification rules and the search key value it outputs to the corresponding parallel comparison module. For example, the search module 720-1 outputs the at most M message classification rules and the search key value it outputs to the parallel comparison module 730-1, and the search module 720-2 outputs the at most M message classification rules and the search key value it outputs to the parallel comparison module 730-2, and so on.
[0100] Each parallel comparison module is used to perform a bit-level precise comparison between the obtained at most M message classification rules and the search key value, generate a second message classification rule set, and then output the message classification rules in the second message classification rule set to the priority arbitration module 740 .
[0101] The priority arbitration module 740 is used to perform priority arbitration on each message classification rule in the second message classification rule set, for example, to compare the priority levels of each message classification rule and sort them according to the priority levels to obtain the message classification rule with the highest priority, and use the message classification rule with the highest priority to classify the messages to be classified to obtain the message type of the messages to be classified.
[0102] In this embodiment, multiple search modules are used to search the multi-bit dwarf tree to obtain multiple groups of message classification rules corresponding to the search key value; then, the search key value and its corresponding message classification rule are output to the corresponding parallel comparison module, so that multiple parallel comparison modules can accurately compare the message classification rules at the same time, improve processing efficiency, and quickly generate a second message classification rule set; finally, a priority arbitration module is used to perform priority arbitration on the rules in the second message classification rule set to obtain the message classification rule with the highest priority, and then use the message classification rule with the highest priority to classify the messages to be classified, so as to ensure the accuracy of message classification.
[0103] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, or a part of a physical unit, or can be implemented in a combination of multiple physical units. In addition, the present invention is not limited to the specific configurations and processes described in the above embodiments and shown in the figures. For the convenience and brevity of description, a detailed description of the known methods is omitted here. And the specific working process of the above-described systems, modules and units can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0104] Figure 8 This is a structural diagram of an exemplary hardware architecture of an electronic device according to the message classification method and device in the embodiments of the present application.
[0105] like Figure 8 As shown, the electronic device 800 includes an input device 801, an input interface 802, a central processing unit 803, a memory 804, an output interface 805, and an output device 806. The input interface 802, the central processing unit 803, the memory 804, and the output interface 805 are connected to each other via a bus 807, and the input device 801 and the output device 806 are connected to the bus 807 via the input interface 802 and the output interface 805, respectively, and then connected to other components of the electronic device 800.
[0106] Specifically, the input device 801 receives input information from the outside and transmits the input information to the central processing unit 803 through the input interface 802; the central processing unit 803 processes the input information based on the computer executable instructions stored in the memory 804 to generate output information, temporarily or permanently stores the output information in the memory 804, and then transmits the output information to the output device 806 through the output interface 805; the output device 806 outputs the output information to the outside of the computing device 800 for user use.
[0107] In one embodiment, Figure 8 The electronic device 800 shown can be implemented as a network device, which may include: a memory configured to store a program; a processor configured to run the program stored in the memory to execute any one of the message classification methods described in the above embodiments.
[0108] According to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network, and / or installed from a removable storage medium.
[0109] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0110] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for limiting purposes. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly noted, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present application as set forth in the appended claims.
Claims
1. A message classification method, wherein include: Searching a first multi-bit dwarf tree set according to a search key value of the message to be classified, and obtaining a second message classification rule set matching the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, and the multi-bit dwarf tree is a decision tree constructed according to the rules in the first message classification rule set; Classifying the message to be classified according to the rules in the second message classification rule set to determine the message type of the message to be classified; Before the step of searching the first multi-bit dwarf tree set according to the search key value of the message to be classified to obtain a second message classification rule set matching the message to be classified, the step further includes: Using a bit decision method, processing the rules in the first message classification rule set inputted, and determining the storage address information of each layer node in the multi-bit dwarf tree; According to the storage address information of each layer node in the multi-bit dwarf tree, the rules in the first message classification rule set are filled into the multi-bit dwarf tree.
2. The method according to claim 1, in, The step of searching the first multi-bit dwarf tree set according to the search key value of the message to be classified to obtain a second message classification rule set matching the message to be classified includes: Searching the multi-bit short trees in the first multi-bit short tree set in parallel according to the search key value to obtain a primary message classification rule set; The search key value is compared with the rules in the primary message classification rule set to screen and obtain the second message classification rule set, wherein the rules in the second message classification rule set are rules that match the message to be classified.
3. The method according to claim 1, in, The classifying the to-be-classified message according to the rule in the second message classification rule set to determine the message type of the to-be-classified message includes: Obtaining the priority level corresponding to the rule in the second message classification rule set; According to the priority level, performing priority arbitration on the rules in the second message classification rule set to generate an optimal message classification rule; The to-be-classified messages are classified according to the optimal message classification rule to determine the message type of the to-be-classified messages.
4. The method according to any one of claims 1 to 3, in, Before the step of classifying the to-be-classified message according to the rule in the second message classification rule set and determining the message type of the to-be-classified message, the method further includes: Insert the message classification rule to be updated into the multi-bit dwarf tree, or delete the message classification rule to be deleted stored in the multi-bit dwarf tree.
5. The method according to claim 4, in, The inserting the message classification rule to be updated into the multi-bit dwarf tree comprises: Traversing the first multi-bit short tree set, and determining whether a multi-bit short tree in the first multi-bit short tree set has space to store the message classification rule to be updated; If yes, construct a new multi-bit dwarf tree according to the message classification rule to be updated; Otherwise, the message classification rule to be updated and the multi-bit shrubs in the first multi-bit shrub set are rebuilt according to the filling degree of the multi-bit shrub and a preset filling degree threshold.
6. The method according to claim 5, in, The reconstructing the message classification rule to be updated and the multi-bit shrubs in the first multi-bit shrub set according to the filling degree of the multi-bit shrubs and a preset filling degree threshold comprises: Selecting a second multi-bit short tree set from the first multi-bit short tree set, wherein the filling degree of the multi-bit short trees in the second multi-bit short tree set is less than the preset filling degree threshold; Reconstructing the message classification rule to be updated and the second multi-bit short tree set to obtain a third multi-bit short tree set; Determining whether the number of multi-bit shrubs in the third multi-bit shrub set is less than or equal to the number of multi-bit shrubs in the first multi-bit shrub set; If so, it is determined that the reconstruction is successful, and the multi-bit dwarf trees in the second multi-bit dwarf tree set are partially updated according to the message classification rule to be updated.
7. The method according to claim 4, in, The deleting the message classification rule to be deleted stored in the multi-bit dwarf tree comprises: According to the message classification rule to be deleted, matching and searching each branch node of the multi-bit dwarf tree in the first multi-bit dwarf tree set to determine the branch node to be deleted; The message classification rule to be deleted stored in the branch node to be deleted is deleted.
8. The method according to any one of claims 1 to 3, in, Before the step of classifying the to-be-classified message according to the rule in the second message classification rule set and determining the message type of the to-be-classified message, the method further includes: Modify the attribute information of the branch nodes in the multi-bit dwarf tree, wherein the attribute information includes the priority level corresponding to the rules in the first message classification rule set.
9. The method according to claim 1, in, The bit decision method is used to process the rules in the first message classification rule set input to determine the storage address information of each layer node of the multi-bit dwarf tree, including: Using the bit decision method, the rules in the first message classification rule set are processed to determine K decision bits of the root node of the multi-bit dwarf tree, where K is an integer greater than or equal to 1; Determine storage address information of a first-layer tree node of the multi-bit dwarf tree according to the K decision bits of the root node; Divide the rules in the first message classification rule set into 2 k A subset of first-level message classification rules; The first-level message classification rule subset is used as the new first message classification rule set, and the rules in the first-level message classification rule subset are processed by continuing to adopt the bit decision method to determine K decision bits of the first-level tree nodes of the multi-bit dwarf tree; Determine the storage address information of the second layer tree node according to the storage address information of the first layer tree node and the K decision bits of the first layer tree node; At this point, the storage address information of the tree nodes at each layer of the multi-bit dwarf tree is obtained.
10. The method according to claim 9, in, The method of using the bit decision method to process the rules in the first message classification rule set to determine K decision bits of the root node of the multi-bit dwarf tree includes: Respectively calculating the bit discrimination between each bit of the rule in the first message classification rule set to obtain a bit discrimination set; The K decision bits are determined according to the positions of the respective bits, a bit discrimination threshold and the respective bit discriminations in the bit discrimination set.
11. The method according to claim 1, in, The step of filling the rules in the first message classification rule set into the multi-bit dwarf tree according to the storage address information of each layer node in the multi-bit dwarf tree comprises: Grouping the rules in the first message classification rule set to generate N groups, where N is an integer greater than or equal to 1; Determine whether the number of rules in each of the groups is greater than a preset rule number threshold, and obtain a first determination result; Extract the group for which the first judgment result is yes, and generate a filling set; According to the storage address information of each layer node in the multi-bit dwarf tree, M message classification rules in each group in the filling set are filled into the multi-bit dwarf tree, wherein M is less than or equal to the preset number threshold of the rules.
12. The method according to claim 11, in, After the step of filling the M message classification rules in each group in the filling set into the multi-bit dwarf tree according to the storage address information of each layer node in the multi-bit dwarf tree, the method further includes: Recycling the remaining message classification rules other than the M message classification rules in each group; The remaining message classification rules are used to fill the next multi-bit dwarf tree.
13. A message classification device, include: A search module, configured to search a first multi-bit dwarf tree set according to a search key value of a message to be classified, and obtain a second message classification rule set matching the message to be classified, wherein the first multi-bit dwarf tree set includes a multi-bit dwarf tree, and the multi-bit dwarf tree is a decision tree constructed according to the rules in the first message classification rule set; A classification module, used to classify the message to be classified according to the rules in the second message classification rule set, and determine the message type of the message to be classified; The message classification device further includes: A filling module is used to process the rules in the first message classification rule set input by bit decision-making, determine the storage address information of each layer of nodes in the multi-bit dwarf tree; and fill the rules in the first message classification rule set into the multi-bit dwarf tree according to the storage address information of each layer of nodes in the multi-bit dwarf tree.
14. An electronic device, wherein include: one or more processors; A storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 12.
15. A computer readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Multi-dimensional packet classification
US20170244642A1