Reinforcement learning data packet classification method based on Saras algorithm
By applying reinforcement learning method based on Saras algorithm in packet classification, the rule set is trained to achieve reasonable division of tuple space, and the problem of traversal search and overlapping rules in the tuple space search scheme is solved, which significantly improves the speed and performance of packet classification.
Patent Information
- Application Number
- CN202311666961.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
The existing tuple space search scheme has traversal search characteristics in packet classification, which affects the classification search performance. Merging tuples will cause overlapping rules and hash collisions, reducing the classification search performance.
The reinforcement learning packet classification method based on Saras algorithm is adopted, and the rule set is trained, the environment state, action space and return functions are designed, and the rule prefix length range output by the training model is used to metacombine and merge, so as to achieve reasonable division of tuple space.
The time for traversing packets in packet classification is reduced, the packet classification rate is greatly improved, the constant-level time complexity of rule updates is realized, and the performance of packet classification is improved.
Smart Images

Figure CN120123807A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data packet classification scheme design, and particularly aims at the problem of reasonably constructing a tuple space for fast packet classification in tuple space - based data packet classification. Background Art
[0002] Data packet classification technology refers to establishing a data packet classifier using a given rule set. The data packet searches through the classifier to find the matching rule and execute the action corresponding to the rule. With the continuous development of the Internet, the amount of data that needs to be processed per unit time has increased sharply, and many network services rely on data packet classification. This situation has brought great challenges to the design of high - performance data packet classifiers and has attracted extensive research attention from many institutions at home and abroad.
[0003] The design of high - performance data packet classification methods needs to adhere to the following two principles. First, the designed classifier should support fast rule set updates. Second, the designed scheme should achieve fast data packet classification to cope with the increasing amount of data.
[0004] Current research has proposed many data packet classification methods. They mainly include two categories: decision tree search and tuple space search (TSS). The decision - tree - based scheme can achieve high classification speed, but the problem of rule duplication increases memory occupancy and hinders support for fast updates. The tuple - space - search - based scheme partitions the rule set, and each subspace is called a tuple and stored using a hash table. This method supports fast rule updates, but its traversal search characteristic affects the classification lookup performance.
[0005] In view of the characteristics of tuple space search, currently, some tuple - space - based schemes propose to improve TSS by merging tuples with similar characteristics to reduce the time of traversing tuples. However, merging tuples will bring rule overlap and more hash collisions, which will lead to a decline in the classification lookup performance of large - scale rule sets. In short, the main research results for improving TSS are to trade - off between the number of tuples and rule overlap, that is, to design the most reasonable tuple partitioning method. Therefore, there is an urgent need to propose a new classification scheme that retains the fast - update advantage of TSS while reasonably partitioning the tuple space as much as possible to improve the data packet classification search speed to adapt to the current situation of Internet development. Summary of the Invention
[0006] In view of the above prior art, the present invention designs a reinforcement learning packet classification method based on the Saras algorithm. By training the rule set, the present invention obtains a meta-combination method that balances the number of rule groups and hash collisions within the group. This method can reduce the time for traversing groups in packet classification and greatly improve the packet classification rate; the present invention uses a hash function to ensure a constant time complexity for rule update and realizes the rapid update of classification rules.
[0007] Meanwhile, the specific steps are as follows:
[0008] Step 1: Divide the rules to be processed into four rule subsets according to the IP field. Each subset is stored in a hash table in the tuple space manner. Set the length of the tuple IP field as the environmental state, and changing the length of the tuple IP field as the action. Design a reward function according to the sum of the number of rule overlaps and hash collisions in the four tuples obtained after taking actions under different states.
[0009] Step 2: Use the Saras algorithm to build a tuple space partitioning model with the rules to be processed as the input and the final tuple partitioning state as the output.
[0010] Step 3: According to the tuple partitioning method obtained in Step 2, divide the rules into four rule subsets. Determine whether to perform iterative partitioning on the subset according to the number of rules in each subset; if partitioning is required, transfer this subset as the new rules to be processed to Step 1, otherwise, transfer to Step 4.
[0011] Step 4: Insert the rules into different tuples according to the tuple partitioning method obtained in Step 2 as the rule classifier.
[0012] Step 5: Input the data packet to be classified into the rule classifier. Take the IP value of the data packet corresponding to the tuple according to the length of the IP field of each tuple for hashing, search for matching rules in the tuple, and after traversing all tuples, use the matching rule with the highest priority as the output and execute the actions corresponding to the rules.
[0013] Compared with the prior art, the beneficial effects of the present invention are:
[0014] A reinforcement learning packet classification method based on the Saras algorithm proposed by the present invention retains the advantages of the tuple class classification scheme in supporting fast update of the rule set. At the same time, the present invention makes full use of the characteristics of reinforcement learning, which studies how to maximize the obtained reward in a complex environment. Based on the Saras algorithm, the rule set is trained, the environmental state, action space and reward function are designed, and the rule prefix length range output by the training model is used for tuple merging, realizing a reasonable partition of the tuple space under different rule sets, that is, a partition method that balances the number of tuples and the internal rule overlap and hash collision of tuples. Experimental results show that under the 1k and 10k rule sets, compared with TSS, the present invention achieves a packet classification speed 5.1 times and 2.4 times that of TSS. It effectively solves the problem of poor packet classification performance of the tuple space search scheme and has broad application prospects in network services relying on packet classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the overall flowchart of the algorithm of the reinforcement learning packet classification method based on the Saras algorithm of the present invention;
[0016] Figure 2 is the flowchart of building the tuple space partition model of the present invention;
[0017] Figure 3 is the architecture diagram of the rule classifier constructed based on the learning-based tuple space partition method of the present invention.
[0018] Figure 4 is the rule set used in the example of the present invention.
[0019] ( Figure 1 is the attached drawing of the abstract) DETAILED DESCRIPTION OF THE INVENTION
[0020] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments, but the following embodiments are by no means any limitation to the present invention.
[0021] A reinforcement learning packet classification method based on the Saras algorithm proposed by the present invention includes a tuple space partition method and a rule classifier construction method. Based on the Saras algorithm, the rule set is trained, the environmental state, action space and reward function are designed, and the rule prefix length range output by the training model is used for tuple merging, and a tuple space partition method that balances the number of tuples and the internal rule overlap and hash collision of tuples is proposed, and then a more reasonable rule classifier is established.
[0022] As Figure 1 shown, the specific steps of the reinforcement learning packet classification method based on the Saras algorithm proposed by the present invention are as follows:
[0023] Step 1: Divide the rules to be processed into four rule subsets according to the IP field. Each subset is stored in a hash table in the tuple space manner. Set the length of the tuple IP field as the environmental state, and change the length of the tuple IP field as the action. Design a reward function based on the sum of the number of rule overlaps and hash collisions in the four tuples obtained after taking actions in different states.
[0024] Step 2: Use the Saras algorithm to build a tuple space partitioning model with the rules to be processed as the input and the final tuple partitioning state as the output.
[0025] Step 3: According to the tuple partitioning method obtained in Step 2, divide the rules into four rule subsets. Determine whether to perform iterative partitioning on the subset based on the number of rules in each subset; if partitioning is required, transfer this subset as the new rules to be processed to Step 1, otherwise, transfer to Step 4.
[0026] Step 4: Insert the rules into different tuples according to the tuple partitioning method obtained in Step 2 as a rule classifier.
[0027] Step 5: Input the data packet to be classified into the rule classifier, take the IP value of the data packet corresponding to the tuple according to the length of the IP field of each tuple for hashing, find the matching rule in the tuple, and use the matching rule with the highest priority as the output after traversing all tuples, and execute the corresponding rule action.
[0028] The flowchart of the tuple space partitioning learning model based on the Saras algorithm is as Figure 2 shown.
[0029] Initialize Q(S, A) = 0 and the initial state S, set the ε value, and use the ε-greedy strategy to generate actions. The action A taken by the agent at time t interacts with the environment to obtain a new state S' at time t + 1, calculate the reward function R under S', use the policy to select A' from S' again, and update the Q function using the reward function value and the Q value according to the Q update formula. The formulas used are as follows:
[0030] R = -(overlap + collision)
[0031] Q(S, A) ← Q(S, A) + α(R + rQ(S', A') - Q(S, A))
[0032] S ← S'; A ← A'
[0033] where α is the learning rate and r is the discount factor of the next action value function.
[0034] Perform the reinforcement learning process within the set number of iterations. After reaching the maximum number of iterations, output the final state to obtain the optimal tuple partitioning method.
[0035] The overall architecture of the tuple space data packet classification method based on reinforcement learning proposed by the present invention is as shown in Figure 3 Figure. The rule set obtains an optimized tuple space partitioning method through iterative reinforcement learning, and the rules are inserted into the corresponding hash tables to form a rule classifier. The method of inserting a rule into the corresponding tuple is as follows:
[0036] Obtain the lengths (x, y) of the IP fields of the rule. Let the optimal partitioning state finally output be (x 1 , y 1 ), and the minimum prefix length of the rule be (x 0 , y 0 ). When the rule set is the initial rule set, x 0 = y 0 = 0;
[0037] When x < x 1 and y < y 1 , insert the rule into the (x 0 , y 0 ) tuple. Take the first x 0 bits of the source IP and the first y 0 bits of the destination IP as the key value for hashing, and insert it into the hash table of this tuple;
[0038] When x < x 1 and y > y 1 , insert the rule into the (x 0 , y 1 ) tuple. Take the first x 0 bits of the source IP and the first y 1 bits of the destination IP as the key value for hashing, and insert it into the hash table of this tuple;
[0039] When x > x 1 and y < y 1 , insert the rule into the (x 1 , y 0 ) tuple. Take the first x 1 bits of the source IP and the first y 0 bits of the destination IP as the key value for hashing, and insert it into the hash table of this tuple;
[0040] When x > x 1 and y > y 1 , insert the rule into the (x 1 , y 1 ) tuple. Take the first x 1 bits of the source IP and the first y 1 bits of the destination IP as the key value for hashing, and insert it into the hash table of this tuple.
[0041] Example:
[0042] In the present invention, an example of establishing a rule classifier for a rule set to be processed through reinforcement learning is as Figure 3 shown, and the rule set used in the example is as Figure 4 shown. Let the threshold threshold = 3. Input rules 0 - 9 into the tuple space partitioning learning model based on the Saras algorithm, and output the first partitioning result (20, 5). Insert the rules into four tuples according to the IP length. Among them, in Tuple2, after the IP lengths of rules 3 and 5 are reduced to (0, 5), the key values are the same, resulting in rule overlap. Rule overlap inevitably occurs in tuple merging, and the tuple space partitioning learning model designed in the present invention minimizes such situations as much as possible. At this time, the number of rule subsets in Tuple4 is greater than the set threshold. Use the rule set in this tuple as the rule set to be processed and iteratively use the tuple space partitioning learning model to output the second partitioning result. After this partitioning, the number of rules in all tuples is less than the threshold, and the rule classifier is established. After the received data packet is input, all tuples will be traversed for matching search, and finally the matching rule with the highest priority will be output and the actions corresponding to the rule will be executed.
[0043] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A reinforcement learning data packet classification method based on the Saras algorithm, the method comprises: Step 1: Divide the rules to be processed into four rule subsets according to the IP field, and each subset is stored in a hash table in the tuple space manner. Set the length of the tuple IP field as the environmental state, change the length of the tuple IP field as the action, and design a reward function according to the value of the sum of the rule overlap and hash collision times in the four tuples obtained after taking actions under different states. Step 2: Use the Saras algorithm to build a tuple space partitioning model with the rules to be processed as the input and the final tuple partitioning state as the output. Step 3: According to the tuple partitioning method obtained in Step 2, divide the rules into four rule subsets, and judge whether to perform iterative partitioning on the subset according to the number of rules in each subset; if partitioning is required, this subset is used as the new rules to be processed and transferred to Step 1, otherwise, transfer to Step 4. Step 4: Insert the rules into different tuples according to the tuple partitioning method obtained in Step 2 as the rule classifier. Step 5: Input the data packet to be classified into the rule classifier, take the IP value of the data packet corresponding to the tuple according to the length of the IP field of each tuple for hashing, find the matching rule in the tuple, and after traversing all tuples, use the matching rule with the highest priority as the output and execute the actions corresponding to the rule.
2. A reinforcement learning data packet classification method according to claim 1, characterized in that the Step 2 specifically includes: (1) In the tuple IP field length (x, y), let the minimum value of x be lenA 0 , and the minimum value of y be lenB 0 . In the initial partition, since the lengths of the rule IP fields are all 0 - 32 bits, so lenA 0 = lenB 0 = 0. When iteratively partitioning the already partitioned rule subsets, the values of lenA 0 and lenB 0 inherit the minimum IP length values of the subsets. (2) If the random initial state is (lenA 1 , lenB 1 ), then the large rule set is divided into 4 subsets, and the corresponding tuple IP lengths are (lenA 0 , lenB 0 ), (lenA 0 , lenB 1 ), (lenA 1 , lenB 0 ), (lenA 1 , lenB 1 ). The possible actions in this state include (lenA 1 + 1, lenB 1 ), (lenA 1 - 1, lenB 1 ), (lenA 1 , lenB 1 + 1), (lenA 1 , lenB 1 - 1); (3) Use the ε-greedy strategy to generate actions. The action A taken by the agent at time t interacts with the environment to obtain a new state S' at time t + 1, and calculate the reward function under S': R = -(overlap + collision) where overlap is the sum of the rule overlap numbers in each tuple under S', and collision is the total number of hash collisions. The smaller the negative feedback caused by rule overlap and hash collision, the larger the R value. Use the strategy to select A' from S' again, and learn and update the Q function according to the new state and reward value. Initialize Q(S, A) = 0, and the update formula is: Q(S, A) ← Q(S, A) + α(R + rQ(S', A') - Q(S, A)) S ← S'; A ← A' where α is the learning rate and r is the discount factor of the next action value function; (4) Perform the reinforcement learning process within the set number of iterations. After reaching the maximum number of iterations, complete the tuple space partitioning, and output the final state to obtain the optimal tuple partitioning method.
3. A reinforcement learning data packet classification method according to claim 1, characterized in that in the Step 3, judge whether to perform iterative rule partitioning on the rules to be processed according to the number of rules in the rules to be processed set; specifically includes: Count the number of rules in the rules subset to be processed, set the iterative partitioning threshold according to the size of the total rule set. When the number of rules is greater than the threshold, the rule set needs to be partitioned; otherwise, partitioning is not required.
4. A reinforcement learning data packet classification method according to claim 1, characterized in that In step 4, insert the rules into different tuples according to the tuple partitioning method obtained in step 2; specifically including: (1) Obtain the length of the rule IP field (x, y), and set the optimal division state to be output finally as (x 1 , y 1 ), and the minimum IP length of the rule is (x 0 , y 0 ). When the rule set is the initial rule set, x 0 = y 0 = 0; (2) When x < x 1 and y < y 1 , insert the rule (x 0 , y 0 ) tuple. Take the first x 0 bits of the source IP and the first y 0 bits of the destination IP as the key value for hashing, and insert the tuple into the hash table; (3) When x < x 1 and y > y 1 , insert the rule (x 0 , y 1 ) tuple. Take the first x 0 bits of the source IP and the first y 1 bits of the destination IP as the key value for hashing, and insert the tuple into the hash table; (4) When x > x 1 and y < y 1 , insert the rule (x 1 , y 0 ) tuple. Take the first x 1 bits of the source IP and the first y 0 bits of the destination IP as the key value for hashing, and insert the tuple into the hash table; (5) When x > x 1 and y > y 1 , insert the rule (x 1 , y 1 ) tuple. Take the first x 1 bits of the source IP and the first y 1 bits of the destination IP as the key value for hashing, and insert the tuple into the hash table.
5. According to a reinforcement learning data packet classification method described in claim 1, characterized in that, In step 5, input the data packet to be classified into the rule classifier, and take the IP value of the data packet corresponding to the tuple according to the IP field length of each tuple for hashing; specifically including: The data packet to be classified will traverse each tuple, take the first x1 bits of the source IP and the first y1 bits of the destination IP in (x1, y1) as the key value for hashing. If there is a rule at the indexed position, continue to match the source port, destination port, and protocol fields. If all match, this rule is the matching rule for the data packet. Then, traverse the remaining tuples in the same way. If there is a rule with a higher priority, replace the previous matching rule until the traversal is complete, output the final matching rule, and execute the corresponding action.