Multi-service traffic packet scheduling method for P4 switch
By combining the Sketch algorithm and Bloom filter for feature extraction, the decision tree conditions are simplified, and the variable priority packet scheduling method is adopted, the problems of rough feature engineering, insufficient hardware adaptability and insufficient dynamic adjustment of P4 switches in service flow classification and packet scheduling are solved, and efficient traffic processing and service quality assurance are achieved.
Patent Information
- Application Number
- CN202510131815.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-06
AI Technical Summary
When using P4 switches for service flow classification and grouping scheduling, the prior art has problems such as rough feature engineering, insufficient hardware adaptability, insufficient dynamic adjustment and high hash collision rate, making it difficult to achieve efficient traffic processing and service quality assurance in multiple business scenarios.
A multi-service traffic packet scheduling method for P4 switches is proposed. By combining the Sketch algorithm and the feature extraction method of Bloom filter, the conflict detection mechanism is optimized; business flow classification is performed based on the decision tree, and the decision tree conditions are simplified to adapt to the resource limitation of P4 switches; the packet scheduling method AD-PIFO is designed to dynamically adjust the packet forwarding priority.
It improves network packet processing efficiency and service quality, meets the QoS needs of different services, enhances the overall performance and efficiency of the switch, and significantly improves the performance in high-priority traffic throughput, low latency, medium-low priority traffic throughput, stability and practicality.
Smart Images

Figure CN119996310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technology and traffic classification, and in particular to a multi-service traffic grouping scheduling method based on a P4 switch. Background Art
[0002] The Programmable Protocol Independent Processor (P4) technology provides an open language and architecture for programming network packet processors. The emergence of P4 technology has greatly improved the programmability of SDN, making it possible to schedule packets in the network. However, the current research on using P4 for service flow classification and packet scheduling still faces many difficulties, and there are obvious shortcomings compared with related patent technologies.
[0003] Among the existing related technologies, some use traditional decision trees or deep learning models for traffic classification and scheduling. However, these models generally have the problem of insufficient hardware adaptability, and they are not optimized for the resource limitations of P4 switches. P4 switches have limited resources, such as limited memory size and computing resources, and traditional decision tree models may require a large amount of memory when building and storing. Deep learning models have high computational complexity, and running them on P4 switches will seriously affect their performance, resulting in inefficient processing of data packets and difficulty in meeting actual application needs.
[0004] At the same time, the existing technology lacks in terms of dynamics. Many scheduling algorithms lack a dynamic priority adjustment mechanism based on real-time network status. In actual SDN networks, the network status is complex and changeable, and the transmission delay of data packets, link load, etc. may change at any time. However, the existing technology cannot dynamically adjust the forwarding priority of data packets according to these real-time changing network status, and it is difficult to adapt to the dynamic network environment and the QoS requirements of diversified services. For example, when the network is congested, the priority of key business data packets cannot be increased in time, resulting in a decline in service quality.
[0005] In addition, in terms of feature extraction, the efficiency of existing technologies is low. Most methods do not combine Sketch with the time threshold detection mechanism, resulting in a high hash conflict rate. When processing a large number of data packets, the accuracy and efficiency of feature extraction are greatly reduced due to frequent hash conflicts. For example, when using Bloom filters to store features, if the conflict detection mechanism is not optimized, the misjudgment rate will increase, which will affect the accuracy of business flow classification and ultimately reduce the performance of the algorithm in multi-classification scenarios.
[0006] Although P4 technology has brought new possibilities for network traffic scheduling, there are still two major problems in the research on using P4 for service flow classification and packet scheduling: First, the classification algorithm based on P4 is rough in feature engineering, and it is difficult to extract effective features in resource-constrained P4 switches, resulting in poor algorithm performance in multi-classification scenarios; second, the scheduling algorithm based on P4 is insufficient in dynamic adjustability, and cannot adaptively adjust the scheduling strategy according to the actual delay of the data packet in the SDN network, making it difficult to adapt to the dynamic network environment and diversified service QoS requirements.
[0007] In response to the above challenges, the present invention proposes a service packet scheduling method for P4 switches, which aims to improve network packet processing efficiency and quality of service (QoS) through refined feature extraction, service flow classification and dynamic packet scheduling strategy, meet the QoS requirements of different services, and improve the overall performance and efficiency of the switch. This method optimizes the conflict detection mechanism by using a feature extraction method based on statistical intervals and Sketch, designs a packet scheduling method AD-PIFO based on variable priority, and dynamically adjusts the packet forwarding priority to solve the existing problems. The method also verifies its advantages in high priority traffic throughput, low latency, medium and low priority traffic throughput, stability and practicality in experiments. Summary of the invention
[0008] In view of the above problems, the present invention proposes a multi-service traffic group scheduling method for P4 switches. The implementation path is constructed from all aspects from feature extraction, service flow classification and group scheduling. In the feature extraction link, the service flow is locked according to the five-tuple of the data packet header, and the features are stored in the register by the Sketch algorithm. After the quantity is sufficient, the model is input for classification and the label is transmitted, during which the Bloom filter is optimized and key features are selected; when classifying the service flow, the decision tree algorithm is selected according to the P4 characteristics and simplified and optimized; the group scheduling constructs a ranking calculation and group scheduling module architecture, dynamically adjusts the priority according to the delay information, and optimizes the queue management by the AD-PIFO algorithm. The present invention advances from feature capture, classification decision to group scheduling in sequence, effectively improving the performance of the P4 switch.
[0009] To achieve the purpose of the present invention, the technical solution of the present invention is as follows:
[0010] (1) Extract and store features based on the Sketch algorithm, use the five-tuple in the packet header to determine the service flow, and store the packet features in a register through a Bloom filter based on the Sketch algorithm. When the feature data reaches a certain amount, it is formed into a feature vector and input into the pre-trained machine learning model for classification.
[0011] The P4 switch has powerful programmable capabilities and can flexibly process data packets. The feature extraction and storage based on the Sketch algorithm can efficiently complete the extraction and storage of data packet features on the data plane with the help of the programmable characteristics of P4, providing a data basis for subsequent business flow classification, which is in line with the characteristics of the P4 switch that can customize the data packet processing logic according to different needs.
[0012] (2) Based on the decision tree, the service flow classification is performed. The decision tree node range compression method based on bit operation is used to adjust the decision tree node range, reduce the number of entry matches, adapt to the P4 switch data plane requirements, and execute the classification process through the P4 match-action pipeline to achieve service flow classification. Bit operation directly operates on binary bits, and right shift operation plays a key role in this method.
[0013] The data plane of P4 switches usually has limited resources and is not suitable for complex calculations. The decision tree node range compression method based on bit operations can effectively reduce the number of entry matches, reduce the computational complexity, and adapt well to the resource limitations of the data plane of P4 switches. At the same time, P4's match-action pipeline provides a natural implementation path for the execution of the decision tree classification process, which can efficiently complete the business flow classification task.
[0014] (3) Packet scheduling is performed based on variable priorities. When a data packet enters the first P4 switch, a custom header is added to record the priority and accumulated queuing delay. A priority adjustment module is connected after the delay statistics module to modify the packet forwarding priority based on the state classification of the accumulated delay ratio. That is, the data packet is divided into idle state, normal state and congested state according to the ratio of the accumulated queuing delay of the data packet to the ideal queuing delay, and the priority is adjusted based on this state classification. At the same time, the dynamic boundary threshold update rule is followed, and the PIFO algorithm is used to select the queue according to the priority and adjust the sorting.
[0015] P4 supports flexible customization and modification of packet headers, which enables custom headers to be added to record priority and accumulated queuing delay when packets enter the switch. Moreover, the programmability of P4 facilitates the implementation of priority adjustment modules and dynamic boundary threshold update rules. Combined with the PIFO algorithm, it can dynamically adjust the priority and select the queue according to the real-time status of the packet, fully demonstrating the flexibility and programmability of the P4 switch in packet scheduling.
[0016] In actual application scenarios, the order of magnitude and service types involved in multi-service traffic are diverse. The order of magnitude of service flows can reach thousands or even higher. In large-scale network scenarios, such as metropolitan area networks or large-scale data center networks, the order of magnitude of service flows is likely to expand to tens of thousands or even higher to meet the needs of many users and complex services. In terms of service types, it covers a variety of service types with different characteristics. For example, real-time interactive services such as VoIP, data transmission services such as FTP services, and multimedia services, including streaming video and streaming audio, have a large amount of service data. In addition, in actual networks, there may also be traffic generated by IoT devices and various business system traffic within the enterprise.
[0017] As an improvement of the present invention, the specific method of step (1) is as follows:
[0018] (1.1) First, the five-tuple in the packet header (source IP, destination IP, protocol number, source port number, and destination port number) is used to determine the service flow. Then, the five-tuple is hashed by a hash function to determine the service flow and update the traffic features. In this process, the Sketch algorithm is introduced to efficiently calculate the features. The Sketch algorithm uses multiple hash functions to map the key content of the data packet to a fixed-size Sketch structure, thereby generating a compact and effective feature representation and completing the feature extraction. The specific calculation process is as follows: In the feature extraction method based on the Bloom filter, first input the packet timestamp t n , data packet length l, maximum data packet length l m , number of packets s, cumulative packet length l s , time threshold t resold , packet number threshold len t , quintuple k and bit array B. Use three hash functions H1, H2, and H3 to calculate the quintuple k, and obtain the mapping position of k in the bit array B h1=B[H1(k)], h2=B[H2(k)], and h3=B[H3(k)]. If at least one of h1, h2, and h3 is 0, it is determined to be a new flow, and the corresponding bit is set to 1. The F_Initialize function is called to initialize the new flow. In this function, the current data packet timestamp t n Timestamp t assigned to the first packet of the flow f and the timestamp of the previous data packet t l , assign the packet length l to the packet size cumulative sum register l s , set the packet count register s to 1, and assign the packet length l to the maximum packet length register l m , and finally returns the initialized statistical value. If h1, h2, and h3 are not 0, the subsequent statistics will be calculated based on the difference between the packet timestamp and the preset time threshold tresold The comparison result is processed accordingly (the relevant judgment logic is described later in the article, but it is not fully presented here). When the number of packets s reaches the preset packet number threshold len t When the output flow feature F(t n ,t l ,t f ,t i ).
[0019] (1.2) In the feature selection stage, a set of features is selected, which includes the sum of the packet lengths of the flow segments within the statistical interval, the overall time interval, and the maximum packet length. These features are optimized through a balancing strategy while maintaining low complexity.
[0020] (1.3) Next, Bloom filter is used to store these features. Bloom filter consists of a bit array, a set of hash functions and a query mechanism. In order to reduce the probability of hash collision, a time threshold detection mechanism is introduced. First, all the bit arrays are set to 0, indicating an empty set, to prepare space for storing elements. Then, for the elements to be stored, the position in the bit array is calculated by the hash function, and the corresponding position is set to 1. Since multiple hash functions are used, the probability of collision and misjudgment can be effectively reduced. When querying an element, the position is obtained through the hash function and the bit value is checked. If all are 1, the element may exist (there is a low false alarm rate); if there is 0, the element must not exist. This is the basis for the rapid existence judgment of the Bloom filter.
[0021] (1.4) Finally, the storage of service flow features is completed. For data packets in the service flow, the location of the packets in the Sketch structure and Bloom filter is located by the hash function based on the five-tuple, and the features are stored. For new flows, the bit is initialized; for conflicting flows, the register is updated or reset by comparing the timestamp with the threshold. When the preset packet number threshold is reached, features are extracted from the Sketch structure and Bloom filter to form a feature vector, which is then input into the pre-trained machine learning model for classification.
[0022] In this way, the time correlation of data packets and the dynamic changes of traffic are combined, and the conflict detection mechanism of the Bloom filter based on the Sketch structure is optimized. Based on the Sketch method, the data packets are coarsely classified using hash technology to obtain the estimated value of the measurement, which meets the needs of real-time measurement in high-speed environments and saves computing and space overhead. It not only ensures the effective extraction of features, but also improves the adaptability and accuracy of the algorithm in a limited resource environment. In the process of updating the hash table, the difference between the timestamp of the current data packet and the previous data packet is compared to determine whether it is a conflict, which significantly reduces the probability of hash conflicts.
[0023] As an improvement of the present invention, the specific method of step (2) is as follows:
[0024] Traffic classification is implemented on the P4 switch by simplifying the decision tree structure, avoiding creating matching items one by one for large numerical interval values, using bit operations to simplify conditions, and reducing hardware resource usage without losing classification performance.
[0025] That is, the following processing is performed. Taking the original decision tree branch condition processing as an example, for example, when processing the transition condition "<8192" from branch 1 to branch 2, because its value is large and measured in microseconds, a large number of entries are required in the matching table. Innovative bit operations are used to shift all decision conditions and feature values right by 3 bits, reducing this condition to "[0,1024]", greatly reducing the size of the matching table. After rigorous experimental verification, this move significantly reduces the complexity of model implementation, improves resource utilization efficiency and classification response speed, and makes the decision tree run more robustly and efficiently in the P4 environment while ensuring the basic stability of classification performance1.
[0026] By improving the decision tree structure, that is, simplifying the decision conditions through bit operations, the number of required matching items is reduced. Experimental results show that it does not significantly affect the performance of the classifier. It can significantly reduce the complexity of the decision tree model implementation in the P4 switch while maintaining the classifier performance.
[0027] As an improvement of the present invention, the specific method of step (3) is as follows:
[0028] (3.1) The ranking calculation based on the cumulative queuing delay includes the following steps:
[0029] The ingress switch uses the five-tuple as the session basis for an independent network. When a new network session requests a custom header at the ingress switch, the controller queries the network topology to identify the expected transmission path and configures the flow table rules of the edge switch.
[0030] For the intermediate switch, the data packets are divided into idle state, normal state and congested state according to the ratio of the cumulative queuing delay of the data packets to the ideal queuing delay. The priority of the data packet is determined by the link load. When the ratio is less than 0.5, it is determined to be in idle state, and the priority of the current data packet is reduced by one level. The priority of the data packet is reduced by one level in the current switch forwarding stage. When the ratio is between 0.5-1, it is determined to be in normal state and no priority processing is performed. When the ratio is greater than 1, it is determined to be in congested state and the priority of the data packet is increased by one level.
[0031] The egress switch is similar to the intermediate switch. It needs to determine the transmission status of the data packet and adjust the priority. At the same time, to ensure that the data packet is forwarded normally in the external network, the custom header needs to be deleted.
[0032] (3.2) Packet scheduling is performed based on AD-PIFO. When a data packet enters the forwarding link of the switch, the priority of the data packet is determined by the information carried in the custom header. The algorithm starts from the highest priority FIFO queue, checks the boundary threshold of each queue in sequence, and compares it with the priority of the data packet until the data packet meets the queue entry conditions. The data packet is pushed down. When the data packet is pushed down, it is detected whether there is a reverse order phenomenon. If there is a reverse order, the algorithm will reduce the boundary threshold of the corresponding queue and lower the priority upper limit to ensure that high-priority data packets are forwarded first.
[0033] This step designs an adjustable PIFO (AD-PIFO) algorithm, which uses PIFO blocks to assign priorities to data packets. Its queue division is based on priority standards rather than flow boundaries. By dynamically adjusting the boundary threshold, data packets of different business types flow between multiple queues according to priority, improving the utilization efficiency of network resources. It can also detect the reverse order phenomenon when the data packet is pushed down and adjust the queue threshold to reduce the chance of low-priority data packets entering the high-priority queue and ensure the priority forwarding of high-priority data packets.
[0034] A multi-service traffic group scheduling method for P4 switches is proposed. The services are classified at the application level to solve the problems of coarse feature engineering of existing P4 classification algorithms and insufficient dynamic adjustability of scheduling algorithms, thereby improving network data packet processing efficiency and service quality (QoS).
[0035] Compared with the prior art, the advantages of the present invention are as follows:
[0036] 1. Advantages of feature extraction and storage
[0037] (1) Innovative combination and optimization: The present invention innovatively combines the Sketch algorithm with the Bloom filter, and introduces a time threshold detection mechanism to optimize conflict detection. Compared with the prior art that simply extracts TCP stream packet features, the present invention comprehensively considers the time correlation of packets and the dynamic changes of traffic, and can more efficiently and accurately extract and store features in resource-constrained P4 switches, adapting to complex network environments. For example, in a multi-tenant scenario in a data center network, it can accurately capture the features of different business flows and ensure classification accuracy. Other documents do not use similar optimized feature extraction and storage methods.
[0038] (2) Targetedness and balance of feature selection: The present invention selects the sum of the packet lengths of the flow segments within the statistical interval, the overall time interval, and the maximum packet length as features, taking into account low complexity and effectiveness, and optimizing the selection through a balanced strategy. The existing technology mainly performs scheduling based on the network quintuple and the time limit and size of the business flow, and the focus of feature selection is different. The feature selection of the present invention is more targeted at traffic classification in the P4 switch environment, and can effectively support business flow classification and scheduling decisions.
[0039] 2. Advantages of business flow classification
[0040] (1) Decision tree optimization and hardware adaptation: In view of the characteristics of the P4 switch data plane, the present invention simplifies the decision tree conditions through bit operations, reduces the number of matching entries, reduces hardware resource usage, and ensures classification performance. For example, the decision tree branch condition "<8192" is optimized to "[0,1024]", which improves resource utilization efficiency and classification response speed. Although the prior art also performs structural transformation on the decision tree, it does not perform similar optimization for the resource limitations of specific hardware such as the P4 switch.
[0041] (2) Excellent performance in multi-classification scenarios: The present invention performs well in multi-classification scenarios. Through improved feature extraction and decision tree optimization, it can effectively distinguish multiple service flows. Experiments show that it is superior to similar methods in terms of accuracy, precision, and recall. For example, compared with the SMASH and IDEAFIX models, it has obvious advantages in processing complex network traffic classification. The existing technology mainly focuses on the application of CNN and LSTM-based classification models in QoS queue scheduling, and is not as targeted and performant as the present invention in multi-classification scenarios of P4 switches.
[0042] 3. Advantages of group scheduling
[0043] (1) Dynamic priority adjustment mechanism: The present invention designs a variable priority packet scheduling method AD PIFO based on the accumulated queuing delay and the ideal queuing delay of the service, which can dynamically adjust the forwarding priority according to the real-time status of the data packet in the network. When the network is congested or the load changes, the priority of the data packet can be adjusted in time to ensure the QoS requirements of high-priority services. In contrast, the existing technology mainly schedules based on the time limit and size of the service flow, and does not fully combine the real-time delay feedback of the data packet in the network transmission to adjust the priority;
[0044] (2) Efficient queue management and resource utilization: The AD PIFO algorithm divides queues based on priority standards. By dynamically adjusting the boundary threshold, packets of different service types can flow reasonably between multiple queues, thereby improving the efficiency of network resource utilization. At the same time, it can detect and handle the reverse order phenomenon when packets are pushed down, ensuring that high-priority packets are forwarded first. This queue management method has more advantages in improving the overall performance and efficiency of the network than traditional flow-based queue management strategies and other methods that do not specify similar efficient queue management mechanisms in other documents.
[0045] 4. Comprehensive performance advantages
[0046] (1) Strong hardware adaptability: This paper is specially designed for P4 switches, taking into full account their resource limitations and operation modes, and realizing efficient traffic classification and scheduling in the P4 switch environment. Compared with other studies that do not focus on P4 switches, this paper can better leverage the programmable advantages of P4 switches, improve their performance in complex network traffic processing, and meet the needs of fine-grained scheduling in data packets in SDN networks.
[0047] (2) Multi-scenario stability and practicality: It is stable and effective under different load conditions, and the additional time overhead for data packet transmission is reasonable. Through the example verification of multi-tenant scenarios in data center networks and traffic management in different business periods in metropolitan area networks, it can ensure the service quality of various services in a complex and changeable network environment. Compared with other studies, the method of the present invention has stronger adaptability and practicality in actual network scenarios, providing a more reliable solution for network optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A method for scheduling service groups in a P4 switch.
[0049] Figure 2 This is a schematic diagram of the processing flow of the Bloom filter.
[0050] Figure 3 is the original decision tree model instance,
[0051] Figure 4 To improve the decision tree model instance,
[0052] Figure 5 The overall architecture of the variable priority group scheduling method is shown in Figure 1.
[0053] Figure 6 The specific process of the ranking calculation method based on the cumulative queuing delay is as follows:
[0054] Figure 7 This is the specific process of the AD-PIFO algorithm. DETAILED DESCRIPTION
[0055] In order to deepen the recognition and understanding of the present invention, the present invention is described in detail below with reference to the accompanying drawings.
[0056] Example: See Figure 1 — Figure 7 , a multi-service traffic packet scheduling method for P4 switches, builds an implementation path from all aspects of feature extraction, service flow classification and packet scheduling.
[0057] The service flow data packet first enters the data plane, and the data packet is parsed by using the five-tuple splitting method. Then, the sum of the data packet lengths of the flow segments within the statistical interval, the overall time interval, and the maximum data packet length are selected. The feature engineering based on the sketch algorithm is completed in the Bloom filter, and the feature vector is extracted. The decision tree is trained and the structure is improved in the deployment classification model. The extracted feature vector is classified in the improved decision tree-based service classification model, and then the priority is corrected by delay calculation. Then enter the packet scheduling module, complete the five-tuple upload according to the custom header, perform service flow priority judgment, and then complete the flow table rule update and rule distribution. Finally, boundary perception, data packet push and boundary update are completed in sequence to complete data packet forwarding.
[0058] The method comprises the following steps:
[0059] (1) Extract and store features based on the Sketch algorithm, use the five-tuple in the packet header to determine the service flow, and store the packet features in a register through a Bloom filter based on the Sketch algorithm. When the feature data reaches a certain amount, it is formed into a feature vector and input into the pre-trained machine learning model for classification.
[0060] The P4 switch has powerful programmable capabilities and can flexibly process data packets. The feature extraction and storage based on the Sketch algorithm can efficiently complete the extraction and storage of data packet features on the data plane with the help of the programmable characteristics of P4, providing a data basis for subsequent business flow classification, which is in line with the characteristics of the P4 switch that can customize the data packet processing logic according to different needs.
[0061] (2) Based on the decision tree, the service flow classification is performed. The decision tree node range compression method based on bit operation is used to adjust the decision tree node range, reduce the number of entry matches, adapt to the P4 switch data plane requirements, and execute the classification process through the P4 match-action pipeline to achieve service flow classification. Bit operation directly operates on binary bits, and right shift operation plays a key role in this method.
[0062] The data plane of P4 switches usually has limited resources and is not suitable for complex calculations. The decision tree node range compression method based on bit operations can effectively reduce the number of entry matches, reduce the computational complexity, and adapt well to the resource limitations of the data plane of P4 switches. At the same time, P4's match-action pipeline provides a natural implementation path for the execution of the decision tree classification process, which can efficiently complete the business flow classification task.
[0063] (3) Packet scheduling is performed based on variable priorities. When a data packet enters the first P4 switch, a custom header is added to record the priority and accumulated queuing delay. A priority adjustment module is connected after the delay statistics module to modify the packet forwarding priority based on the state classification of the accumulated delay ratio. That is, the data packet is divided into idle state, normal state and congested state according to the ratio of the accumulated queuing delay of the data packet to the ideal queuing delay, and the priority is adjusted based on this state classification. At the same time, the dynamic boundary threshold update rule is followed, and the PIFO algorithm is used to select the queue according to the priority and adjust the sorting.
[0064] P4 supports flexible customization and modification of packet headers, which enables custom headers to be added to record priority and accumulated queuing delay when packets enter the switch. Moreover, the programmability of P4 facilitates the implementation of priority adjustment modules and dynamic boundary threshold update rules. Combined with the PIFO algorithm, it can dynamically adjust the priority and select the queue according to the real-time status of the packet, fully demonstrating the flexibility and programmability of the P4 switch in packet scheduling.
[0065] The specific method of step (1) is as follows:
[0066] (1.1) First, the five-tuple in the packet header (source IP, destination IP, protocol number, source port number, and destination port number) is used to determine the service flow. Then, the five-tuple is hashed by a hash function to determine the service flow and update the traffic features. In this process, the Sketch algorithm is introduced to efficiently calculate the features. The Sketch algorithm uses multiple hash functions to map the key content of the data packet to a fixed-size Sketch structure, thereby generating a compact and effective feature representation and completing the feature extraction. The specific calculation process is as follows: In the feature extraction method based on the Bloom filter, first input the packet timestamp t n , data packet length l, maximum data packet length l m , number of packets s, cumulative packet length l s , time threshold t resold , packet number threshold len t , quintuple k and bit array B. Use three hash functions H1, H2, and H3 to calculate the quintuple k, and obtain the mapping position of k in the bit array B h1=B[H1(k)], h2=B[H2(k)], and h3=B[H3(k)]. If at least one of h1, h2, and h3 is 0, it is determined to be a new flow, and the corresponding bit is set to 1. The F_Initialize function is called to initialize the new flow. In this function, the current data packet timestamp t n Timestamp t assigned to the first packet of the flow f and the timestamp of the previous data packet t l, assign the packet length l to the packet size cumulative sum register l s , set the packet count register s to 1, and assign the packet length l to the maximum packet length register l m , and finally returns the initialized statistical value. If h1, h2, and h3 are not 0, the subsequent statistics will be calculated based on the difference between the packet timestamp and the preset time threshold t resold The comparison result is processed accordingly (the relevant judgment logic is described later in the article, but it is not fully presented here). When the number of packets s reaches the preset packet number threshold len t When the output flow feature F(t n ,t l ,t f ,t i ).
[0067] (1.2) In the feature selection stage, a set of features is selected, which includes the sum of the packet lengths of the flow segments within the statistical interval, the overall time interval, and the maximum packet length. These features are optimized through a balancing strategy while maintaining low complexity.
[0068] (1.3) Next, Bloom filter is used to store these features. Bloom filter consists of a bit array, a set of hash functions and a query mechanism. In order to reduce the probability of hash collision, a time threshold detection mechanism is introduced. First, all the bit arrays are set to 0, indicating an empty set, to prepare space for storing elements. Then, for the elements to be stored, the position in the bit array is calculated by the hash function, and the corresponding position is set to 1. Since multiple hash functions are used, the probability of collision and misjudgment can be effectively reduced. When querying an element, the position is obtained through the hash function and the bit value is checked. If all are 1, the element may exist (there is a low false alarm rate); if there is 0, the element must not exist. This is the basis for the rapid existence judgment of the Bloom filter.
[0069] (1.4) Finally, the storage of service flow features is completed. For data packets in the service flow, the location of the packets in the Sketch structure and Bloom filter is located by the hash function based on the five-tuple, and the features are stored. For new flows, the bit is initialized; for conflicting flows, the register is updated or reset by comparing the timestamp with the threshold. When the preset packet number threshold is reached, features are extracted from the Sketch structure and Bloom filter to form a feature vector, which is then input into the pre-trained machine learning model for classification.
[0070] In this way, the time correlation of data packets and the dynamic changes of traffic are combined, and the conflict detection mechanism of the Bloom filter based on the Sketch structure is optimized. Based on the Sketch method, the data packets are coarsely classified using hash technology to obtain the estimated value of the measurement, which meets the needs of real-time measurement in high-speed environments and saves computing and space overhead. It not only ensures the effective extraction of features, but also improves the adaptability and accuracy of the algorithm in a limited resource environment. In the process of updating the hash table, the difference between the timestamp of the current data packet and the previous data packet is compared to determine whether it is a conflict, which significantly reduces the probability of hash conflicts.
[0071] The specific method of step (2) is as follows:
[0072] Traffic classification is implemented on the P4 switch by simplifying the decision tree structure, avoiding creating matching items one by one for large numerical interval values, using bit operations to simplify conditions, and reducing hardware resource usage without losing classification performance.
[0073] That is, the following processing is performed. Taking the original decision tree branch condition processing as an example, for example, when processing the transition condition "<8192" from branch 1 to branch 2, because its value is large and measured in microseconds, a large number of entries are required in the matching table. Innovative bit operations are used to shift all decision conditions and feature values right by 3 bits, reducing this condition to "[0,1024]", greatly reducing the size of the matching table. After rigorous experimental verification, this move significantly reduces the complexity of model implementation, improves resource utilization efficiency and classification response speed, and makes the decision tree run more robustly and efficiently in the P4 environment while ensuring the basic stability of classification performance1.
[0074] By improving the decision tree structure, that is, simplifying the decision conditions through bit operations, the number of required matching items is reduced. Experimental results show that it does not significantly affect the performance of the classifier. It can significantly reduce the complexity of the decision tree model implementation in the P4 switch while maintaining the classifier performance.
[0075] The specific method of step (3) is as follows:
[0076] (3.1) The ranking calculation based on the cumulative queuing delay includes the following steps:
[0077] The ingress switch uses the five-tuple as the session basis for an independent network. When a new network session requests a custom header at the ingress switch, the controller queries the network topology to identify the expected transmission path and configures the flow table rules of the edge switch.
[0078] For the intermediate switch, the data packets are divided into idle state, normal state and congested state according to the ratio of the cumulative queuing delay of the data packets to the ideal queuing delay. The priority of the data packet is determined by the link load. When the ratio is less than 0.5, it is determined to be in idle state, and the priority of the current data packet is reduced by one level. The priority of the data packet is reduced by one level in the current switch forwarding stage. When the ratio is between 0.5-1, it is determined to be in normal state and no priority processing is performed. When the ratio is greater than 1, it is determined to be in congested state and the priority of the data packet is increased by one level.
[0079] The egress switch is similar to the intermediate switch. It needs to determine the transmission status of the data packet and adjust the priority. At the same time, to ensure that the data packet is forwarded normally in the external network, the custom header needs to be deleted.
[0080] (3.2) Packet scheduling is performed based on AD-PIFO. When a data packet enters the forwarding link of the switch, the priority of the data packet is determined by the information carried in the custom header. The algorithm starts from the highest priority FIFO queue, checks the boundary threshold of each queue in sequence, and compares it with the priority of the data packet until the data packet meets the queue entry conditions. The data packet is pushed down. When the data packet is pushed down, it is detected whether there is a reverse order phenomenon. If there is a reverse order, the algorithm will reduce the boundary threshold of the corresponding queue and lower the priority upper limit to ensure that high-priority data packets are forwarded first.
[0081] This step designs an adjustable PIFO (AD-PIFO) algorithm, which uses PIFO blocks to assign priorities to data packets. Its queue division is based on priority standards rather than flow boundaries. By dynamically adjusting the boundary threshold, data packets of different business types flow between multiple queues according to priority, improving the utilization efficiency of network resources. It can also detect the reverse order phenomenon when the data packet is pushed down and adjust the queue threshold to reduce the chance of low-priority data packets entering the high-priority queue and ensure the priority forwarding of high-priority data packets.
[0083] like Figure 1 As shown, a multi-service traffic group scheduling method in a P4 switch is shown.
[0084] Among them, in the data plane, the business flow is locked according to the five-tuple of the data packet header, and the features are stored in the register by the Sketch algorithm. After the quantity is sufficient, the model is input for classification and label transmission, and the Bloom filter is optimized and the key features are selected during the process; when classifying the business flow, the decision tree algorithm is selected according to the P4 characteristics and simplified and optimized, and the priority is corrected and the delay is calculated based on the priority scheduling algorithm based on the accumulated queuing delay; the packet scheduling module is determined by the three steps of boundary perception, packet push down and boundary update to construct a ranking calculation and packet scheduling module architecture, and the priority is dynamically adjusted according to the delay information in the control plane, and the queue management is optimized by the AD-PIFO algorithm. The present invention advances from feature capture, classification decision to packet scheduling in sequence, effectively improving the performance of the P4 switch.
[0085] In order to better illustrate the service group scheduling method for P4 switches, the processing flow of Bloom filter and its combination with Sketch algorithm are introduced first. Figure 2 As shown in the figure, the Bloom filter uses a predefined bit array to represent the elements in the set. In the initialization phase, all bits are set to 0, indicating that the set is empty. At the same time, multiple hash functions required by the Sketch algorithm are also predefined. These hash functions will be used to map the characteristics of the data packet to the Sketch structure and the bit array of the Bloom filter.
[0086] When adding an element (i.e., a feature of a packet) to a Bloom filter, the feature value of the packet is first calculated using the Sketch algorithm. The Sketch algorithm uses multiple hash functions to map the key content of the packet (such as a five-tuple, packet length, etc.) into a fixed-size Sketch structure to generate a compact feature representation. This feature value and the original packet (or part of its features) are then used as input and mapped to a specific position in the bit array through one or more hash functions of the Bloom filter. Each hash function calculates the input and sets the bit at the corresponding position in the bit array to 1.
[0087] For the query operation, the Sketch algorithm is also used to calculate the feature value of the packet to be queried. Then, the Bloom filter uses a set of hash functions to calculate the hash value of this feature value, and then checks the bits at these positions in the bit array. If all the bits at the calculated positions are 1, then the element (i.e., the feature of the packet) may exist in the set; if there is a bit at any position that is 0, it can be determined that the element is not in the set.
[0088] During the update process of the hash table (i.e., the bit array of the Bloom filter), if a hash conflict occurs (i.e., the characteristics of multiple data packets are mapped to the same bit through the hash function, resulting in the bits at these positions being all 1), an additional judgment mechanism needs to be introduced. Specifically, the difference in timestamp between the current data packet and the previous data packet can be compared to determine whether it is a true conflict. If the time interval is less than the preset threshold, it is considered to be a data packet of the same business flow, and the hash table is updated (i.e., the bits at these positions are kept as 1); if the time interval is greater than or equal to the threshold, it is considered to be the beginning of a new business flow. At this time, the register needs to be reset (i.e., the bits at these positions are reset to 0), and the Sketch structure and Bloom filter are updated to reflect the new business flow characteristics.
[0089] To explain the multi-service traffic group scheduling method for P4 switches in detail, the present invention will introduce the improvement of the decision tree model. In the data plane of P4 switches, machine learning is applied with limited resources. Although the decision tree is suitable for this environment, the original decision tree is Figure 3 As shown in the figure, some decision conditions have large values, which will make the P4 switch matching table large, because it needs to create a matching item for each value in the interval [0,8192], which increases the difficulty of hardware implementation and resource consumption. To address the above problems, an improved decision tree structure is designed ( Figure 4 ), all decision conditions are divided by 8 to simplify the condition judgment. Although the P4 switch does not support floating point and multiplication and division operations, it can be implemented through bit operations, that is, all decision conditions and eigenvalues are shifted right by 3 bits, so that the original judgment condition [0,8192] is converted to [0,1024], which greatly reduces the number of matching items, optimizes the balance between resource utilization and classification efficiency, and ensures the funny operation of the network and accurate traffic management.
[0090] To further illustrate the multi-service traffic group scheduling method for P4 switches, the overall architecture of the variable priority group scheduling method is introduced. Figure 5 — Figure 7The workflow of the variable priority packet scheduling method is rigorous and efficient. Its overall architecture relies on the coordinated operation of two core modules: ranking calculation and PIFO-based packet scheduling. When the data packet arrives at the switch, the packet parsing submodule takes the lead, accurately extracts the five-tuple hash value and uploads it to the control plane. The control plane quickly updates the flow table rules according to the preset rules, and then sends them to the data planes of each switch. At this moment, the ingress switch assigns a custom header to the data packet, records the priority and cumulative queuing delay in detail, and starts the transmission journey of the data packet in the network. In the delay calculation submodule, the system continuously compares the cumulative queuing delay with the ideal queuing delay, and the results are seamlessly transmitted to the priority adjustment submodule, which becomes the key basis for adjusting the forwarding priority of the data packet. Once the data packet enters the packet scheduling module, the priority adjustment submodule flexibly corrects the priority according to the delay comparison result, and the packet forwarding submodule accurately screens the queue and properly adjusts the order according to the PIFO algorithm. The data packet is properly encapsulated by the reverse parser and forwarded robustly. In the forwarding process, the custom header information is continuously and dynamically updated to reflect the status changes of the data packet in real time. The egress switch shoulders a heavy responsibility. It not only accurately determines the data packet transmission status and adjusts the priority in a timely manner, but also decisively deletes the custom header when the data packet is about to enter the external network to ensure that the data packet is transmitted smoothly in the external network environment.
[0091] The present invention is described below in conjunction with specific examples and the accompanying drawings. The invention has the following steps:
[0092] (1) Extract and store features based on the Sketch algorithm, use the five-tuple in the packet header to determine the service flow, and store the packet features in a register through a Bloom filter based on the Sketch algorithm. When the feature data reaches a certain amount, it is formed into a feature vector and input into the pre-trained machine learning model for classification.
[0093] The P4 switch has powerful programmable capabilities and can flexibly process data packets. The feature extraction and storage based on the Sketch algorithm can efficiently complete the extraction and storage of data packet features on the data plane with the help of the programmable characteristics of P4, providing a data basis for subsequent business flow classification, which is in line with the characteristics of the P4 switch that can customize the data packet processing logic according to different needs.
[0094] (2) Based on the decision tree, the service flow classification is performed. The decision tree node range compression method based on bit operation is used to adjust the decision tree node range, reduce the number of entry matches, adapt to the P4 switch data plane requirements, and execute the classification process through the P4 match-action pipeline to achieve service flow classification. Bit operation directly operates on binary bits, and right shift operation plays a key role in this method.
[0095] The data plane of P4 switches usually has limited resources and is not suitable for complex calculations. The decision tree node range compression method based on bit operations can effectively reduce the number of entry matches, reduce the computational complexity, and adapt well to the resource limitations of the data plane of P4 switches. At the same time, P4's match-action pipeline provides a natural implementation path for the execution of the decision tree classification process, which can efficiently complete the business flow classification task.
[0096] (3) Packet scheduling is performed based on variable priorities. When a data packet enters the first P4 switch, a custom header is added to record the priority and accumulated queuing delay. A priority adjustment module is connected after the delay statistics module to modify the packet forwarding priority based on the state classification of the accumulated delay ratio. That is, the data packet is divided into idle state, normal state and congested state according to the ratio of the accumulated queuing delay of the data packet to the ideal queuing delay, and the priority is adjusted based on this state classification. At the same time, the dynamic boundary threshold update rule is followed, and the PIFO algorithm is used to select the queue according to the priority and adjust the sorting.
[0097] P4 supports flexible customization and modification of packet headers, which enables custom headers to be added to record priority and accumulated queuing delay when packets enter the switch. Moreover, the programmability of P4 facilitates the implementation of priority adjustment modules and dynamic boundary threshold update rules. Combined with the PIFO algorithm, it can dynamically adjust the priority and select the queue according to the real-time status of the packet, fully demonstrating the flexibility and programmability of the P4 switch in packet scheduling.
[0098] In order to explain in detail the working steps of a multi-service traffic group scheduling method for P4 switches, three use cases are used to explain the different working modes of the method under different measurements.
[0099] Specific Example 1: Traffic Scheduling in a Multi-Tenant Data Center Network Scenario
[0100] The data center network provides services for multiple tenants with various types of services, including big data analysis (high priority), cloud computing services (medium priority), and regular office applications (low priority). The network traffic load is constantly changing, and tenant business peaks occur at different times, which places extremely high demands on network resource allocation and scheduling accuracy, and at the same time, traffic interference between different tenant services must be strictly controlled.
[0101] In the feature extraction and classification stage, when big data analysis business traffic is generated, the system uses the five-tuple in the packet header to lock the business flow and accumulates and stores features in the register through the Sketch algorithm. For example, in the data-intensive computing stage, a large number of data packets are transmitted in a short period of time. The system accurately extracts features such as the sharp increase in the total length of data packets, the stable transmission time interval, and the large maximum data packet length, and forms a feature vector input decision tree model classification. Bloom filters efficiently store features with a low probability of conflict. Based on the time threshold detection mechanism, they accurately handle the storage requirements of frequently updated traffic features to ensure accurate classification of big data business traffic.
[0102] When cloud computing services are running, such as during the creation of virtual machine instances and data interaction phases, the system continuously tracks data packet features. Based on the five-tuple hash positioning feature storage and update, the Bloom filter ensures the accuracy of feature queries, and through the optimized hash conflict handling mechanism, accurately distinguishes cloud computing services from other business traffic, providing an accurate basis for subsequent scheduling.
[0103] When regular office application traffic appears, the system operates according to the feature extraction process, based on the regularity of the time interval between data packet arrivals and the relatively stable length of data packets, combined with the Bloom filter storage query function, to accurately determine the service type and priority, and effectively distinguish different priority services in the feature extraction and classification stages.
[0104] In the scheduling phase, packets enter the switch, and the AD-PIFO algorithm dominates the scheduling process. When the data packet of the big data analysis service arrives at the ingress switch, it is marked with a custom header with priority and accumulated queuing delay. The control plane updates the flow table based on the network topology and business rules. In the delay statistics module, if the queuing delay of the data packet exceeds the ideal value, the algorithm increases its priority in the switch, prioritizes the allocation of bandwidth resources, and ensures data analysis efficiency.
[0105] Cloud computing service data packets are scheduled based on their priority. When network congestion causes queuing delays to change, the algorithm adjusts the priority based on the delay ratio. If the queuing delay of a medium-priority cloud computing service packet is 1.2 times the ideal value, the priority is increased by one level; if it is less than 0.5 times, the priority is decreased by one level, ensuring a dynamic balance in service response performance and maintaining the overall service quality of the data center.
[0106] Conventional office application data packets are scheduled in low-priority queues, and resource allocation ratios are adjusted dynamically according to network load. When high- and medium-priority business loads are low, their bandwidth share is appropriately increased; when high- and medium-priority businesses are busy, bandwidth is strictly limited to prevent interference with key businesses and ensure the orderly and efficient operation of multi-tenant businesses in the data center.
[0107] Specific Example 2: Traffic Management in Different Service Periods of Metropolitan Area Network
[0108] The metropolitan area network covers a wide area and provides rich service types, including daytime commercial office (high priority such as enterprise dedicated line services, medium priority such as ordinary enterprise office networks, and low priority such as street shop networks) and nighttime residents' entertainment (high priority such as online games, medium priority such as high-definition video, and low priority such as social media browsing). Network traffic shows obvious alternating characteristics of daytime and nighttime business peaks, and is affected by regional activity patterns. The traffic load distribution in each region is uneven, requiring flexible allocation of network resources.
[0109] During the daytime business office hours, the enterprise dedicated line business traffic is active, and the system identifies business flows based on quintuples. Taking the high-frequency transaction data transmission of financial institutions as an example, this type of data packet has the required characteristics of high frequency, small batch, and low latency. The system captures these characteristics through the Sketch algorithm, which maintains a compact data structure to approximately track the frequency of occurrence of elements in the data stream, thereby capturing the key statistical information of the data packet in a limited space. Specifically, the Sketch algorithm uses several hash functions to map the characteristics of the data packet into a bit array, so that the distribution of the data packet characteristics can be quickly estimated at a low storage cost. The application of this algorithm enables Bloom filters to accurately store and quickly query classification, ensuring that financial business data packets can be quickly identified and processed, achieving low-latency, highly reliable network transmission, and giving priority to the allocation and scheduling of key business network resources.
[0110] When ordinary corporate office network traffic emerges, the system extracts data packet characteristics such as file sharing and email sending and receiving. Based on the characteristics of moderate data packet size and uncertain time intervals, combined with Bloom filter optimization conflict handling function, business flows are accurately classified and scheduled in order according to priority, improving the overall efficiency of the office network and reducing the cost of corporate network operations.
[0111] During the nighttime entertainment period, when online gaming services are at their peak, the system focuses on packet features, such as bursts of real-time game packets and latency-sensitive features. The Sketch algorithm efficiently extracts features, and the Bloom filter ensures feature accuracy, providing accurate classification for the AD-PIFO algorithm, ensuring low-latency experience for gamers and improving user satisfaction.
[0112] In the transmission of high-definition video service traffic, the system accurately classifies data packets based on their large length and stable transmission characteristics. Bloom filters help store and query features, provide a basis for scheduling, allocate bandwidth based on network load, ensure smooth video playback, and improve the quality of residents' entertainment services.
[0113] When processing social media browsing business traffic, the system classifies data packets based on their scattered nature and long time intervals, and combines it with a feature storage query mechanism to rationally allocate network resources and ensure a smooth online social experience for residents.
[0114] In terms of scheduling strategy adjustment, the AD-PIFO algorithm dynamically schedules based on the alternating business peaks during the day and night in the metropolitan area network. During the daytime commercial office peak, priority is given to protecting enterprise dedicated lines and high-demand office business bandwidth resources, increasing the boundary threshold of high-priority business queues, and limiting the influx of low-priority business traffic. For example, the priority of enterprise dedicated line business packets is adjusted according to latency, and priority is given to forwarding during congestion to ensure the smooth flow of key businesses such as financial transactions.
[0115] During the peak entertainment hours at night, the algorithm optimizes the scheduling of high-priority services such as online games and high-definition videos. When the network is congested, the algorithm increases the priority of gaming services and reduces the bandwidth of low-priority packets for video services to ensure real-time interactive gaming experience. Based on changes in network load, the thresholds and resource allocation ratios of each service queue are flexibly adjusted. For example, resources are dynamically adjusted in different areas based on user density and business needs to improve metropolitan area network resource utilization and user experience.
[0116] Experiment 1: Experimental evaluation of traffic flow classification method
[0117] Purpose:
[0118] The classification performance of the business flow classification method for limited resource scenarios is comprehensively evaluated on the business flow dataset extracted from the real network environment, and compared with other existing classification methods, the advantages of the proposed method are confirmed.
[0119] Experimental procedures
[0120] 1. We selected the Bloom filter with the highest query efficiency in Sketch as the feature access structure. At the same time, combined with the time correlation of data packets, we improved the conflict judgment method to reduce the misjudgment rate. In this process, Sketch captures and stores the key features of data packets at a lower storage cost by maintaining a probabilistic data structure, and maps the data packet features to bit arrays through hash functions. When the data packet features reach a certain amount, the information of these bit arrays is used to form feature vectors, which are then input into the pre-trained machine learning model for classification. This method not only improves the efficiency of feature extraction, but also enhances the accuracy of the model's classification of data packets. With the SMASH and IDEAFIX models as the main comparison objects, our method shows better performance in feature extraction and classification.
[0121] 2. Use BMV2 virtual software switch and Mininet to set up a simple network topology, including 2 virtual hosts, 1 P4 software switch and 1 central controller. All experiments are run on high-performance servers with configurations such as CPU: Intel Core i9-13900KF@3GHz, memory: 64GB, etc.
[0122] 3. Use ISCXVPN2016 and self-collected traffic to build a mixed data set, exclude insufficient data of some business flows, supplement self-collected traffic balance samples, 70% for training set, 30% for test set, and use 5-fold cross-validation. Evaluate KNN, SVM, ANN and decision tree algorithms to determine the best hyperparameters. The decision tree model is converted to P4 language for data plane classification, and the remaining algorithms are classified in the control plane. Test the performance of packet-by-packet and flow-by-flow classification models, compare the proposed method with existing methods in terms of accuracy, precision and recall, and evaluate the performance loss of the decision tree model shift operation.
[0123] Experiment 2: Experimental evaluation of packet scheduling method
[0124] Purpose:
[0125] Verify whether the scheduling scheme based on cumulative queuing delay can provide differentiated services for data flows of different priorities, compare the key performance indicators of existing priority scheduling schemes in different network environments, and test the performance of the AD-PIFO algorithm in the P4 switch and the impact of the number of logical queues and custom header time overhead.
[0126] Experimental steps:
[0127] 1. Based on the cumulative queuing delay, a priority adjustment algorithm and a packet scheduling algorithm based on AD-PIFO are proposed, and the SP-PIFO, AIFO and RIFO algorithms based on P4 are used as comparison objects. It is expanded on the basis of the 4.1.1 network topology, including 3 virtual hosts, 4 P4 software switches and 1 central controller, using the same server configuration.
[0128] 2. Use the 4.1.2 traffic data set, use tcpreplay and tcprewrite to adjust the packet sending rate, use the iperf tool to generate TCP traffic and add identifiers to distinguish priorities, randomly mix equal amounts and sizes into data sets, set the experimental link bottleneck to 100Mbps, configure the service flow classification model, and divide the service flow into high, medium, and low priorities. The default forwarding priorities are 1, 4, and 7 for the experiment.
[0129] 3. Measure the end-to-end latency and throughput of TCP packets generated by iperf under different algorithms in a congested network environment. The ideal latency is about 2ms when sending TCP traffic with the same load at 10Mbps. Use "tcpdump-i eth0" and the packet capture command to analyze the pcap file and calculate the latency and throughput.
[0130] 4. Use a specific data set to measure the dynamic changes in the throughput of the intermediate switch s2 and the average transmission time of packets in switch s2 under different loads at an 80Mbps sending rate and a 20-second data flow duration for the AD-PIFO algorithm, and compare the four algorithms.
[0131] 5. Evaluate the impact of the number of logical queues on the performance of the AD-PIFO algorithm using flow completion time as an indicator, and compare the performance of the SP-PIFO algorithm with that of the 8-queue and 16-queue configurations. Send service traffic at 20Mbps for 30 seconds, and calculate the average processing time of the ingress, intermediate, and egress switches (based on the difference between the average packet transmission time and the P4 program metadata field time) to evaluate the custom header time overhead.
[0132] It should be noted that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Equivalent replacements or substitutions made on the basis of the above technical solutions all fall within the protection scope of the present invention.
Claims
1. A multi-service traffic group scheduling method for a P4 switch, characterized in that: The method comprises the following steps: (1) Extract and store features based on the Sketch algorithm, use the five-tuple in the packet header to determine the service flow, and store the packet features in the register through a Bloom filter based on the Sketch algorithm and combined with a time threshold conflict detection mechanism. Construct a joint feature storage method of Sketch and Bloom filter. When the feature data reaches a certain amount, it is formed into a feature vector and input into the pre-trained machine learning model for classification. (2) Based on the decision tree, the service flow classification is carried out. The decision tree node range compression method based on bit operation is used to adjust the decision tree node range, reduce the number of entry matching, adapt to the P4 switch data plane requirements, and execute the classification process through the P4 match-action pipeline to realize service flow classification. Bit operation directly operates on binary bits. (3) Packet scheduling is performed based on variable priority. When a data packet enters the first P4 switch, a custom header is added to record the priority and cumulative queuing delay. A priority adjustment module is connected after the delay statistics module. The packet forwarding priority is modified based on the state classification of the cumulative delay ratio. That is, the data packet is divided into idle state, normal state and congested state according to the ratio of the cumulative queuing delay of the data packet to the ideal queuing delay. The priority is then adjusted based on this state classification. At the same time, the dynamic boundary threshold update rule is followed and the PIFO algorithm is used to select the queue according to the priority and adjust the sorting.
2. The multi-service traffic group scheduling method for a P4 switch according to claim 1 is characterized in that: The specific method of step (1) is as follows: (1.1) First, the five-tuple in the packet header (source IP, destination IP, protocol number, source port number, and destination port number) is used to determine the business flow. Then, the five-tuple is hashed by a hash function to determine the business flow and update the traffic features. In this process, the Sketch algorithm is introduced to efficiently calculate the features. The Sketch algorithm uses multiple hash functions to map the key content of the data packet to a fixed-size Sketch structure, thereby generating a compact and effective feature representation and completing the feature extraction. The specific calculation process is as follows: In the feature extraction method based on Bloom filter, first input the packet timestamp t n , data packet length l, maximum data packet length l m , number of packets s, cumulative packet length l s , time threshold t resold , packet number threshold len t , quintuple k and bit array B, use three hash functions H1, H2, H3 to calculate the quintuple k, and get the mapping position of k in bit array B h1=B[H1(k)], h2=B[H2(k)], h3=B[H3(k)] respectively. If at least one of h1, h2, h3 is 0, it is determined to be a new flow, and the corresponding bit is set to 1. The F_Initialize function is called to initialize the new flow. In this function, the current data packet timestamp t n Timestamp t assigned to the first packet of the flow f and the timestamp of the previous data packet t l , assign the packet length l to the packet size cumulative sum register l s , set the packet count register s to 1, and assign the packet length l to the maximum packet length register l m Finally, the initialized statistical value is returned. If h1, h2, and h3 are not 0, the statistical value will be calculated based on the difference between the packet timestamp and the preset time threshold t resold The comparison result is processed accordingly. When the number of packets s reaches the preset packet number threshold len t When the output flow feature F(t n ,t l ,t f ,t i ), (1.2) In the feature selection stage, a set of features including the sum of the packet lengths of the flow segments within the statistical interval, the overall time interval, and the maximum packet length are selected. These features are optimized through a balance strategy while maintaining low complexity. (1.3) Next, Bloom filter is used to store these features. Bloom filter consists of bit array, hash function set and query mechanism. In order to reduce the probability of hash collision, a time threshold detection mechanism is introduced. First, all bit arrays are set to 0, indicating an empty set, to prepare bit space for storing elements. Then, for the elements to be stored, the position in the bit array is calculated by hash function, and the corresponding position is set to 1. Since multiple hash functions are used, the probability of collision and misjudgment is effectively reduced. When querying elements, the position is obtained by hash function and the bit value is checked. If all are 1, the element may exist (with a low false alarm rate); if there is 0, the element must not exist. This is the basis for the rapid existence judgment of Bloom filter. (1.4) Finally, the storage of service flow features is completed. For data packets in the service flow, the position of the packets in the Sketch structure and Bloom filter is located by the hash function based on the five-tuple, and the features are stored. For new flows, the bit is initialized. For conflicting flows, the register is updated or reset by comparing the timestamp with the threshold. When the preset packet number threshold is reached, features are extracted from the Sketch structure and Bloom filter to form a feature vector, which is then input into the pre-trained machine learning model for classification.
3. The multi-service traffic group scheduling method for a P4 switch according to claim 1, characterized in that: The specific method of step (2) is as follows: Traffic classification is implemented on the P4 switch by simplifying the decision tree structure. With the help of the decision tree node range compression method based on bit operation, matching items are avoided one by one for large numerical interval values. Bit operation is used to simplify conditions, and hardware resource usage is reduced without loss of classification performance. For example, the condition "<8192" from branch 1 to branch 2 is processed. Because of its large value, measured in microseconds, the matching table requires a large number of entries. The right shift operation is innovatively adopted to shift all decision conditions and characteristic values right by 3 bits. Shifting one bit right is equivalent to dividing the value by 2 (integer division), and shifting 3 bits right is equivalent to dividing by 8, thereby reducing this condition to "[0,1024]", greatly reducing the size of the matching table.
4. The multi-service traffic group scheduling method for a P4 switch according to claim 1, characterized in that: The specific method of step (3) is as follows: (3.1) The ranking calculation based on the cumulative queuing delay includes the following steps: The ingress switch uses the five-tuple as the session basis for an independent network. When a new network session requests a custom header at the ingress switch, the controller queries the network topology to identify the expected transmission path and configures the flow table rules of the edge switch. For the intermediate switches, the state classification based on the cumulative delay ratio is carried out. According to the ratio of the cumulative queuing delay of the data packet to the ideal queuing delay, the data packet is divided into the idle state, the normal state and the congested state. When the ratio is less than 0.5, it is judged as the idle state and the priority of the current data packet is reduced by one level; when the ratio is between 0.5-1, it is judged as the normal state and no priority processing is performed; when the ratio is greater than 1, it is judged as the congested state and the priority of the data packet is increased by one level. At the same time, following the dynamic boundary threshold update rules, the boundary thresholds of the above status classification can be updated according to dynamic factors such as link load and queue status. The egress switch needs to determine the transmission status of the data packet and adjust the priority based on the above rules. At the same time, in order to ensure the normal forwarding of the data packet in the external network, the custom header needs to be deleted. (3.2) Packet scheduling is performed based on AD-PIFO. When a data packet enters the forwarding link of the switch, the priority of the data packet is determined by the information carried in the custom header. The algorithm starts from the FIFO queue with the highest priority, checks the boundary threshold of each queue in sequence, and compares it with the priority of the data packet until the data packet meets the queue entry conditions and pushes the data packet down. When the data packet is pushed down, it detects whether there is a reverse order phenomenon. If there is a reverse order, the algorithm follows the dynamic boundary threshold update rule and reduces the boundary threshold of the corresponding queue, lowering the priority upper limit to ensure that high-priority data packets are forwarded first. At the same time, the boundary threshold can be further dynamically adjusted according to the real-time status of the network to optimize the scheduling strategy.
Citation Information
Patent Citations
Network multi-dimensional data traffic simulation device based on composite two-dimensional Sketch
CN112134738A
Whole network flow measurement method and system in data center network and packet loss detection method
CN112822077A
Low-overhead continuous infrequent flow accurate identification architecture and method
CN118018440A
Sketch real-time network measurement method based on programmable hardware DPU, electronic equipment and medium
CN118467151A
Network overpoint measurement method and device based on multi-stage filtering Sketch
CN118590290A
Cited By
Satellite communication high-speed data transmission method and system
CN120415550A
UPF function acceleration method based on programmable hardware
CN120640318A
UPF Function Acceleration Method Based on Programmable Hardware
CN120640318B
Resource scheduling method and device, equipment and storage medium
CN120675954A
High-round-trip-delay flow detection method and system for high-speed network flow
CN121486252A