Fast and accurate packet classification method based on learning index
By optimizing the learning indexing method through distribution distance partitioning and cardinality-index representation, combined with a parallel multi-model architecture, the contradiction between model complexity and linear search range is resolved, achieving efficient and fast packet classification.
Patent Information
- Application Number
- CN202510773704.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-16
AI Technical Summary
Existing learning indexing methods have difficulty balancing model complexity and linear search range, resulting in high computational delay for high-complexity models and large search range for low-complexity models, which affects packet classification efficiency.
It uses distributed distance partitioning (DDP) and cardinality-index integer representation (BI), combined with a parallel multi-model architecture, to optimize model complexity and search range through fine-grained trend partitioning and integer modeling, and achieve parallel query.
Significantly reduce model complexity, compress linear search range, improve search speed and throughput, and meet real-time packet classification needs in high-concurrency scenarios.
Smart Images

Figure CN120654109A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer networks and machine learning, and in particular relates to a fast and accurate packet classification method based on learning indexes. Background Art
[0002] In computer networks, a fast and accurate packet classification rule matching mechanism is the core and foundation for efficient packet classification on large-scale network devices. The goal of packet classification is to identify the highest-priority rule that matches a packet within a preset rule set and then determine the action to be applied to the packet based on the content of that rule. However, with the continuous growth of network scale and traffic complexity, the high memory consumption and low query efficiency of traditional classification methods have become a major bottleneck for network speed improvements, severely impacting packet classification performance. Existing packet classification primarily uses two approaches to implement rule matching: hardware-based TCAM architectures and software-based architectures. In TCAM architectures, all packets are matched against rules using dedicated hardware. While this architecture offers fast query speeds, its high power consumption and high cost limit scalability, making it particularly vulnerable to large rule set scenarios. In contrast, software-based architectures (such as decision trees) utilize in-memory indexing structures and improve query efficiency through algorithmic optimization. This overcomes the scalability issues of TCAM architectures and has become a common solution for modern network devices, offering lower hardware costs and greater flexibility.
[0003] As the number and dimensionality of rules increase, software-based classification methods face increasing pressure on memory and computing power. To address this, various machine learning-based data structures have been proposed to optimize memory access and improve query efficiency. Learning indexes are an emerging solution that directly leverages data distribution characteristics for prediction, reducing memory consumption while avoiding the complex maintenance of traditional index structures. Existing work falls into two categories. The first, represented by NuevoMatch, uses the RMI (Recursive Model Index) model to accurately predict rule subscripts and compress the error range, but this model is complex, resulting in significant floating-point latency and resource overhead. The second, represented by NeuTree, replaces RMI with a binary model (RBMI). This model is less complex, but its coarse-grained segmentation leads to a large amount of linear scans. Both approaches struggle to strike a balance between model complexity and the size of the linear search range. Summary of the Invention
[0004] In response to the defects and shortcomings of the existing technologies, the present invention provides an efficient packet classification method based on dynamic trend perception and hardware-free optimization. Through the collaborative design of distribution distance division, integer modeling and parallel multi-model architecture, it systematically solves the core contradiction that has long existed in the field of learning indexing, that is, the inability to balance model complexity and linear search range.
[0005] Existing learning-based indexing methods (such as NuevoMatch and NeuTree) face a dilemma: high-complexity models narrow the search range but significantly increase computational latency; low-complexity models, while responsive, require larger linear scans. The EffiMatch framework proposed in this paper overcomes this bottleneck through three levels of innovation.
[0006] First, the Distributed Distance Partitioning (DDP) technique is introduced to achieve refined trend perception during the rule preprocessing stage. This method performs a least-squares linear fit on the key-value sequence of a subset of rules. It iteratively calculates the normalized geometric error and dynamically merges consecutive points with similar distribution trends to form highly cohesive segments. This adaptive segmentation mechanism significantly improves local model fitting accuracy, compressing the linear search range by over 26.84%, laying the foundation for subsequent lightweight modeling.
[0007] Secondly, the innovative radix-index integer representation (BI) completely eliminates the floating-point bottleneck. This method encodes the slope and intercept parameters of a piecewise linear model as integer pairs in base-exponent form, replacing floating-point multiplication and addition with exponent-aligned shift and accumulation operations. A special 0.5 offset rounding mechanism ensures integer approximation accuracy, improving model efficiency by four orders of magnitude on embedded devices without floating-point units while also reducing the risk of accuracy fluctuations.
[0008] Furthermore, a parallel multi-model collaborative architecture is constructed to unleash the potential of hardware. The rule set is divided into four subsets based on source IP / destination IP prefix length combinations (long-long, long-short, short-long, and short-short), with multiple lightweight segmentation models deployed within each subset. During queries, all segmentation models execute predictions synchronously and in parallel, combining associated compact rule groups (≤20 rules) to achieve millisecond-level retrieval. A two-level priority decision-making mechanism is employed—first selecting the highest-priority rule within the subset, then aggregating and comparing results to output the global optimal result. This ultimately achieves a sixfold increase in throughput, breaking through the throughput limitations of traditional serial queries.
[0009] Three major innovations form a closed-loop technical collaboration: the low-overlap segmented samples provided by DDP enable the BI integer model to maintain high precision at low complexity; the compact rule group design supports the full parallel execution of the BI model, releasing hardware computing power; ultimately, while reducing model complexity, a collaborative breakthrough is achieved, achieving a 26.84% reduction in search range, a 6-fold increase in throughput, and a 4-order-of-magnitude reduction in build time, providing a high-real-time packet classification solution for resource-constrained network devices.
[0010] The technical solution specifically adopted by the present invention to solve the technical problem is:
[0011] A fast and accurate packet classification method based on learned indexes:
[0012] The global rule set is divided into multiple rule subsets, and the key value sequence of each subset is dynamically divided into segments with similar distribution trends based on the linear fitting error;
[0013] The floating-point slope parameters and floating-point intercept parameters of each piecewise linear model are encoded as integer pairs in base-exponent form to construct a lightweight prediction model that supports pure integer operations.
[0014] When receiving a data packet to be classified, all segmentation models in each rule subset perform predictions in parallel and output the compact rule group index associated with the target segment;
[0015] Based on the index, the rule groups associated with each segment are retrieved in parallel, and the highest priority rule is selected from all rule group matching results to complete the data packet classification.
[0016] Furthermore, the key value sequence of each subset is dynamically divided into segments with similar distribution trends based on the linear fitting error, including:
[0017] Calculate the normalized geometric error between the fitted line and the data points for the key value-position pair;
[0018] The error threshold is set dynamically, and consecutive points with errors below the threshold are iteratively merged to form segments.
[0019] Furthermore, the global rule set is divided into a plurality of subsets according to the prefix lengths of the source IP address and the destination IP address;
[0020] Each subset selects the field with the lowest interval overlap rate as the main query dimension.
[0021] Furthermore, the rule subset is divided according to the combination relationship between the source IP address prefix length and the target IP address prefix length;
[0022] It is divided into the following four categories:
[0023] The source IP address has a long prefix and the destination IP address has a long prefix;
[0024] The source IP has a long prefix and the destination IP has a short prefix;
[0025] The source IP has a short prefix and the destination IP has a long prefix;
[0026] The source IP has a short prefix and the destination IP has a short prefix.
[0027] Furthermore, after encoding the floating-point slope parameter and the floating-point intercept parameter as integer pairs in base-exponent form, the calculation process includes:
[0028] According to the input value, the slope parameter and the intercept parameter are shifted and aligned according to the exponential size relationship and then accumulated;
[0029] The accumulated result is added with an offset of 0.5, shifted, and rounded to the output index value.
[0030] Furthermore, each segment-associated compact rule group contains no more than 20 rules.
[0031] Furthermore, the selection of the highest priority rule from all rule group matching results includes:
[0032] Select the local highest priority rule from the matching rules of each rule subset;
[0033] Compare the local results of all rule subsets and output the global highest priority rule.
[0034] And, a fast and accurate packet classification system based on learning index, comprising:
[0035] Rule set partitioning module: configured to partition the global rule set into multiple rule subsets;
[0036] Dynamic segmentation module: configured to dynamically divide the key value sequence of each subset into segments with similar distribution trends based on the linear fitting error;
[0037] Integer modeling module: This module is configured to encode the floating-point slope and intercept parameters of each piecewise linear model into integer pairs in base-exponent form, thus building a lightweight prediction model that supports pure integer operations.
[0038] Parallel query engine: When receiving a packet to be classified, it is configured to drive all segmentation models in each rule subset to perform predictions in parallel and output a compact rule group index associated with the target segment;
[0039] Rule decision module: It is configured to retrieve the rule groups associated with each segment in parallel based on the index, and select the highest priority rule from all rule group matching results to complete the data packet classification.
[0040] And, a computer device includes a memory, a processor and a computer program stored in the memory, and the processor implements the above method when executing the computer program.
[0041] A non-transitory computer-readable storage medium stores a computer program, which implements the method described above when executed by a processor.
[0042] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:
[0043] Breaking the trade-off between model complexity and search scope
[0044] By collaboratively applying distributed distance partitioning (DDP) and cardinality-index representation (BI), the linear search range is compressed while significantly reducing the model complexity, solving the contradiction between "high-complexity models have a small search range but high latency" and "low-complexity models have low latency but a large search range" in existing learning indexing methods, and achieving simultaneous optimization of two indicators.
[0045] Eliminate hardware deployment bottlenecks
[0046] The original BI integerization mechanism replaces floating-point calculations with pure integer shift and addition operations, allowing the model to run efficiently without relying on floating-point units, greatly improving the deployment capability on resource-constrained hardware such as embedded devices and network processors.
[0047] Unleashing the potential of parallel architecture
[0048] The parallel query architecture based on rule subset partitioning and segmentation model, combined with compact rule group design, significantly improves system throughput and meets the stringent requirements for real-time packet classification in high-concurrency scenarios.
[0049] Improve local modeling accuracy
[0050] DDP technology generates highly cohesive segments by dynamically merging regular points with similar distribution trends, enhancing the fitting ability of local linear models, reducing prediction errors, and providing a high-quality data foundation for subsequent lightweight modeling.
[0051] Ensuring global optimality of decision-making
[0052] The two-level priority aggregation mechanism (intra-subset optimization → cross-subset decision) ensures efficient parallel retrieval while outputting classification results with optimal matching accuracy and priority, avoiding rule omissions that may result from traditional parallel architectures.
[0053] Controlling memory and computational overhead
[0054] The compact rule group (≤20 rules) design significantly reduces the memory access volume of a single retrieval, while BI integer operations further reduce the number of computing instructions, jointly achieving efficient classification with low resource consumption.
[0055] Performance closed-loop optimization: DDP provides low-overlap segmented samples, BI implements low-complexity integer operations, and compact rule groups support high parallel efficiency. The three work together to simplify the model structure while achieving comprehensive improvements in throughput, accuracy, and hardware adaptability.
[0056] Enhanced system robustness: Through full-link innovation from data preprocessing (DDP), model lightweighting (BI), to execution architecture (parallelism), the risk of performance fluctuations under large-scale rule sets is significantly reduced, ensuring stable services in complex network environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0058] Figure 1 The figure shows the overall structure and flow chart of the embodiment of the present invention.
[0059] Figure 2 This is an example diagram of an implementation method of the Distribution-Distance Partitioning (DDP) algorithm according to an embodiment of the present invention.
[0060] Figure 3 This is a flow chart of an implementation method of the base-index representation (BI) of floating-point numbers according to an embodiment of the present invention.
[0061] Figure 4 This is a flowchart for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are given for detailed description:
[0063] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs.
[0064] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0065] Embodiments of the present invention provide a fast and accurate packet classification method based on learned indexes, specifically designed to balance the trade-off between the complexity of the recursive model index (RMI) model and the linear search range. This packet classification method based on learned indexes reduces memory overhead by using a lightweight recursive model index structure to limit the search range, followed by linear matching. However, this type of method faces a trade-off: while a complex recursive model index structure allows for a smaller search range, it also reduces search speed; whereas a simpler structure allows for faster searches but requires a larger scan range.
[0066] To this end, this paper proposes EffiMatch, a parallel multi-model search architecture designed to address the trade-off between the complexity of recursive model index structures and the linear search range in learned indexing systems. This architecture is used to support efficient packet classification in network devices. By designing the Distribution-Distance Partitioning (DDP) algorithm and the Base-Index Representation (BI) floating-point representation method, combined with parallel querying, it effectively reduces the linear search range, minimizes memory consumption, and improves search speed, providing a solid technical foundation for fast and accurate packet classification.
[0067] The present invention proposes two key designs:
[0068] 1) We designed a partitioning strategy based on distribution distance, which groups data points with similar trends into the same segment. Combined with parallel search, this method can effectively reduce the linear search range while maintaining high search speed.
[0069] 2) A finer-grained binarization method, the radix-index representation, is proposed. This method uses integers to approximate floating-point operations, further reducing the linear search range without increasing the complexity of the model.
[0070] The embodiment of the present invention reduces the linear search range by 26.84% when using a low-complexity RMI model, increases the search speed by up to 6 times, and reduces the construction time by up to 4 orders of magnitude compared to the most advanced learning indexing method.
[0071] It maps multi-field rules into key-value pairs, designs a distribution distance-based partitioning technology, divides data points with similar distribution trends into the same segment, adopts a cardinality-index representation, supports the use of integers to approximate floating-point operations, and ensures the combination of a parallel multi-model search architecture, thereby solving the trade-off between the recursive model index complexity and the linear search range.
[0072] like Figure 4 As shown, the implementation of this solution includes the following steps:
[0073] 1) Use distribution distance partitioning to divide the global rule set into multiple segments. This process involves iteratively fitting a straight line between key-value and position pairs. The straight line is approximated using the least squares method and points with similar distribution trends are grouped together according to a distribution distance threshold to reduce the linear search range.
[0074] 2) Using the cardinality-index representation, the parameters of the segmented line are encoded as a cardinality and an index ( This allows the use of pure integer operations to approximate floating-point operations, retaining high precision, thereby further reducing the search range without increasing the complexity of the model;
[0075] 3) Implementing a parallel multi-model search architecture where each segment is assigned a deployable linear model (called Range Binarized Model Index - RBIMI), which performs local segment fitting and supports independent parallel search;
[0076] 4) Within each segment, compact rule subsets (RuleSubsets), which typically contain fewer than twenty local rules, are pre-associated with key values to significantly narrow the scope of subsequent rule matching operations during the inference phase;
[0077] The distribution distance-based partitioning technique adaptively groups keys with similar local distribution trends into the same segment by evaluating the linear fitting error of key-position pairs point by point. It uses least squares regression to obtain a fitting line reflecting the linear trend and calculates the distribution distance to estimate the fit quality, thereby improving the accuracy of the linear approximation and reducing the linear search range.
[0078] The radix-index representation technique encodes floating-point numbers into an expressive pure integer format, where each weight is represented as This is done by using integers to perform calculations (e.g. for affine transformations ), while preserving the value precision, it eliminates the need for floating-point operations and ensures that the complexity of the model is not increased.
[0079] The parallel multi-model search architecture first partitions the entire rule set into a small number of rule subsets (RuleSubsets). Within each rule subset, a distribution-based distance partitioning technique is applied to generate local segments, and a corresponding model (RBIMI) is trained for each segment using a radix-index representation. During inference, the system serially traverses the rule subsets and uses M pre-trained RBIMI models for parallel search within each active subset to locate the target segment and retrieve matching rules from its associated rule group:
[0080] Within the parallel multi-model framework, the trend-based regular segmented distribution distance partitioning and the cardinality-index representation for accurate pure integer model evaluation are synergistically applied. Even if a lower complexity RMI model is used, the linear search range can be reduced, thereby significantly improving the search speed and model building time, effectively resolving the contradiction between the existing work's inability to balance the complexity of the RMI model and the linear search range.
[0081] Compared with existing learning indexing work, this method can improve the lookup speed by up to 6 times and reduce the construction time by up to 4 orders of magnitude while reducing the linear search range by 26.84% when using a lower complexity RMI model.
[0082] The EffiMatch framework proposed in this paper is suitable for implementing fast and accurate rule matching operations under large-scale rule sets in resource-constrained network devices. It effectively solves the trade-off between model complexity and linear search range size in existing learning-based classification methods. While ensuring search accuracy, it significantly reduces the error range and model construction overhead, effectively improving the system's classification throughput performance.
[0083] To achieve the above objectives, embodiments of the present invention utilize the Distribution-Distance Partitioning (DDP) algorithm and the Base-Index Representation (BI) method for floating-point numbers, combined with a multi-model parallel query approach, to simultaneously reduce model complexity, compress error intervals, and significantly improve query performance. First, the EffiMatch architecture partitions the global rule set into multiple rule subsets and, using algorithms such as Distribution-Distance Partitioning (DDP), performs fine-grained trend segmentation on each subset into multiple segments. This allows each rule subset to be independently queried on a single, specific dimension by a local, lightweight model, effectively adapting to the multi-model parallel query architecture. Second, EffiMatch utilizes the Base-Index Representation (BI) floating-point representation method to integerize floating-point parameters in the model, eliminating a large number of unnecessary floating-point operations during the query process. This significantly reduces the search error propagation problem caused by model simplification and ensures low model complexity and high hardware deployability. To further improve resource utilization efficiency, EffiMatch designs a compact rule group structure for each segment, achieving a high-speed packet classification system with controllable memory usage and efficient throughput.
[0084] In one embodiment of the present invention, the method includes the following steps: 1) performing distribution-distance partitioning (DDP): Key fields in a rule subset are first sorted into a monotonically increasing sequence. A greedy segmentation algorithm based on distribution trends is then used to partition key values with similar distribution characteristics into several linear segments using least squares fitting and error metrics. Rules within each segment can be individually modeled, effectively reducing search errors and increasing search parallelism, thereby adapting to a multi-model collaborative search architecture.
[0085] 2) Applying Base-Index Representation (BI): Quantize the linear model coefficients a and b trained in each segment to integer form, replacing floating-point multiplication and addition with integers. This representation supports fast lookups that rely entirely on integer shifts and additions, significantly reducing hardware computing overhead and improving model deployment efficiency.
[0086] In one embodiment of the present invention, the complete rule set is first divided into four rule subsets based on the prefix lengths of the source and destination IP addresses. Each subset selects a field as the query dimension to minimize the overlap of rule intervals along that dimension, thereby improving the accuracy of linear fitting and enhancing model query efficiency. This operation lays the foundation for the subsequent construction of a local index model, making the rule mapping within each subset more regular.
[0087] In one embodiment of the present invention, to further reduce sub-model complexity and ease the difficulty of fitting sub-models, the Distribution-Distance Partitioning (DDP) algorithm is used to perform fine-grained trend partitioning on the sequence. Based on least-squares fitting and geometric error evaluation, this algorithm iteratively extracts consecutive segments whose linear fitting errors are within a set threshold, thereby partitioning the rules into multiple linear segments based on distribution trends. Rules within each segment have similar key-value mapping patterns, facilitating the training of streamlined models and improving prediction accuracy. DDP partitioning not only enhances the model's local fitting capabilities but also significantly narrows the prediction range of each model, effectively reducing the burden of subsequent linear searches.
[0088] In one embodiment of the present invention, in order to further improve the computational efficiency and platform adaptability of the model, a Base-Index Representation (BI) integer encoding mechanism is introduced. This mechanism quantizes the floating-point parameters a and b in each model into integer form. and , reconstructing the original values as powers of 2, allowing floating-point operations to be replaced by integer addition, integer multiplication, and binary shifts. During inference, based on the input value x, the system constructs the linear prediction value using different shift strategies according to the exponential relationship between the two parameters. The final output is indexed using bit-aligned weighted accumulation and rounding. The BI mechanism significantly improves the execution speed of the index model and avoids the precision fluctuations and hardware deployment obstacles caused by floating-point numbers, providing a technical foundation for the rapid and stable operation of the model in actual systems.
[0089] In one embodiment of the present invention, a parallelized multi-model search architecture is designed to achieve high-throughput rule search. The system establishes multiple RBIMI sub-models corresponding to the DDP partition segments in each rule subset, and each model is responsible for predicting the approximate position of the input value in a linear segment. All RBIMI models run simultaneously during the query and return the index of the candidate rule group corresponding to each. Subsequently, the system retrieves specific rules from each rule group in parallel and finds the one with the highest matching priority as the matching result of the subset. Finally, the rule with the highest global priority is selected from the candidate results of the four subsets to complete the classification. This parallel structure fully utilizes the independence advantage brought by model segmentation, greatly improves the query throughput, and ensures the accuracy of classification and the controllability of response delay.
[0090] By introducing the Distribution-Distance Partitioning method, this paper achieves more refined distribution trend partitioning during the rule preprocessing phase, resulting in higher linear fitting accuracy for each model segment and significantly reducing the search range. Furthermore, using the Base-Index Representation floating-point representation method, the precision and magnitude of floating-point numbers are converted into shift, addition, and multiplication operations, eliminating a large number of unnecessary floating-point operations and effectively improving the model's operational efficiency and consistency on platforms without floating-point units. Furthermore, the parallel multi-model search architecture designed in this paper supports high-speed and accurate packet classification queries while maintaining extremely low model complexity. Experimental results show that EffiMatch reduces build time by 2 to 4 orders of magnitude compared to NuevoMatch and improves it by 63.56% compared to NeuTree for a 10k rule set. EffiMatch also improves search throughput by 30.71% compared to NeuTree and by 6 times compared to NuevoMatch, demonstrating its high efficiency and practicality for large rule sets.
[0091] A more specific application example is provided below to further demonstrate and illustrate the embodiment of the present invention:
[0092] The present invention provides a high-speed learning packet classification method based on distribution trend partitioning and integer modeling optimization. By designing and introducing the distribution distance partitioning algorithm Distribution-Distance Partitioning (DDP), the floating point base-index representation method Base-Index Representation, and the parallel multi-model search mechanism, it solves the bottlenecks of existing learning indexing methods in terms of model fitting accuracy, platform computing power adaptation, and large-scale search throughput. Figures 1 to 3 The specific implementation process of the present invention is described.
[0093] (1) Overall Overview
[0094] Please refer to Figure 1 The overall design structure of the method of the present invention is divided into two main stages: rule set partitioning, model structure establishment and parallel deployment. The figure shows the complete process from data preprocessing, index construction to query execution.
[0095] First, the system partitions the original rule set into four subsets based on the source and destination address prefix lengths. For each subset, the field with the minimum interval overlap ratio is selected as the primary query dimension. Subsequently, algorithms such as the distributed distance partitioning algorithm (DDP) are applied to each subset to sort key field values and evaluate fitting errors, generating multiple linear segments for local modeling (Phase 1). Next, in Phase 2, the system uses Base-Index Representation (BIR) to integerize the linear model parameters within the segment to improve the model's computational efficiency on the hardware platform. All segment models and rule groups are organized to construct a multi-model search path that can run in parallel, achieving high-speed, high-precision packet classification.
[0096] (2) Implementation of the Partial Distance Partitioning Algorithm
[0097] In this invention, applying the DDP algorithm is a key step in the first stage. The system extracts the primary query field from each subset, sorts its key values, and then inputs them into the segmentation module. This module locally models the value range using a least-squares linear fitting method and calculates the normalized geometric error between each sample point and the fitted line. If the error does not exceed a set threshold, the point is added to the current segment; otherwise, it is used as the starting point of a new segment and the segmentation continues.
[0098] Specific as Figure 2 As shown in the figure, the partitioning process continues until all key values are covered. The output list of segments constitutes the trend modeling structure for that subset. The similar distribution characteristics of the rules within each segment facilitate the construction of a concise linear model index. This method effectively reduces the fitting complexity of each segment model and narrows the linear search range during queries, providing a high-quality sample foundation for subsequent integer model training.
[0099] (3) Construction of floating-point base-exponent representation
[0100] like Figure 3 As shown in Figure 1, in order to reduce the cost of floating-point operations on the hardware platform, in the second stage of the present invention, the Base-Index Representation floating-point representation method is used to convert the model parameters into integers. Each linear model consists of two parameters a and b, which represent the slope and intercept in the learning index model respectively. The encoding mechanism converts it into and Two pairs of integers such that , .
[0101] In the query phase, the system calculates the value x based on the input The integer approximation of is achieved by shift-add with index alignment: if , then Left shift and Otherwise, Left shift and Align and sum. Finally, add a true value to the obtained base The offset is then shifted based on the sign and magnitude of the final exponent to restore the integer portion, rounding the value to obtain the index within the segment, which is used to search for the rule group. This representation simplifies the arithmetic logic in the inference path, significantly reduces prediction latency, and improves the model's efficiency on processors without floating-point units.
[0102] (4) Parallel multi-model structure deployment optimization
[0103] After completing the modeling and encoding of all segments, the system organizes the segment models under all subsets into a parallelizable search structure. The predicted results of each segment model are directly located in the corresponding rule group (RuleGroup). During a query, the four subsets are activated simultaneously, each calling the model of its own segment in parallel and returning a candidate matching index. The matching engine retrieves the highest-priority match from the rule group indicated by the predicted index as the subset output. The system then selects the highest-priority match from the four subset results as the final decision. This structure fully leverages the independence between model segments, enabling the entire query process to be fully parallelized, significantly improving the system's query throughput and response consistency.
[0104] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0105] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, performs the above-described method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0106] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
[0108] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of a fast and accurate packet classification method based on learning index under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A fast and accurate packet classification method based on learning index, characterized by: The global rule set is divided into multiple rule subsets, and the key value sequence of each subset is dynamically divided into segments with similar distribution trends based on the linear fitting error; The floating-point slope parameters and floating-point intercept parameters of each piecewise linear model are encoded as integer pairs in base-exponent form to construct a lightweight prediction model that supports pure integer operations. When receiving a data packet to be classified, all segmentation models in each rule subset perform predictions in parallel and output the compact rule group index associated with the target segment; Based on the index, the rule groups associated with each segment are retrieved in parallel, and the highest priority rule is selected from all rule group matching results to complete the data packet classification.
2. The fast and accurate packet classification method based on learning index according to claim 1, characterized in that: The key value sequence of each subset is dynamically divided into segments with similar distribution trends based on the linear fitting error, including: Calculate the normalized geometric error between the fitted line and the data points for the key value-position pair; The error threshold is set dynamically, and consecutive points with errors below the threshold are iteratively merged to form segments.
3. The fast and accurate packet classification method based on learning index according to claim 1, characterized in that: The global rule set is divided into multiple subsets according to the prefix lengths of the source IP address and the destination IP address; Each subset selects the field with the lowest interval overlap rate as the main query dimension.
4. The fast and accurate packet classification method based on learning index according to claim 3, characterized in that: The rule subset is divided according to the combination relationship between the source IP address prefix length and the target IP address prefix length; It is divided into the following four categories: The source IP address has a long prefix and the destination IP address has a long prefix; The source IP has a long prefix and the destination IP has a short prefix; The source IP has a short prefix and the destination IP has a long prefix; The source IP has a short prefix and the destination IP has a short prefix.
5. The fast and accurate packet classification method based on learning index according to claim 1, characterized in that: After encoding the floating-point slope and intercept parameters as integer pairs in base-exponent form, the operation consists of: According to the input value, the slope parameter and the intercept parameter are shifted and aligned according to the exponential size relationship and then accumulated; The accumulated result is added with an offset of 0.5, shifted, and rounded to the output index value.
6. The fast and accurate packet classification method based on learning index according to claim 1, characterized in that: Each segment-associated compact rule group contains no more than 20 rules.
7. The fast and accurate packet classification method based on learning index according to claim 1, characterized in that: The highest priority rule is selected from all rule group matching results, including: Select the local highest priority rule from the matching rules of each rule subset; Compare the local results of all rule subsets and output the global highest priority rule.
8. A fast and accurate packet classification system based on learning index, characterized by: include: Rule set partitioning module: configured to partition the global rule set into multiple rule subsets; Dynamic segmentation module: configured to dynamically divide the key value sequence of each subset into segments with similar distribution trends based on the linear fitting error; Integer modeling module: This module is configured to encode the floating-point slope and intercept parameters of each piecewise linear model into integer pairs in base-exponent form, thus building a lightweight prediction model that supports pure integer operations. Parallel query engine: When receiving a packet to be classified, it is configured to drive all segmentation models in each rule subset to perform predictions in parallel and output a compact rule group index associated with the target segment; Rule decision module: It is configured to retrieve the rule groups associated with each segment in parallel based on the index, and select the highest priority rule from all rule group matching results to complete the data packet classification.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
AWG waveform compression transmission algorithm for quantum computing
CN121860081A