Flow table parallel matching method and system of programmable switching chip

By implementing the parallel matching method of flow tables on the programmable switching chip, including generating subtables, configuring matching control units and optimizing system configuration parameters, the performance bottlenecks and resource optimization problems of traditional flow table matching methods are solved, and efficient and accurate packet processing and resource utilization are achieved.

CN120196961AInactive Publication Date: 2025-06-24SHENZHEN SCODENO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510407251.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional flow table matching method has performance bottlenecks when facing large-scale flow tables, which cannot meet the needs of high-speed data packet processing, and lacks an effective parallel processing mechanism and a systematic resource optimization mechanism.

Method used

By implementing the flow table parallel matching method on the programmable switching chip, it includes determining the shard boundary points to generate a subtable, establishing the storage area of ​​the subtable and configuring a matching control unit, connecting the storage area with a cross switch structure, mathematical modeling and particle swarm optimization based on the subtable parallel processing architecture, and optimizing the system configuration parameters to achieve efficient parallel matching.

Benefits of technology

It realizes accurate matching result sorting, improves the accuracy of matching result, significantly improves the system's data processing throughput, optimizes resource utilization, reduces matching delay, and improves the system's load balancing capabilities and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196961A_ABST
    Figure CN120196961A_ABST
Patent Text Reader

Abstract

The invention relates to a flow table parallel matching method and system of a programmable switching chip. The method comprises the following steps: determining an input flow table based on the programmable switching chip as a fragment boundary point, and generating a plurality of sub-tables; respectively establishing storage areas of a plurality of sub-tables, configuring matching control units in the storage areas, and connecting the storage areas to obtain a sub-table parallel processing architecture; mathematical modeling is carried out based on the processing unit number constraint, the storage space constraint and the bandwidth constraint of the subtable parallel processing architecture, and a flow table parallel matching mathematical model is generated; performing particle swarm optimization on the system throughput parameter, the resource utilization rate parameter and the matching time delay parameter to obtain a system configuration parameter group; and obtaining an input data packet, distributing the input data packet to the corresponding sub-table unit according to the system configuration parameter group, and performing parallel matching operation in the sub-table unit to obtain a data packet forwarding strategy. According to the invention, accurate matching result sorting is realized, and the accuracy of the matching result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of flow table parallel matching, and particularly to a method and system for flow table parallel matching of a programmable switch chip. Background Art

[0002] Traditional flow table matching methods adopt a single serial processing architecture, which has serious performance bottlenecks when facing large-scale flow tables and cannot meet the requirements of high-speed packet processing.

[0003] During the flow table matching process of current programmable switch chips, due to the large scale and complex structure of the flow table, the matching delay increases and the resource utilization rate is low. The traditional flow table organization method fails to fully consider the similarity characteristics between flow table entries, resulting in repeated matching operations, wasting processing resources. At the same time, there is a lack of an effective parallel processing mechanism and the hardware parallel computing ability cannot be fully utilized. In addition, the existing flow table matching methods lack a systematic resource optimization mechanism and it is difficult to achieve optimal configuration under multi-dimensional constraints such as the number of processing units, storage space, and bandwidth. Summary of the Invention

[0004] The main object of the present invention is to provide a method and system for flow table parallel matching of a programmable switch chip, and the present invention realizes accurate sorting of matching results and improves the accuracy of matching results.

[0005] To achieve the above object, the present invention provides a method for flow table parallel matching of a programmable switch chip, including the following steps: Determine sharding boundary points based on the input flow table of the programmable switch chip and generate multiple sub-tables; Establish storage areas for the multiple sub-tables respectively, configure matching control units in the storage areas, and connect the storage areas using a crossbar switch structure to obtain a sub-table parallel processing architecture; Perform mathematical modeling based on the processing unit number constraint, storage space constraint, and bandwidth constraint of the sub-table parallel processing architecture to generate a flow table parallel matching mathematical model; Perform particle swarm optimization on the system throughput parameter, resource utilization parameter, and matching delay parameter of the flow table parallel matching mathematical model to obtain a system configuration parameter group; Obtain an input packet, and allocate the input packet to the corresponding sub-table unit according to the system configuration parameter group, and perform parallel matching operations in the sub-table unit to obtain a packet forwarding strategy.

[0006] The present invention also provides a system for flow table parallel matching of a programmable switch chip, including: A generation module, configured to determine sharding boundary points based on the input flow table of the programmable switch chip and generate multiple sub-tables; Configuration module, which is used to respectively establish storage areas for the multiple sub-tables, configure a matching control unit within the storage areas, and connect each of the storage areas using a crossbar switch structure to obtain a sub-table parallel processing architecture; Modeling module, which is used to perform mathematical modeling based on the processing unit number constraint, storage space constraint, and bandwidth constraint of the sub-table parallel processing architecture to generate a flow table parallel matching mathematical model; Optimization module, which is used to perform particle swarm optimization on the system throughput parameter, resource utilization parameter, and matching delay parameter of the flow table parallel matching mathematical model to obtain a system configuration parameter group; Operation module, which is used to obtain an input data packet, and distribute the input data packet to a corresponding sub-table unit according to the system configuration parameter group, and perform parallel matching operations within the sub-table unit to obtain a data packet forwarding policy.

[0007] In summary, the technical solution provided by the present invention realizes intelligent grouping of flow tables through flow table sharding based on cosine similarity and a hierarchical matching tree structure, reduces the matching complexity, and improves the matching efficiency. By adopting a sub-table parallel processing architecture and a global clock synchronization mechanism, parallel matching operations of multiple sub-tables are realized, significantly improving the data processing throughput of the system. The multi-objective particle swarm optimization algorithm is introduced to achieve the optimal configuration of system resources under multi-dimensional constraint conditions such as the number of processing units, storage space, and bandwidth. A parallel matching mechanism based on a hierarchical matching tree is designed, combined with a fast lookup tree structure and a cache organization strategy, to reduce the data access latency during the matching process. Through a dynamic data packet distribution mechanism and a multi-level parallel matching strategy, the load balancing ability of the system is improved, ensuring stable matching performance. The priority encoding and weight calculation mechanism are adopted to achieve accurate sorting of matching results, improving the accuracy of matching results. Through a hierarchical cache structure design and a result indexing mechanism, the data access efficiency is optimized, reducing the memory access overhead. Description of the Drawings

[0008] Figure 1 is a schematic diagram of the steps of the flow table parallel matching method of the programmable switch chip in an embodiment of the present invention; Figure 2 is a block diagram of the structure of the flow table parallel matching system of the programmable switch chip in an embodiment of the present invention.

[0009] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments

[0010] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0011] Referring to Figure 1 , this embodiment provides a method for parallel matching of flow tables of a programmable switching chip, including the following steps: S1. Determine the shard boundary points based on the input flow table of the programmable switching chip, and generate multiple sub-tables; Among them, an initial feature matrix is constructed for the input flow table according to the number of flow table entries \(n\) and the number of flow table fields \(m\). This matrix uses flow table entries as rows and corresponding fields as columns to form a matrix expression describing the flow table matching rules. On this basis, each element in the initial feature matrix is binarized. The values that can be used for exact matching are marked as 1, while the fields with wildcards or unspecified values are marked as 0 to generate a numerical feature matrix. Normalization operations are performed on each row vector of the numerical feature matrix to eliminate the dimensional differences of data in different dimensions, obtaining a standardized feature matrix. Based on the standardized feature matrix, the inner product of adjacent row vectors is calculated, and the vector inner product result is used as the cosine similarity between adjacent rows. The cosine similarity is used to measure the feature similarity between two flow table rules. The closer its value is to 1, the more similar the matching features between the rules are, and vice versa, the greater the difference. By traversing the entire feature matrix, a similarity sequence is calculated to reflect the similarity distribution among flow table entries. Based on the similarity sequence, in order to determine the sharding boundary points of the flow table, a threshold parameter is set, and the sequence positions with similarities less than the threshold are selected as candidate sharding points. The candidate sharding points are potential boundaries with large feature differences between different rules in the flow table. After obtaining the candidate sharding points, density clustering analysis is performed on these points to extract effective sharding boundary points. The core of density clustering is to analyze the distribution characteristics of candidate sharding points, especially the clustering interval, and select the points with distances greater than the preset distance from each other as effective sharding boundary points to form a sharding boundary set, representing the natural sharding boundaries of the flow table. According to the determined sharding boundary set above, the input flow table is segmented, and the flow table entries located between adjacent sharding boundary points are divided into the same sub-table to obtain an initial sub-table set. Through boundary point division, the input flow table is decomposed into several sub-tables, and each sub-table contains entries with relatively consistent features, reducing the complexity of a single sub-table and creating conditions for subsequent parallel processing. Feature similarity verification is performed on each sub-table. The average similarity between entries within each sub-table is calculated to evaluate the rule consistency within the sub-table. If the average similarity of a certain sub-table exceeds the preset similarity threshold, it is considered that the entries within the sub-table have sufficient similarity, and the sub-table is retained as the final result. Otherwise, the sub-tables that do not meet the threshold requirements are excluded. Through this verification process, it is ensured that the generated multiple sub-tables have high consistency in internal features.

[0012] S2. Storage areas for multiple sub-tables are established respectively. Matching control units are configured within the storage areas, and each storage area is connected using a crossbar switch structure to obtain a sub-table parallel processing architecture; Specifically, calculate the required storage capacity based on the number of entries and field lengths of each sub-table to generate a storage requirement matrix. The storage requirement matrix is the basis for the allocation of the sub-table storage area, indicating the capacity size required for each sub-table in the physical storage unit. Based on this matrix, divide the physical storage unit, allocate continuous storage space to each sub-table, and ensure that the storage areas of each sub-table do not overlap, obtaining an optimized sub-table storage area allocation scheme. Configure an independent read-write controller for each sub-table storage area. The read-write controller is used to manage the storage access operations of the sub-table and is directly connected to the memory through the storage address bus to form a storage access control structure. This structure enables the storage space of each sub-table to be independently read and written, avoiding access conflicts between sub-tables and improving the efficiency of data parallel processing. On the basis of the storage access control structure, configure a matching control unit for each storage area. The core of the matching control unit is to integrate a matching sequence control module and a priority arbitration module. The matching sequence control module is used to coordinate the execution order of data matching, while the priority arbitration module is used to resolve the priority conflicts between multiple matching requests to ensure the efficient execution of matching tasks. By integrating these modules into the matching control unit, a matching control architecture is constructed. On the basis of the matching control architecture, establish a data path between each matching control unit to form an interconnected topology structure. The interconnected topology structure realizes data interaction between sub-table storage areas through a flexible connection method and provides high-bandwidth communication support for parallel matching tasks. To ensure the synchronization of the system, design a global clock distribution network based on the interconnected topology structure. The core of the clock distribution network is to transmit the global clock signal to each storage area and control unit in a low-latency and high-precision manner to establish a unified synchronous clock system. Perform clock skew compensation on the synchronous clock system. Insert a delay matching unit on the target path of the clock signal, and balance the clock arrival time of each path by precisely controlling the delay amount to construct a balanced clock tree. Integrate the balanced clock tree with the interconnected topology structure to form a sub-table parallel processing architecture. This architecture ensures the efficiency and parallelism of flow table matching tasks through optimized storage allocation, independent storage access control, flexible matching control, and precise clock synchronization.

[0013] S3. Based on the processing unit number constraint, storage space constraint, and bandwidth constraint of the sub-table parallel processing architecture, perform mathematical modeling to generate a flow table parallel matching mathematical model; It should be noted that the resources in the parallel processing architecture of sub-tables are statistically analyzed, including the specific resource distribution of processing units, storage areas, and data paths. Through statistics, a resource parameter matrix is obtained, which reflects the core numbers of each processing unit in the system, the spatial distribution of the storage area, and the transmission bandwidth of each data path. Based on the resource parameter matrix, the difference between the total number of processor cores of each sub-table unit and the maximum number of cores of the chip is used as a constraint variable to obtain a processing resource constraint equation. Assuming that each sub-table unit is allocated a specific number of processor cores, the total number of these cores is compared with the maximum number of cores supported by the hardware, and the difference is taken as the constraint variable, thereby restricting the allocation range of processing unit resources to ensure that the core resource allocation within the chip meets the actual requirements of the system. At the same time, the storage area of the sub-tables is analyzed, the actual occupied space of each storage area is statistically calculated, compared with the physical storage capacity, the difference is calculated, and this difference is used as the core variable of the storage resource constraint equation. In this way, the occupation of storage resources by each sub-table is restricted to avoid exceeding the range of the physical storage capacity. The establishment of the storage resource constraint equation helps to optimize the allocation of storage space so that it can fully meet the requirements of the flow table matching task. Based on the interconnection topology structure of the sub-table parallel processing architecture, the actual transmission rate of each data path is compared with the link bandwidth, the difference is calculated, and this difference is used as the variable of the bandwidth resource constraint equation. The goal of this constraint condition is to limit the bandwidth occupation of the data path to ensure that during the parallel matching process, the data transmission rate does not exceed the maximum support range of the link, thereby avoiding data bottlenecks caused by insufficient bandwidth. The processing resource constraint equation, the storage resource constraint equation, and the bandwidth resource constraint equation are linearly combined to construct a system resource constraint model. Based on the system resource constraint model, the system throughput function, the resource utilization function, and the matching delay function are set as optimization objectives to form a multi-objective optimization equation set. Among them, the system throughput function reflects the number of matching tasks that can be completed per unit time, the resource utilization function measures the usage efficiency of the chip hardware resources, and the matching delay function describes the time required to complete a single flow table matching. By comprehensively considering these objective functions, the relationship between system performance and resource allocation is balanced. To achieve multi-objective optimization, weight allocation is performed on each objective function to establish a comprehensive evaluation function. By assigning different weights to different objectives, their importance in the optimization process is reflected. The comprehensive evaluation function is combined with the system resource constraint model to form a flow table parallel matching mathematical model.

[0014] S4. Perform particle swarm optimization on the system throughput parameter, resource utilization parameter, and matching delay parameter of the flow table parallel matching mathematical model to obtain a set of system configuration parameters; Specifically, normalize the system throughput parameter, resource utilization parameter, and matching delay parameter in the flow table parallel matching mathematical model to eliminate the influence of dimensions between different parameters, convert them into a unified numerical range, and generate an optimization objective vector, including the standardized expressions of maximizing system throughput, maximizing resource utilization, and minimizing matching delay. Set the initial position and velocity of the particle swarm according to the optimization objective vector. The initial position of the particle swarm is generated by a randomization method and mapped to the specific configuration values in the parameter space to form the initial particle swarm. The position of each particle in the parameter space represents a potential system configuration scheme, and its initial velocity determines the direction and step size of the particle's movement in the next generation. After the generation of the initial particle swarm is completed, calculate the fitness value of each particle. The fitness value is the core index of the optimization algorithm and is used to measure the quality of the configuration scheme corresponding to the particle. Through fitness calculation, match the position of each particle with the optimization objectives in the target vector to obtain the performance of the current particle. On the basis of fitness evaluation, compare the fitness value of each particle with its historical best value and the global best value, update the individual best position and the global best position, and obtain the initial value of the optimization iteration. Retain the optimal solutions found by the particles in the historical search process and combine the optimal solution information in the global scope to provide a reference for the update of the particle swarm in the next generation. Based on the initial value of the optimization iteration, calculate the new velocity vector of each particle. The update of the particle velocity is based on the velocity update formula, which introduces three key parameters: the inertia weight, the individual cognitive factor, and the social cognitive factor. Among them, the inertia weight is used to balance the continuity of the current velocity of the particle, the individual cognitive factor reflects the trend of the particle approaching its own historical best position, and the social cognitive factor represents the driving force for the particle to approach the global best position. By synthesizing these three factors, calculate the new velocity vector of the particle and update the position of the particle based on this vector. If the position of the particle exceeds the boundary of the search space, adopt a position limit strategy for correction to ensure that the position of the particle is always within a reasonable range. The updated particle swarm constitutes a new generation of particle swarm, and perform fitness calculation on it again. By evaluating the performance of the system configuration scheme corresponding to the position of the new generation of particles, optimize the solution quality of the particle swarm. On the basis of fitness calculation, use the non-dominated sorting method to stratify the particle swarm and classify the particles according to their non-dominated levels. The core of non-dominated sorting is to select those solutions that are not completely dominated by other particles in any one objective to form the Pareto optimal solution set. The Pareto optimal solution set represents a set of candidate configuration schemes that achieve a balance in multi-objective optimization. In order to screen and retain particles from the Pareto optimal solution set, perform crowding degree calculation. The crowding degree calculation is used to evaluate the distribution density of each particle in the solution space in the Pareto solution set, and the purpose is to preferentially retain those solutions located in the sparse region to improve the diversity of the solutions.After the crowding degree calculation is completed, the particles with the same non-dominated rank are sorted in descending order of crowding degree, and the particles with larger crowding degree are selected and retained for the next generation, so as to avoid too many redundancies in the solution set while maintaining the performance of the optimal solution. Through the iterative optimization of multiple generations of particle swarms, a set of system configuration parameter groups are obtained based on the Pareto optimal solution set and crowding degree screening.

[0015] S5. Obtain the input data packet, and allocate the input data packet to the corresponding sub-table unit according to the system configuration parameter group, and perform parallel matching operations within the sub-table unit to obtain the data packet forwarding strategy.

[0016] Among them, protocol analysis is performed on the input data packet. By identifying the protocol type of the data packet, splitting its header fields and data payload segments, the data packet is converted into a structured data form to obtain the feature sequence of the data packet. Among them, the header field group is used for subsequent matching operations, while the data payload segment remains unchanged for the final forwarding process. Based on the distribution rules of the system configuration parameter group, classification marking is performed on the feature sequence of the parsed data packet to generate an allocation priority table. The allocation priority table associates the feature sequence of the data packet with the processing capabilities of different sub-table units by analyzing the rule constraint conditions in the system configuration, and determines the priority allocation sub-table and its priority for each data packet feature sequence. In this way, a reasonable distribution path is formulated for each data packet to ensure the efficient utilization of resources and the balanced allocation of matching tasks. According to the allocation priority table, the feature sequence of the data packet is mapped to the storage address space of the corresponding sub-table unit to generate a sub-table allocation scheme. According to the resource configuration and bandwidth allocation of the sub-table unit, dynamic mapping is performed on the feature sequence of the data packet and it is allocated to a suitable storage area to complete the physical mapping of the data packet feature to the sub-table unit. After the sub-table allocation scheme is generated, the cache space of the sub-table unit is divided, and a matching buffer area and a result buffer area are established inside each sub-table unit. The functions of these buffer areas are to temporarily store the feature sequence of the data packet to be matched and store the intermediate and final results of the matching operation respectively. By writing the feature sequence of the data packet into the matching buffer area, an efficient buffer area organizational structure is constructed. After the buffer area organization is completed, in order to improve the efficiency of flow table matching, a fast lookup tree is constructed based on the buffer area organizational structure. The fast lookup tree is a hierarchical matching tree that constructs a hierarchical index structure for the flow table entries in the sub-table unit according to the coverage range of the matching fields. In this matching tree, each node represents a field matching range of the flow table entry and is assigned a value according to the weight of the matching field, so as to quickly filter out the optimal matching entry during the matching operation. Through the design of hierarchical indexing, the matching tree realizes the gradual precise matching from coarse-grained to fine-grained. The feature sequence of the data packet is input into the hierarchical matching tree, and field matching operations are performed simultaneously on each level of nodes. The parallel matching operation significantly accelerates the matching process through the hierarchical organizational structure between nodes. During the matching process, the matching status and weight value of each node are recorded in real time, and the node information of the successfully matched nodes is written into the result buffer area to generate a node matching result set containing all the matched node information. The forwarding policy of the data packet is generated based on the matching items in the node matching result set. By parsing the matching result, a forwarding instruction set including port selection instructions, quality of service marking instructions, and message modification instructions is generated. These instructions are combined to form a complete data packet forwarding policy, where the port selection instruction determines the physical forwarding path of the data packet, the quality of service marking instruction is used to identify the priority and service requirements of the data packet, and the message modification instruction modifies or rewrites the content of the data packet when necessary.

[0017] Group the fields of the data packet feature sequence. According to the hierarchical structure of the hierarchical matching tree, the feature fields are divided into a main matching field group and an auxiliary matching field group. The main matching field group contains the core fields with the most significant recognition meaning, and these fields have a higher priority in the matching rules, while the auxiliary matching field group contains secondary fields for refining the matching. Sort the main matching field group according to its matching priority to generate a field priority sequence. Input the main matching field group in the field priority sequence into the root node of the hierarchical matching tree. At the root node, perform an exact matching operation based on the preset matching rules. The matching rules of the root node are global and are used to quickly filter out the most likely matching paths. For the field combinations that match successfully, mark them as the first-level matching results to form an initial matching set. The initial matching set contains the preliminary matching relationship between the data packet features and the flow table rules. Based on the initial matching set, determine the nodes that need to enter the next-level matching, input the auxiliary matching field group into these nodes, and perform field comparison according to the matching rules of the nodes. During this process, record the matching status of each field in real time to generate intermediate matching results. The intermediate matching results reflect the matching progress of the current-level nodes. Assign weight coefficients to each field in the intermediate matching results. The distribution of the weight coefficients is based on the coverage range and matching accuracy of the fields. By dynamically adjusting these coefficients, calculate the comprehensive weight value of each node to generate a node weight matrix. Based on the node weight matrix, encode the weight values and matching status of the successfully matched nodes to generate a matching status sequence. The matching status sequence is an efficient data expression method used to parallelize the field matching operations in different-level nodes of the hierarchical matching tree. Under the parallel processing mechanism, the nodes at different levels perform matching simultaneously, thus significantly improving the matching efficiency. Through multi-level parallel matching, generate multi-level matching results containing complete field matching information. Based on the multi-level matching results, establish a result index table as a storage and query tool for the matching results. The result index table associates the status descriptors and weight information of the successfully matched nodes and writes them into the corresponding address space of the result buffer to generate a node matching result set. The node matching result set stores the status information and weight evaluation results of all successfully matched nodes in a highly structured form.

[0018] In one example, based on the input flow table of the programmable switching chip, it is determined as the shard boundary point and multiple sub-tables are generated, including: Divide the input flow table into a matrix according to the number of flow table entries n and the number of flow table fields m to obtain an initial feature matrix, and perform binarization processing on each element in the initial feature matrix, mark the specific matching value as 1, and mark the wildcard as 0 to obtain a numerical feature matrix; Normalize each row vector of the numerical feature matrix to obtain a standardized feature matrix, and calculate the inner product of adjacent row vectors based on the standardized feature matrix. Take the inner product as the cosine similarity between adjacent rows to obtain a similarity sequence; Set a threshold parameter based on the similarity sequence, take the similarity positions less than the threshold parameter as candidate shard points to obtain a candidate shard sequence, and perform density clustering analysis on the candidate shard sequence. Determine the candidate shard points with a clustering interval greater than the preset distance as effective shard boundary points to obtain a shard boundary set; Segment the input flow table according to the shard boundary set, and divide the flow table entries between adjacent shard boundary points into the same sub-table to obtain an initial sub-table set; Perform feature similarity verification on each sub-table in the initial sub-table set, and retain the sub-table when the average similarity of the entries inside the sub-table is greater than the preset similarity threshold to obtain multiple sub-tables.

[0019] In this example, for the input flow table, assume it contains entries, and each entry has fields. Construct an initial feature matrix , where the element of the matrix represents the matching value of the th field of the rd flow table rule. For the specific matching value, it is a specific numerical value or range, and the wildcard field represents any value. To facilitate subsequent calculations of the matrix, perform binarization on it. It is stipulated that when the field has a specific matching value, assign 1, that is when the field is a wildcard, assign 0, that is . Through this step, a numerical feature matrix is obtained, where is the binarized feature value. To make the feature vectors between different entries have a unified scale in the numerical space, normalize each row vector of the matrix . The normalization uses the Euclidean norm, that is, for each row , its normalized result is: ; where represents the normalized vector, is the Euclidean norm of the original vector. After normalization, a standardized feature matrix is obtained, where represents the standardized feature value. To evaluate the similarity between flow table entries, calculate the inner product of adjacent row vectors based on the standardized feature matrix For the Row vector and the row vector , and the formula for its inner product is: ; where represents the cosine similarity between the th row and the th row. Arrange the calculation results of the cosine similarities of all adjacent rows into a sequence , which is the similarity sequence. Identify the characteristic changes between flow table entries by analyzing the similarity sequence. Based on the similarity sequence , set a threshold parameter , mark the positions less than this threshold as candidate shard points, and generate a candidate shard sequence , where represents the index of the candidate shard point. To optimize the selection of shard points, perform density clustering analysis on the candidate shard sequence. Group the candidate shard points with close distances into one class, and define the distance between each candidate point as . If the distance between two candidate points is greater than the preset clustering interval , then they are considered to belong to different clusters. The candidate points with a clustering interval greater than obtained from the clustering analysis are determined as valid shard boundary points, forming a shard boundary set . According to the shard boundary set , perform segmentation processing on the input flow table. Divide the flow table entries between adjacent shard boundary points and into a sub-table, forming an initial sub-table set . Each sub-table contains a group of flow table entries with similar characteristics. To verify the quality of the initial sub-tables, calculate the feature similarity of the entries inside each sub-table . For any two flow table entries in the sub-table , calculate their similarity: ; After the calculation is completed, count the average similarity of all entry pairs inside the sub-table : ; where represents the number of entries in the sub-table . If AvgSim , where is the preset similarity threshold, then retain the sub-table . After this process, the finally retained sub-tables form an optimized sub-table set.

[0020] In one example, storage areas for multiple sub-tables are established respectively, a matching control unit is configured within the storage areas, and each storage area is connected using a crossbar switch structure to obtain a sub-table parallel processing architecture, including: Perform storage space allocation for multiple sub-tables, calculate the required storage capacity according to the number of entries and field lengths of each sub-table to obtain a storage requirement matrix, and divide the physical storage units based on the storage requirement matrix, allocating continuous storage space to each sub-table to obtain a sub-table storage area allocation scheme; Configure an independent read / write controller for each storage area in the sub-table storage area allocation scheme, and connect the read / write controller to the storage address bus to obtain a storage access control structure; Construct a matching control unit based on the storage access control structure, and integrate the matching sequence control module and the priority arbitration module into the matching control unit to obtain a matching control architecture; Establish a data path between the control units in the matching control architecture to obtain an interconnection topology structure, and design a clock distribution network based on the interconnection topology structure to distribute the global clock signal to each storage area and control unit to obtain a synchronous clock system; Perform clock skew compensation on the synchronous clock system, insert delay matching units on the target path to obtain a balanced clock tree, and integrate the balanced clock tree with the interconnection topology structure to obtain a sub-table parallel processing architecture.

[0021] In this example, the storage requirements for each sub-table are calculated according to the characteristics of multiple sub-tables. Assume the system contains sub-tables, and each sub-table contains flow table entries, and each entry has fields. Then the storage requirement for sub-table is: ; Where represents the storage capacity required for sub-table (in bytes), and represents the storage bit width occupied by each field (usually in bytes). Arrange the storage requirements of all sub-tables into a storage requirement matrix: ; This matrix provides the storage distribution information for all sub-tables. Based on the storage requirement matrix , divide the physical storage units . Assume the total physical storage capacity is , and the physical storage address space is divided into continuous blocks , which need to meet the following constraints: ; where represents the actual storage capacity allocated to the sub - table . Through linear programming or greedy algorithms, optimize the division of storage units and ensure that each sub - table has continuous storage space. After completing the storage division, map the storage area of each sub - table to the storage address space to form a sub - table storage area allocation scheme. Based on the sub - table storage area allocation scheme, configure an independent read - write controller for each storage area. The core function of the read - write controller is to manage the interaction between packet characteristics and flow - table storage space, including address resolution, data writing, and data reading operations. Each controller is connected to the storage unit through the storage address bus . Ensure that read - write operations can be executed efficiently. Assume that the bandwidth of the bus is , and the read - write rate requirement for a single storage area is , then the following conditions need to be met: ; After completing this design, form a storage access control structure, where the read - write operations of each sub - table storage area are independent through an independent controller. Build a matching control unit based on the storage access control structure. The matching control unit integrates a matching order control module and a priority arbitration module. The function of the matching order control module is to determine the matching order and target sub - table according to the characteristics of the input packet, and the priority arbitration module is used to handle conflicts of multiple matching requests. Assume that the maximum number of matching tasks supported by the system for parallel processing is , and the matching task queue of the input packet is , then the arbitration module needs to allocate resources according to the following rules: Priority order: . Based on the matching control architecture, to achieve the interconnection between sub - table units, establish a data path to form an interconnection topology structure. The design goal of the topology structure is to maximize data transmission efficiency and ensure that parallel matching between different sub - tables is not blocked due to bandwidth conflicts. Assume that there are data paths in the system, the bandwidth requirement for each path is , and the global bandwidth upper limit is , then it is necessary to meet: ; By optimizing the configuration of the data path, build an efficient interconnection topology structure. Design a global clock distribution network based on the interconnection topology structure to distribute clock signals to each storage area and control unit to form a synchronous clock system. Assume that the clock frequency is , and the propagation delay of the signal on different paths is , then the synchronization condition of the global clock signal is: ; where is the maximum allowable clock skew. By analyzing the path delay distribution, delay matching units are inserted for paths with larger delays , so that the delays of all paths are balanced, and a balanced clock tree is constructed. The balanced clock tree is integrated with the interconnection topology to form a complete sub-table parallel processing architecture.

[0022] In one example, based on the processing unit number constraint, storage space constraint, and bandwidth constraint of the sub-table parallel processing architecture, mathematical modeling is performed to generate a flow table parallel matching mathematical model, including: Statistical analysis is performed on the processing units in the sub-table parallel processing architecture to obtain a resource parameter matrix; Based on the resource parameter matrix, the difference between the total number of processor cores of each sub-table unit and the maximum number of cores of the chip is used as a constraint variable to obtain a processing resource constraint equation; For the sub-table storage area, the difference between the actual occupied space of each storage area and the physical storage capacity is used as a constraint variable to obtain a storage resource constraint equation; Based on the interconnection topology of the sub-table parallel processing architecture, the difference between the actual transmission rate of each data path and the link bandwidth is used as a constraint variable to obtain a bandwidth resource constraint equation; A linear combination of the processing resource constraint equation, storage resource constraint equation, and bandwidth resource constraint equation is performed to obtain a system resource constraint model; Based on the system resource constraint model, the system throughput function, resource utilization function, and matching delay function are set as optimization objectives to obtain a multi-objective optimization equation set; Weight distribution is performed on the objective functions in the multi-objective optimization equation set, a comprehensive evaluation function is established, and the comprehensive evaluation function is combined with the system resource constraint model to obtain a flow table parallel matching mathematical model.

[0023] In this example, statistics are performed on the resources in the sub-table parallel processing architecture, including the number of processor cores, storage space, and transmission rate of the data path occupied by each sub-table unit. Assume that the system contains sub-table units, and for the th sub-table, the number of processor cores is , the storage space is bytes, and the actual transmission rate of the data path is . Then a resource parameter matrix is constructed: ; where , , respectively represent the core number, storage occupancy, and transmission rate of the th sub-table unit. Based on the resource parameter matrix , define the constraint equations for each resource. For the processor core resource, assume the total number of cores in the chip is , and the total core demand of the sub-table unit is . Then the constraint equation for the processor core is: ; where represents the remaining number of cores. For the storage resource, assume the physical storage capacity of the system is , and the total storage demand of all sub-table units is . Then the constraint equation for the storage resource is: ; where represents the remaining storage capacity. For the bandwidth resource of the data path, assume the total bandwidth of the system link is , and the total data path rate of all sub-table units is . Then the constraint equation for the bandwidth resource is: ; where represents the remaining bandwidth. By linearly combining the processing resource constraint equation, the storage resource constraint equation, and the bandwidth resource constraint equation, the resource constraint model of the system is obtained: ; where , , are weight coefficients, respectively measuring the importance of processing resources, storage resources, and bandwidth resources in the constraint model. Based on the resource constraint model, define the optimization goal of the system performance. The system throughput represents the number of flow table matching tasks completed per unit time, and the calculation formula is: ; The resource utilization rate is an index measuring the resource usage efficiency of the system, defined as: ; The matching delay represents the average time consumed for a single flow table matching operation, and the calculation formula is: ; To achieve multi-objective optimization, the above three performance indicators are constructed into an optimization goal vector : ; where represents minimizing the matching delay. Based on the optimization goal vector And the resource constraint model , construct a multi-objective optimization equation set. To achieve a unified optimization goal, assign weights to each objective function , and define a comprehensive evaluation function: ; where 's value reflects the importance of throughput, resource utilization rate, and matching delay in optimization. Combine the comprehensive evaluation function with the system resource constraint model to obtain the final mathematical model for parallel matching of flow tables: .

[0024] In an example, perform particle swarm optimization on the system throughput parameter, resource utilization rate parameter, and matching delay parameter of the mathematical model for parallel matching of flow tables to obtain a system configuration parameter set, including: Normalize the system throughput parameter, resource utilization rate parameter, and matching delay parameter in the mathematical model for parallel matching of flow tables to obtain an optimization objective vector; Set the initial position and velocity of the particle swarm according to the optimization objective vector, map the initial position to specific configuration values in the parameter space to obtain the initial particle swarm; Calculate the fitness value of each particle based on the initial particle swarm, compare the fitness value with the historical best value and the global best value, and update the individual best position and the global best position to obtain the initial value of the optimization iteration; Perform speed update calculation on the initial value of the optimization iteration, substitute the inertia weight, individual cognitive factor, and social cognitive factor into the speed update formula to obtain a new particle speed vector, and update the particle position based on the new particle speed vector. Correct the particles that exceed the search space boundary using the position limit strategy to obtain a new generation of particle swarm; Perform fitness calculation on the new generation of particle swarm, use the non-dominated sorting method to stratify the particles, and select the Pareto optimal solution set to obtain a candidate configuration scheme set; Perform crowding degree calculation based on the candidate configuration scheme set, sort the particles with the same non-dominated rank in descending order of crowding degree, and select the particles with a larger crowding degree to be retained in the next generation to obtain the system configuration parameter set.

[0025] In this example, the system throughput parameter in the mathematical model for parallel matching of flow tables , resource utilization rate parameter and matching delay parameter are normalized to ensure the consistency of different objective dimensions. The normalized optimization objective vector is: ; where and represent the minimum and maximum values of the system throughput respectively, and is the range of resource utilization, and are the minimum and maximum values of the matching delay. Through normalization, all objective function values are mapped to the interval [0,1]. According to the optimization objective vector, the initial positions and velocities of the particle swarm are initialized. Assume there are particles in the particle swarm, and the dimension of each particle is (corresponding to the number of optimization variables), then the initial position of the -th particle is: ; where is the coordinate of the -th particle in the -th dimension, and the initial velocity is: ; where represents the velocity of the -th particle in the -th dimension. The initial position is mapped to the specific configuration values in the parameter space, including the number of cores, storage allocation, and bandwidth allocation, etc., to generate the initial particle swarm. Based on the initial particle swarm, the fitness value of each particle is calculated. The fitness function is defined based on the optimization objective vector as: ; where are the weights of throughput, resource utilization, and matching delay respectively, is the objective function value corresponding to the -th particle. Through fitness calculation, the performance of each particle is evaluated. The fitness value of the particle is compared with its historical best value and the global best value to update the individual best position and the global best position, and obtain the initial value of the optimization iteration. In the optimization iteration, the next velocity of the particle is calculated based on the velocity update formula: ; where is the inertia weight, and are the individual cognitive factor and the social cognitive factor, is a random number in the range [0,1], is the historical best position of the particle in the -th dimension, is the The global optimal position of the dimension. The particle position update formula is as follows: ; For particles that exceed the boundary of the search space, a position restriction strategy is adopted to correct them within the boundary range. After completing the position and velocity updates, the fitness of the new generation of particle swarms needs to be recalculated. Based on the fitness calculation, a non-dominated sorting method is used to stratify the particles. The core of non-dominated sorting is to divide the particles into different levels of non-dominated solution sets based on the Pareto dominance relationship. Suppose the objective vectors of particle and particle are and respectively. If each objective in is better than or equal to , and at least one objective is strictly better than dominates . After sorting, the particles in the Pareto optimal solution set form a candidate configuration scheme set. Based on the candidate configuration scheme set, the crowding degree of the particles is calculated. The crowding degree measures the sparsity of the distribution of Pareto solutions in the solution space. For particles in the same non-dominated level, calculate their neighborhood distances in each objective dimension: ; where is the crowding degree of particle , and are the neighbor solution values of particle in objective . Arrange the particles with the same non-dominated level in descending order of crowding degree, and preferentially select the particles with larger crowding degrees to be retained in the next generation to improve the diversity of solutions. After multiple iterations, the finally generated system configuration parameter group can achieve an optimal balance among throughput, resource utilization rate, and matching delay.

[0026] In an example, an input data packet is obtained, and the input data packet is allocated to the corresponding sub-table unit according to the system configuration parameter group, and parallel matching operations are performed within the sub-table unit to obtain a data packet forwarding policy, including: Perform protocol parsing on the input data packet, split the input data packet into a header field group and a data payload segment, and obtain a data packet feature sequence; Classify and label the data packet feature sequence according to the distribution rule of the system configuration parameter group to obtain an allocation priority table; Based on the allocation priority table, map the data packet feature sequence to the storage address space of the corresponding sub-table unit according to the allocation priority to obtain a sub-table allocation scheme; Divide the cache space for the sub-table allocation scheme, establish a matching buffer and a result buffer within the sub-table unit, write the data packet feature sequence into the matching buffer, and obtain the buffer organizational structure; Construct a fast lookup tree based on the buffer organizational structure, establish a hierarchical index structure for the flow table entries in the sub-table according to the coverage range of the matching fields, and assign a matching weight to each node to obtain a hierarchical matching tree; Input the data packet feature sequence into the hierarchical matching tree, perform field matching operations simultaneously at each level of nodes, record the matching status and weight values of each node, and write the node information of the successfully matched nodes into the result buffer to obtain a node matching result set; Generate port selection instructions, quality of service marking instructions, and packet modification instructions according to the matching items in the node matching result set, and combine them into a data packet forwarding policy.

[0027] In this example, the input data packet is protocol-analyzed, and the data packet is divided into a header field group and a data payload segment. The header field group contains keyword fields for matching, such as source address, destination address, protocol type, etc., while the data payload segment contains specific transmission data. Assume that the header of the input data packet contains fields, which are respectively , then the data packet feature sequence is represented as a vector: ; where is the specific value of the th header field. This feature sequence serves as the core descriptive information of the data packet during the matching process. Classify and mark the data packet feature sequence according to the distribution rules in the configuration parameter group. The distribution rules exist in the form of a priority rule table, which defines the mapping relationship between each data packet feature and the sub-table unit. Assume that the configuration parameter group defines sub-table units, and the priority rule for each sub-table unit is , then the classification and marking process of the feature sequence is described as: ; where is the priority mark of the feature sequence, and Match indicates whether it conforms to the rule of sub-table . Through this process, an allocation priority table is generated, which records the sub-table corresponding to each data packet feature sequence and its priority. Based on the allocation priority table, map the data packet feature sequence to the storage address space of the corresponding sub-table unit. Assume that the storage address space of each sub-table unit is , then the mapping relationship is: ; wherein is the storage address of the feature sequence . This step completes the mapping from the data packet feature to the storage unit, forming a sub-table allocation scheme. The cache space is partitioned for the sub-table allocation scheme. Each sub-table unit is internally divided into a matching buffer and a result buffer. The matching buffer is used to store the feature sequence to be matched, and the result buffer is used to store the matching result information. Assume that the cache capacity of the sub-table unit is , and the sizes of the matching buffer and the result buffer are and respectively. Then the cache partition satisfies: ; After writing the data packet feature sequence into the matching buffer, a complete cache buffer organizational structure is formed. Based on the cache buffer organizational structure, in order to improve the matching efficiency, a fast search tree is constructed. Within each sub-table unit, the flow table entries are hierarchically indexed according to the coverage range of the matching fields, forming a hierarchical matching tree. Assume that the sub-table contains flow table entries, and each entry has fields. Then the construction process of the hierarchical matching tree divides these entries into layers in turn according to the field priorities. In the matching tree, each node represents a field matching status, and the node weight is defined as: ; where is the weight value of node , which is used to measure the coverage range of the matching. After inputting the data packet feature sequence into the hierarchical matching tree, the field matching operations are performed simultaneously at each level of the nodes. For each node, the matching status and the weight value are recorded, and the node information with successful matching is written into the result buffer. Assume that the node set of the matching tree is , then the node matching result set is: ; Generate the forwarding policy for the data packet according to the matching items in the node matching result set. It includes port selection instructions, quality of service marking instructions, and packet modification instructions. Assume that the node with the highest matching weight in the node matching result set is , and its forwarding policy is: ; where PortSelect , QoSMark , PktModify Specific instructions representing port selection, quality of service marking, and packet modification respectively. These instructions are combined into a complete data packet forwarding policy and applied to the actual forwarding of data packets.

[0028] In one example, the data packet feature sequence is input into a hierarchical matching tree, and field matching operations are performed simultaneously at each level of nodes. The matching status and weight value of each node are recorded, and the information of the nodes with successful matches is written into the result buffer to obtain a node matching result set, including: Group the fields of the data packet feature sequence, divide the feature fields into a primary matching field group and a secondary matching field group according to the hierarchical structure of the hierarchical matching tree, and sort the primary matching field group to obtain a field priority sequence; Input the primary matching field group in the field priority sequence into the root node of the hierarchical matching tree, perform an exact matching operation based on the matching rules of the root node, mark the successfully matched field combinations as the first-level matching results, and obtain an initial matching set; Determine the next-level matching nodes based on the initial matching set, input the secondary matching field group into the next-level matching nodes, perform field comparison according to the node matching rules, and record the field matching status to obtain an intermediate matching result; Assign weight coefficients to each field in the intermediate matching result, adjust the weight coefficients according to the field coverage range and accuracy, calculate the comprehensive weight value of the node, and obtain a node weight matrix; Based on the node weight matrix, encode the weight value and matching status of each successfully matched node to obtain a matching status sequence, and perform parallel processing on the matching status sequence to perform field matching operations simultaneously at different levels of nodes to obtain a multi-level matching result; Establish a result index table based on the multi-level matching result, write the status descriptors and weight information of the successfully matched nodes into the corresponding address space of the result buffer to obtain a node matching result set.

[0029] In this example, for the feature sequence of the input data packet Perform field grouping. According to the hierarchical structure of the hierarchical matching tree, divide the feature fields into a primary matching field group and a secondary matching field group. The primary matching field group Contains fields important for initial screening, such as source address, destination address, etc.; the secondary matching field group Contains fields that play a role in refined matching, such as protocol type or port number. The grouping of fields is based on the priority of the matching rules and the coverage range of the fields. After grouping, sort the primary matching field group according to the importance of the fields to generate a field priority sequence: ; where is the number of primary matching fields, Indicates the th field sorted by priority. The field priority sequence is input into the root node of the hierarchical matching tree. At the root node, an exact matching operation is performed on the main matching field group according to the preset matching rules. Assume the matching rule contained in the root node is , then the matching process is expressed as: ; where and represent the matching items of the field and the rule respectively. The field combinations that match successfully are marked as the first-level matching results, generating the initial matching set . Based on the initial matching set, the next-level matching nodes of the hierarchical matching tree are determined, and the auxiliary matching field group is input into these nodes. In the next-level matching nodes, each auxiliary matching field is compared one by one according to the matching rule of the node, and the field matching status is recorded, where represents the matching result of the field . The matching process is expressed as: ; The successfully matched auxiliary fields are stored as the intermediate matching results . After generating the intermediate matching results, to optimize the weight distribution of the matching results, a weight coefficient is assigned to each field in the intermediate matching results. The calculation of the weight coefficient is based on the coverage range and the matching accuracy , and the specific formula is: ; where and are adjustment parameters used to balance the weight contributions of the coverage range and the matching accuracy. The calculated weight coefficients are used for the comprehensive evaluation of the node weights, generating the node weight matrix , where represents the comprehensive weight of the th node on the th field. Based on the node weight matrix , the weight value and the matching status of each successfully matched node are encoded to generate the matching status sequence. The matching status sequence is represented by , where , containing the weight and the matching status information of the node. The field matching operations are performed simultaneously on the nodes at different levels to obtain the multi-level matching results , where the matching results of each layer are independent of each other, and the operation efficiency is improved through parallel processing. Based on the multi-level matching results, a result index table is established. The role of the index table is to organize the node status descriptors and weight information that match successfully into structured data for subsequent query and operation. For each matching node, its status descriptor and weight are written into the corresponding address space of the result buffer to form a complete set of node matching results .

[0030] Referring to Figure 2 , this embodiment provides a flow table parallel matching system for a programmable switch chip, including: Generation module 1, which is used to determine the shard boundary points based on the input flow table of the programmable switch chip and generate multiple sub-tables; Configuration module 2, which is used to establish storage areas for multiple sub-tables respectively, configure matching control units in the storage areas, and connect each storage area using a crossbar structure to obtain a sub-table parallel processing architecture; Modeling module 3, which is used to perform mathematical modeling based on the processing unit number constraint, storage space constraint, and bandwidth constraint of the sub-table parallel processing architecture to generate a flow table parallel matching mathematical model; Optimization module 4, which is used to perform particle swarm optimization on the system throughput parameter, resource utilization parameter, and matching delay parameter of the flow table parallel matching mathematical model to obtain a system configuration parameter group; Operation module 5, which is used to obtain input data packets and allocate the input data packets to the corresponding sub-table units according to the system configuration parameter group, and perform parallel matching operations within the sub-table units to obtain a data packet forwarding policy.

[0031] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to that described in the above method embodiment, and details will not be elaborated here.

[0032] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, apparatus, article or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article or method including that element.

[0033] The above are only the preferred embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A flow table parallel matching method for a programmable switching chip, characterized in that: The following steps are involved: The input flow table based on the programmable switching chip is determined as the shard boundary point, and multiple sub-tables are generated; Establishing storage areas of the plurality of sub-tables respectively, configuring a matching control unit in the storage area, and connecting the storage areas by using a cross switch structure to obtain a sub-table parallel processing architecture; Perform mathematical modeling based on the processing unit quantity constraint, storage space constraint and bandwidth constraint of the sub-table parallel processing architecture to generate a flow table parallel matching mathematical model; Performing particle swarm optimization on the system throughput parameters, resource utilization parameters and matching delay parameters of the flow table parallel matching mathematical model to obtain a system configuration parameter group; An input data packet is acquired, and the input data packet is allocated to a corresponding sub-table unit according to the system configuration parameter group, and a parallel matching operation is performed in the sub-table unit to obtain a data packet forwarding strategy.

2. The flow table parallel matching method of a programmable switching chip according to claim 1 is characterized in that: The input flow table based on the programmable switching chip is determined as a shard boundary point, and multiple sub-tables are generated, including: The input flow table is matrix-divided according to the number of flow table entries n and the number of flow table fields m to obtain an initial feature matrix, and each element in the initial feature matrix is ​​binarized, and the specific matching value is marked as 1 and the wildcard is marked as 0 to obtain a numerical feature matrix; Normalizing each row vector of the digitized feature matrix to obtain a standardized feature matrix, and calculating the inner product of adjacent row vectors based on the standardized feature matrix, using the inner product as the cosine similarity between adjacent rows to obtain a similarity sequence; A threshold parameter is set based on the similarity sequence, and positions with similarity less than the threshold parameter are taken as candidate fragmentation points to obtain a candidate fragmentation sequence, and a density clustering analysis is performed on the candidate fragmentation sequence, and candidate fragmentation points with a clustering interval greater than a preset distance are determined as valid fragmentation boundary points to obtain a fragmentation boundary set; Segment the input flow table according to the shard boundary set, and divide the flow table entries between adjacent shard boundary points into the same sub-table to obtain an initial sub-table set; A feature similarity verification is performed on each sub-table in the initial sub-table set, and when the average similarity of entries in a sub-table is greater than a preset similarity threshold, the sub-table is retained to obtain a plurality of sub-tables.

3. The flow table parallel matching method of a programmable switching chip according to claim 2 is characterized in that: The storage areas of the plurality of sub-tables are established respectively, a matching control unit is configured in the storage area, and each storage area is connected by a cross switch structure to obtain a sub-table parallel processing architecture, including: Allocating storage space for the multiple sub-tables, calculating the required storage capacity according to the number of entries and field length of each sub-table to obtain a storage requirement matrix, and dividing the physical storage unit based on the storage requirement matrix to allocate continuous storage space to each sub-table to obtain a sub-table storage area allocation plan; configuring an independent read / write controller for each storage area in the sub-table storage area allocation scheme, connecting the read / write controller to a storage address bus, and obtaining a storage access control structure; Building a matching control unit based on the storage access control structure, integrating a matching order control module and a priority arbitration module into the matching control unit to obtain a matching control architecture; Establishing data paths between the control units in the matching control architecture to obtain an interconnection topology, and designing a clock distribution network based on the interconnection topology to distribute the global clock signal to each storage area and control unit to obtain a synchronous clock system; The synchronous clock system is compensated for clock deviation, and a delay matching unit is inserted on the target path to obtain a balanced clock tree, and the balanced clock tree is integrated with the interconnection topology structure to obtain a sub-table parallel processing architecture.

4. The flow table parallel matching method of a programmable switching chip according to claim 3 is characterized in that: The mathematical modeling based on the processing unit quantity constraint, storage space constraint and bandwidth constraint of the sub-table parallel processing architecture to generate a flow table parallel matching mathematical model includes: Performing resource statistics on the processing units in the sub-table parallel processing architecture to obtain a resource parameter matrix; Based on the resource parameter matrix, the difference between the total number of processor cores of each sub-table unit and the maximum number of cores of the chip is used as a constraint variable to obtain a processing resource constraint equation; For the sub-table storage area, the difference between the actual occupied space of each storage area and the physical storage capacity is used as a constraint variable to obtain a storage resource constraint equation; Based on the interconnection topology of the sub-table parallel processing architecture, the difference between the actual transmission rate of each data path and the link bandwidth is used as a constraint variable to obtain a bandwidth resource constraint equation; Linearly combining the processing resource constraint equation, the storage resource constraint equation, and the bandwidth resource constraint equation to obtain a system resource constraint model; Based on the system resource constraint model, the system throughput function, resource utilization function and matching delay function are set as optimization objectives to obtain a multi-objective optimization equation group; The objective functions in the multi-objective optimization equation group are weighted, a comprehensive evaluation function is established, and the comprehensive evaluation function is combined with the system resource constraint model to obtain a flow table parallel matching mathematical model.

5. The flow table parallel matching method of a programmable switching chip according to claim 4 is characterized in that: The particle swarm optimization is performed on the system throughput parameter, resource utilization parameter and matching delay parameter of the flow table parallel matching mathematical model to obtain a system configuration parameter group, including: Normalizing the system throughput parameter, resource utilization parameter and matching delay parameter in the flow table parallel matching mathematical model to obtain an optimized target vector; Setting the initial position and speed of the particle swarm according to the optimization target vector, mapping the initial position to a specific configuration value in the parameter space, and obtaining an initial particle swarm; Calculating the fitness value of each particle based on the initial particle group, comparing the fitness value with the historical optimal value and the global optimal value, updating the individual optimal position and the global optimal position, and obtaining the optimization iteration initial value; Performing speed update calculation on the optimization iteration initial value, substituting inertia weight, individual cognitive factor and social cognitive factor into the speed update formula to obtain a new particle velocity vector, and updating the particle position based on the new particle velocity vector, and correcting the particles beyond the search space boundary by adopting a position restriction strategy to obtain a new generation of particle swarm; Calculating the fitness of the new generation particle swarm, stratifying the particles using a non-dominated sorting method, selecting a Pareto optimal solution set, and obtaining a candidate configuration solution set; The congestion degree is calculated based on the candidate configuration scheme set, particles with the same non-dominated level are arranged in descending order of congestion degree, particles with larger congestion degree are selected and retained to the next generation, and a system configuration parameter group is obtained.

6. The method for parallel matching of flow tables of a programmable switching chip according to claim 5, characterized in that: The step of acquiring an input data packet, allocating the input data packet to a corresponding sub-table unit according to the system configuration parameter group, performing parallel matching operations in the sub-table unit, and obtaining a data packet forwarding strategy includes: Performing protocol parsing on an input data packet, dividing the input data packet into a header field group and a data payload segment, and obtaining a data packet feature sequence; Classify and mark the data packet feature sequence according to the distribution rule of the system configuration parameter group to obtain a distribution priority table; Based on the allocation priority table, the data packet feature sequence is mapped to the storage address space of the corresponding sub-table unit according to the allocation priority to obtain a sub-table allocation scheme; Dividing the sub-table allocation scheme into cache space, establishing a matching cache area and a result cache area in the sub-table unit, writing the data packet feature sequence into the matching cache area, and obtaining a cache area organization structure; A fast search tree is constructed based on the cache area organizational structure, a hierarchical index structure is established for the flow table entries in the sub-table according to the coverage of the matching field, a matching weight is assigned to each node, and a hierarchical matching tree is obtained; Input the data packet feature sequence into the hierarchical matching tree, perform field matching operations at nodes at all levels simultaneously, record the matching status and weight value of each node, write the node information of successful matching into the result buffer, and obtain a node matching result set; A port selection instruction, a service quality marking instruction and a message modification instruction are generated according to the matching items in the node matching result set, and are combined into a data packet forwarding strategy.

7. The flow table parallel matching method of a programmable switching chip according to claim 6, characterized in that: The data packet feature sequence is input into the hierarchical matching tree, field matching operations are performed simultaneously at nodes at each level, the matching status and weight value of each node are recorded, and the node information of the successfully matched node is written into the result buffer area to obtain a node matching result set, including: Field grouping is performed on the characteristic sequence of the data packet, and characteristic fields are divided into a primary matching field group and an auxiliary matching field group according to the hierarchical structure of the hierarchical matching tree, and priority sorting is performed on the primary matching field group to obtain a field priority sequence; Inputting the primary matching field group in the field priority sequence into the root node of the hierarchical matching tree, performing an exact matching operation based on the matching rule of the root node, marking the successfully matched field combination as the first-level matching result, and obtaining an initial matching set; Determine a next-level matching node based on the initial matching set, input the auxiliary matching field group into the next-level matching node, perform field comparison according to the node matching rule, record the field matching status, and obtain an intermediate matching result; Assigning a weight coefficient to each field of the intermediate matching result, adjusting the weight coefficient according to the field coverage and accuracy, calculating the comprehensive weight value of the node, and obtaining a node weight matrix; Based on the node weight matrix, the weight value and matching state of each successfully matched node are encoded to obtain a matching state sequence, and the matching state sequence is processed in parallel, and field matching operations are performed simultaneously on nodes at different levels to obtain multi-level matching results; A result index table is established based on the multi-level matching results, and the state descriptors and weight information of the successfully matched nodes are written into the corresponding address space of the result buffer area to obtain a node matching result set.

8. A flow table parallel matching system for a programmable switching chip, characterized in that: Steps for implementing the flow table parallel matching method of the programmable switching chip according to any one of claims 1 to 7, the flow table parallel matching system of the programmable switching chip comprising: A generation module, used for determining the shard boundary points based on the input flow table of the programmable switching chip and generating multiple sub-tables; A configuration module, used to respectively establish storage areas of the plurality of sub-tables, configure a matching control unit in the storage area, and connect the storage areas using a cross switch structure to obtain a sub-table parallel processing architecture; A modeling module, used to perform mathematical modeling based on the processing unit quantity constraint, storage space constraint and bandwidth constraint of the sub-table parallel processing architecture, and generate a flow table parallel matching mathematical model; An optimization module, used for performing particle swarm optimization on the system throughput parameters, resource utilization parameters and matching delay parameters of the flow table parallel matching mathematical model to obtain a system configuration parameter group; The operation module is used to obtain input data packets, and distribute the input data packets to corresponding sub-table units according to the system configuration parameter group, perform parallel matching operations in the sub-table units, and obtain a data packet forwarding strategy.

Citation Information

Cited By

  • Command execution method and device of unmanaged switch and storage medium

    CN120935015A