Network processor and message processing method
By allowing the matching module and the execution module to share the cache in the network processor, the problem of frequent full write and read out of PS information in the prior art is solved, and the effect of saving bandwidth resources and reducing power consumption is achieved.
Patent Information
- Application Number
- CN202311760350.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
When existing network processors process packets, program status information needs to be frequently written and read out in full, resulting in waste of bandwidth resources and increased power consumption.
In the network processor, the matching module and the execution module share the same cache, reducing the number of read and write times of the full PS information. The PS information is first written into the shared cache. The matching module and the execution module only need to read out some information from the shared cache as needed, and then output it in full after the action is completed.
By reducing the number of read and write times of the full PS information, the bandwidth resources of the network processor are saved and power consumption is reduced.
Smart Images

Figure CN120179577A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a network processor and a packet processing method. Background Art
[0002] A network processor is an editable device used to handle various tasks in the communication field, such as packet processing, protocol analysis, route lookup, etc. There are mainly two architectures for existing network processors: one is the asynchronous Run-to-complete (RTC) architecture, and the other is the synchronous pipeline Pipeline architecture. Among them, the Pipeline architecture is a pipeline architecture composed of numerous pipeline nodes. When processing a packet, the packet will flow through the pipeline nodes on the pipeline in sequence, and each pipeline node is only responsible for completing a part of the entire processing process.
[0003] The pipeline nodes in a network processor usually have a Match-Action (MA) model structure, that is, each pipeline node includes a Match module and an Action module. The Match module is responsible for matching to obtain relevant information of the packet to be processed to determine the processing action of the Action module. The Action module then executes the processing action based on the matching result of the Match module. Usually, cache structures are respectively set in the Match module and the Action module. Among them, the cache in the Match module is used to absorb the query delay of input / output (I / O), and the cache in the Action module is used to absorb the processing delay.
[0004] Based on the above structure, when a pipeline node processes a packet, the program state (PS) information corresponding to the packet needs to be fully written into the cache of the Match module first. After the Match module finishes matching, it needs to be fully read out and fully written into the cache of the Action module. In this way, during the entire processing flow of the packet, the program state information PS will be copied back and forth between the cache of the Match module and the cache of the Action module, that is, continuously written and read out in full, which will seriously waste the bandwidth resources of the network processor and cause huge power consumption. Summary of the Invention
[0005] The embodiments of the present application provide a new network processor and a packet processing method. In this network processor, the Match module and the Action module share the same cache, that is, there is no need to set caches for the Match module and the Action module separately. In this way, when processing packets, the PS information is first written into the shared cache, and the Match module and the Action module only need to read out part of the information from the shared cache as needed. After the Match module and the Action module sequentially complete the matching action and the processing action, the shared cache outputs the updated PS information in full volume. In this way, the PS information does not need to be fully written and read back and forth between the two modules, greatly reducing the number of full-volume read and write operations of the PS information, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0006] In a first aspect, the embodiments of the present application provide a network processor in which the matching module and the execution module share a cache. Among them, the first PS information is first stored in the shared cache. When the packet to be processed is transmitted to the matching module, the matching module updates the first PS information stored in the cache according to the packet to be processed. In this way, the first PS information stored in the cache will become the second PS information including the characteristic information of the packet to be processed. Then, the execution module processes the second PS information in the cache.
[0007] In the above embodiment, after the full volume of the first PS information is stored in the shared cache, the matching module first reads out part of the information in the first PS information as needed, and then updates the first PS information based on the read part of the information. In this way, the first PS information stored in the shared cache is updated to the second PS information carrying the characteristic information corresponding to the packet to be processed. Then, the execution module executes the processing action based on the second PS information in the cache. That is, after the PS information is fully written into the shared cache, it is only fully read out after the matching module and the execution module sequentially complete the relevant actions. In this way, the PS information does not need to be copied back and forth between the matching module and the execution module, and the number of full-volume read and write operations of the PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0008] In an optional embodiment, after the matching module updates the first PS information, it needs to send an update completion notification to the execution module. After receiving the update completion notification, the execution module starts the processing process and begins to process the second PS information stored in the cache. Compared with the existing architecture, after the matching module completes the matching action, it does not need to send the full volume of the second PS information to the execution module, but only needs to send a trigger signal such as an update completion notification. This is because the matching module and the execution module share the cache, and both can directly access the shared cache. In this way, the number of full-volume read operations of the PS information is reduced, and the power consumption of the network processor is reduced.
[0009] In an alternative embodiment, when the matching module performs the matching action, it first reads part of the information as needed from the first PS information stored in the shared cache, and then generates an index based on the part of the information read. Then, based on the index, it queries the feature information of the packet to be processed from the external device, and then writes the obtained feature information into the first PS information in the shared cache. In this way, the first PS information in the shared cache is updated to the second PS information. In this embodiment, the matching module only reads or writes a small amount of information as needed, so that the amount of data read and written is greatly reduced, further reducing the power consumption of the network processor.
[0010] In an alternative embodiment, when the execution module performs the processing action, it first reads part of the information to be processed as needed from the second PS information in the shared cache, and the information to be processed includes the feature information of the packet to be processed previously written by the matching module. Then, the execution module processes the information to be processed and writes the processing result into the second PS information in the shared cache. In this way, the second PS information in the shared cache is updated to the third PS information. Similarly, the execution module also only reads or writes a small amount of information as needed, so that the amount of data read and written is reduced, further reducing the power consumption of the network processor.
[0011] In an alternative embodiment, when the execution module finishes the current round of actions, it can also send feedback information to the matching module to notify the matching module that the current round of update has been completed. After receiving the feedback information, the matching module can also perform the next round of update on the PS information. For example, it reads other parts of the information as needed, generates another index, then queries other feature information of the packet to be processed based on the index, and then updates the PS information based on the other feature information. It can be understood that when the execution module and the matching module perform multiple rounds of update on the PS information, the PS information only needs to be written and read in full in the shared cache once, so that the dynamic power consumption of the network processor is greatly reduced.
[0012] In an alternative embodiment, the network processor is composed of pipeline nodes with a pipeline architecture. Each of the pipeline nodes includes a matching module and an execution module. When the packet to be processed passes through each pipeline node, each pipeline node only performs a part of the entire processing process. Among them, only one cache can be set in each pipeline node, and the matching module and the execution module in the pipeline node share the cache.
[0013] Second aspect, an embodiment of the present application provides a packet processing method, which is applied to a network processor. Among them, the matching module and the execution module in the network processor share a cache. The packet processing method includes: First, the first PS information is stored in the shared cache first. When the packet to be processed is transmitted to the matching module, the matching module updates the first PS information stored in the cache according to the packet to be processed. In this way, the first PS information stored in the cache will become the second PS information including the feature information of the packet to be processed. Then, the execution module processes the second PS information in the cache.
[0014] In the above embodiment, after the full amount of the first PS information is stored in the shared cache, the matching module first reads some information from the first PS information as needed, and then updates the first PS information based on the read partial information. In this way, the first PS information stored in the shared cache is updated to the second PS information carrying the feature information corresponding to the packet to be processed. Then, the execution module performs a processing action based on the second PS information in the cache. That is, after the PS information is fully written into the shared cache, it is only fully read out after the matching module and the execution module have sequentially completed the relevant actions. In this way, the PS information does not need to be copied back and forth between the matching module and the execution module, and the read and write times of the full amount of PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0015] In an alternative embodiment, when the matching module updates the first PS information in the cache, it needs to first read some information from the first PS information stored in the shared cache as needed, and then generate an index based on the read partial information. Then, query the feature information of the packet to be processed from the external device based on the index, and write the obtained feature information into the first PS information in the shared cache. In this way, the first PS information in the shared cache is updated to the second PS information. In this embodiment, the matching module only needs to read or write a small amount of information as needed, so that the amount of data read and written will be greatly reduced, further reducing the power consumption of the network processor.
[0016] In an alternative embodiment, after the matching module updates the first PS information, it needs to send an update completion notification to the execution module. After receiving the update completion notification, the execution module starts the processing process and begins to process the second PS information stored in the cache. Compared with the existing architecture, after the matching module completes the matching action, it does not need to send the full amount of the second PS information to the execution module, but only needs to send a trigger signal such as an update completion notification. This is because the matching module and the execution module share a cache, and both can directly access the shared cache. In this way, the number of reads of the full amount of PS information is reduced, and the power consumption of the network processor is reduced.
[0017] In an alternative embodiment, when the execution module starts to execute the processing flow, it first reads part of the information to be processed as needed from the second PS information in the shared cache, and the information to be processed includes the feature information of the message to be processed previously written by the matching module. Then, the execution module processes the information to be processed and writes the processing result into the second PS information in the shared cache. In this way, the second PS information in the shared cache is updated to the third PS information. Similarly, the execution module only needs to read or write a small amount of information as needed. In this way, the amount of data read and written is reduced, further reducing the power consumption of the network processor.
[0018] In an alternative embodiment, when the execution module finishes the current round of actions, it can also send feedback information to the matching module to notify the matching module that the current round of update has been completed. After receiving the feedback information, the matching module can also perform the next round of update on the PS information. It can be understood that when the execution module and the matching module perform multiple rounds of update on the PS information, the PS information only needs to be written and read in full volume in the shared cache once. In this way, the dynamic power consumption of the network processor will be greatly reduced.
[0019] In a third aspect, an embodiment of the present application provides a chip, which includes: the network processor as described in any item of the first aspect; each module of the network processor is used to run code instructions to execute any one of the message processing methods provided in the second aspect.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program instruction is stored, and the computer program instruction is loaded and executed by a computer to implement any one of the message processing methods provided in the second aspect as described above.
[0021] For the technical effects brought by any implementation manner in the third aspect to the fourth aspect, reference can be made to the technical effects brought by the corresponding implementation manners in the first aspect and the second aspect, which will not be elaborated here.
[0022] In the embodiment of the present application, the matching module and the execution module of the network processor share a cache. In this way, when the network processor executes the processing flow, all the first PS information is first written into the shared cache. Then, the matching module reads part of the information from the first PS information as needed according to the packet to be processed, obtains the characteristic information of the packet to be processed based on this part of the data, and writes the characteristic information into the first PS information to obtain the second PS information. Next, the execution module also reads part of the information (including the characteristic information) in the second PS information as needed, then processes the characteristic information, and finally writes the processing result into the second PS information. It is not until the matching module and the execution module complete the entire processing flow that the final all PS information is read out of the cache. Therefore, the total number of read and write operations of the PS information will be greatly reduced, thus saving the bandwidth resources of the network processor. At the same time, the matching module and the execution module only need to read part of the data as needed, and the amount of data read or written by the two modules is also greatly reduced, which further saves the bandwidth resources of the network processor and reduces the power consumption of the network processor. Description of the Drawings
[0023] Figure 1 Schematic structural diagram of a Pipeline architecture network processor shown in the embodiment of the present application;
[0024] Figure 2A Schematic structural diagram of a pipeline node shown in the embodiment of the present application;
[0025] Figure 2B Schematic structural diagram of another pipeline node shown in the embodiment of the present application;
[0026] Figure 3A Schematic structural diagram of a new network processor provided by the embodiment of the present application;
[0027] Figure 3B Schematic structural diagram of another new network processor provided by the embodiment of the present application;
[0028] Figure 3C Schematic structural diagram of another new network processor provided by the embodiment of the present application;
[0029] Figure 4 Schematic structural diagram of a new pipeline node provided by the embodiment of the present application;
[0030] Figure 5 Schematic structural diagram of another new pipeline node provided by the embodiment of the present application;
[0031] Figure 6 Schematic flow chart of a packet processing method provided by the embodiment of the present application. Detailed Embodiment
[0032] The embodiments of the present application provide a new network processor and a packet processing method. In this network processor, the Match module and the Action module share the same cache, that is, there is no need to set caches for the Match module and the Action module separately. In this way, when processing packets, the PS information is first written into the shared cache, and the Match module and the Action module only need to read out part of the information from the shared cache as needed. After the Match module and the Action module execute the matching action and the processing action in sequence, the shared cache outputs the updated PS information in full volume. In this way, the PS information does not need to be written and read in full volume back and forth between the two modules, greatly reducing the number of read and write operations of the full-volume PS information, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0033] In the description of the present application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B can be singular or plural.
[0034] In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0035] In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit to be different.
[0036] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0037] It can be understood that the "embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that in the various embodiments of the present application, the magnitude of the serial numbers of the various processes does not mean the sequence of execution, and the execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0038] It can be understood that some optional features in the embodiments of the present application can, in some scenarios, be implemented independently without relying on other features, such as the current solution they are based on, to solve the corresponding technical problems and achieve the corresponding effects. In some scenarios, they can also be combined with other features according to requirements. Correspondingly, the devices given in the embodiments of the present application can also implement these features or functions accordingly, which will not be elaborated here.
[0039] In the present application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the present application, if there is no special specification and logical conflict between the various embodiments, the terms and / or descriptions between different embodiments are consistent and can be mutually referred to, and different embodiments can be combined to form new embodiments according to their internal logical relationships. The following embodiments of the present application do not constitute a limitation to the protection scope of the present application.
[0040] To facilitate the understanding of the technical solutions of the embodiments of the present application, a brief introduction to the related technologies of the embodiments of the present application is given below:
[0041] A network processor is an editable device used to handle various tasks in the communication field, such as packet processing, protocol analysis, route lookup, etc. There are mainly two architectures for existing network processors. One is the asynchronous Run-to-complete (RTC) architecture, and the other is the synchronous pipeline Pipeline architecture.
[0042] Among them, the Pipeline architecture is a pipeline architecture composed of numerous pipeline nodes. When processing packets, the packets will flow through the pipeline nodes on the pipeline in sequence, and each pipeline node is only responsible for completing a part of the entire processing process. The pipeline nodes on the pipeline usually adopt the Match-Action model structure. Among them, the editable Match-Action model can implement a protocol-independent network data forwarding processing process, including two parts: the matching logic and the action logic. The matching logic is implemented through a hybrid lookup table of static random access memory and ternary content addressable memory, as well as a combination of counters, traffic statistics tables, and general hash tables, while the action logic is implemented through a set of ALU standard boolean and arithmetic operation units, data packet header modification operations, and hash operations.
[0043] Based on the above description, Figure 1 This is a schematic structural diagram of a Pipeline architecture network processor shown in an embodiment of the present application. Among them, the network processor can be applied to devices that require high-bandwidth and high-performance data communication, such as routers, switches, and base stations. As Figure 1 shown, the network processor includes multiple serial pipeline nodes. When performing packet processing, the packets pass through multiple pipeline nodes in sequence. Among them, the packet first passes through the parser Parser, and the Parser parses the data packet according to the data packet header declared by the programmer to generate an independent data packet header and intermediate information. Then the parsed packet passes through each pipeline node in sequence, and each pipeline node performs a series of packet processing processes in sequence, such as matching the flow table according to the parsed header and meta-information, and performing operations such as modifying, adding, and deleting the header; and each pipeline node only processes a part. In each pipeline node, the Match module is responsible for matching and looking up the table from an external memory to determine the processing action of the Action module. The Action module then executes the processing action based on the matching result of the Match module. Finally, the processed packet is serialized again through the deparser Deparser and then enters the queue Queues and waits for output.
[0044] And Figure 2A This is a schematic structural diagram of a pipeline node shown in an embodiment of the present application. As Figure 2A shown, the pipeline node includes a matching module and an execution module, and the matching module and the execution module are respectively provided with a buffer Buffer for storing the program state PS information of the packet. Among them, the PS information is the state information of the program running, including the Ethernet header, IP header, TCP header, etc. of the packet, the global variables and local variables of the program running, the table entries returned by the IO, and the instruction addresses related to the program control flow, etc.
[0045] The Buffer1 set by the matching module is usually implemented with lower-cost memory to absorb I / O latency. The Buffer2 set by the execution module is usually implemented with registers to absorb processing latency. Based on this structure, when the pipeline node performs the packet processing flow, it first copies the entire PS information to Buffer1 of the matching module. After the matching module finishes the matching operation, it reads out all the PS information in Buffer1 and copies it in full to Buffer2 of the execution module, and then the execution module performs the execution operation based on the PS information in Buffer2.
[0046] As can be seen from the above figure, when the matching module and the execution module perform a matching-processing flow once, all the PS information needs to be copied from Buffer1 to Buffer2. When the pipeline node is loopback, that is, when it needs to perform multiple matching-processing flows, all the PS information will be passed back and forth between the two Buffers of this pipeline node. As Figure 2B shown, assuming that a packet enters the pipeline node and loops back 3 times within the pipeline node, that is, the pipeline node will perform four matching-processing flows for this packet. At this time, the number of memory read and write operations of the two Buffers is 4 times respectively, that is, the two Buffers need to read and write a total of 8 times the PS. Generally, the full PS in the network processor is usually on the order of 512 Bytes, and 8 times the PS read and write means that the processor in this architecture needs to read and write 4096B of memory data. This requires that the bus width of the network processor be large enough and the area of the chip be large enough. At the same time, a large amount of data read from the memory will lead to a high cache miss rate, and the dynamic power consumption of the entire network processor will increase greatly, seriously affecting the performance of the network processor.
[0047] To solve the above problems, the embodiments of the present application provide a new network processor architecture, and propose a packet processing method based on this architecture. In the new network processor, the matching module and the execution module share the same cache, that is, there is no need to set caches for the matching module and the execution module separately. In this way, when processing packets, the PS information is first written into the shared cache, and the matching module and the execution module only need to read out part of the information as needed from the shared cache. After the matching module and the execution module complete the matching operation and the processing operation in sequence, the shared cache outputs all the updated PS information. In this way, the PS information does not need to be fully written and read back and forth between the two modules, greatly reducing the number of read and write operations of all the PS information, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0048] The architecture of the network processor will be introduced first below:
[0049] Figure 3A This is a schematic structural diagram of a new network processor provided by the embodiments of the present application. AsFigure 3A As shown, the network processor is composed of multiple pipeline nodes. Each pipeline node includes a matching module and an execution module, and the matching module and the execution module share a cache. As Figure 3A shown, the pipeline nodes are serially connected. After the packet to be processed is output from the previous pipeline node, it is input into the next pipeline node. Each pipeline node includes a cache, and all the matching modules and execution modules in a pipeline node share this cache.
[0050] Exemplarily, in the network processor, multiple serially connected pipeline nodes can also share the same cache. As Figure 3B shown, multiple pipeline nodes are serially connected. Each pipeline node only includes a matching module and an execution module, and no cache is set. A shared cache is set outside the pipeline nodes, and the matching modules and execution modules of multiple pipeline nodes share the same cache.
[0051] Exemplarily, multiple pipeline nodes in the network processor can also be connected in parallel to form a parallel pipeline architecture. As Figure 3C shown, multiple pipeline nodes are connected in parallel to form multiple parallel branches. In this way, multiple traffic flows can be processed simultaneously. Among them, the pipeline nodes on one branch have a loopback function. For example, node 1 can loop back the packet to be processed, that is, it can execute multiple rounds of matching - processing processes on the packet to be processed and then output the processed packet. In the parallel architecture, the matching module and the execution module of each pipeline node can share the same cache. At the same time, the matching modules and execution modules of multiple pipeline nodes on one branch can also share the same cache, which is not limited here.
[0052] Based on the above description, the processing processes of the matching module and the execution module for a pipeline node are described in detail below:
[0053] Figure 4 It is a schematic structural diagram of a new pipeline node provided by an embodiment of the present application. As Figure 4 shown, the pipeline node includes a matching module and an execution module, and the matching module and the execution module share a Program State Buffer (PSB). The PSB is used to store the PS information of the packet, such as the Ethernet header, IP header, TCP header, etc. of the packet, the global variables and local variables of the program operation, the table entries returned by the IO, and the instruction addresses related to the program control flow, etc.
[0054] First, the first PS information of the message needs to be fully written into the PSB. In this way, when the message to be processed is transmitted to the pipeline node, the matching module reads part of the information as needed, obtains the characteristic information of the message to be processed from the external I / O device based on the partial data, and then writes the characteristic information back into the first PS information of the PSB to complete the update of the first PS information. Finally, the execution module reads part of the information as needed, which includes the newly written characteristic information, and then performs data processing based on the characteristic information to complete the execution action.
[0055] Among them, the process of the matching module reading information is that the matching module first obtains index information from the first PS information stored in the PSB according to the message to be processed, and then generates an index based on the index information. Then, based on the index, the characteristic information corresponding to the message to be processed is obtained from the external I / O device, and the characteristic information is written back into the first PS information of the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.
[0056] When the matching module completes the update of the first PS information, for example, the matching module needs to send an update completion notification to the execution module to notify the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts to process the second PS information in the PSB.
[0057] After the execution module starts the processing process, it also reads information as needed. Specifically, the execution module first reads part of the information to be processed as needed from the second PS information in the PSB, and the information to be processed includes the characteristic information of the message to be processed previously written by the matching module. Then, the execution module processes the information to be processed and writes the processing result into the second PS information in the PSB. In this way, the second PS information in the PSB is updated to the third PS information.
[0058] It can be understood that pipeline nodes can be serially connected to other pipeline nodes. When a pipeline node performs a table lookup operation, that is, executes a round of matching - processing process, after the execution module finishes the processing action, the PSB outputs all the PS information it stores to the next pipeline node.
[0059] The following uses a specific example to illustrate a round of matching - processing process of the pipeline node:
[0060] Assume that the pipeline node is used to modify the IPv4 routing table. Then, the PSB of the pipeline first writes the PS information of the packet. Then, according to the IPv4 routing table program entry carried in the PS, it triggers to enter the matching module to construct the index Key. Exemplarily, the matching module reads out the IPv4 destination IP (32-bit) required to construct the Key and the VPN information (16-bit) to which the packet belongs from the PS of the PSB. Then, based on the two, it constructs the Key of the routing table: VPN + destination IP. Next, the matching module sends the constructed key to the I / O request interface, sets the PS storage location of the I / O return result and the corresponding routing table processing program entry (for subsequent execution modules to use), and then ends.
[0061] After obtaining the corresponding routing table by looking up the routing table through an external I / O device, it will be returned to the matching module through the I / O response interface. After receiving the query result, the matching module will write the obtained routing table to the corresponding location of the PS information according to the PS storage location set in the above operation. After completion, it will send a trigger message to the execution module, and the execution module will execute the subsequent processing flow.
[0062] The execution module starts the processing program for the routing table. First, it reads the PS information required by this program from the PSB (including the routing table written by the matching module, the packet header, and some other information, etc.). Then, it obtains the program through the routing table processing program entry and executes it. For example, it modifies the routing table entries of the routing table, adds new routing table entries to the routing table, etc. After the execution module finishes executing the program, it needs to send the execution result to the local CPU and set the subsequent program processing entry. At the same time, it needs to write the program processing result (such as the modified routing table) back to the PS information in the PSB. Then, if the entry of the subsequent program is not at this node, the modified PS information will be read out in full from the PSB and written into the PSB of the next pipeline node.
[0063] In Figure 4 it, the black solid line is used to represent the data plane of the pipeline node. It can be seen that when the full first PS information is transferred from the previous pipeline node to this pipeline node, it is first written into the PSB in full. After the matching module and the execution module complete their actions, the updated PS information is read out from the PSB in full. In this way, after the pipeline node completes a round of matching-execution actions, the full PS information does not need to be copied back and forth between the matching module and the execution module, and will only be read and written once. In this way, the number of read and write operations of the full PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0064] Meanwhile, when the matching module reads data from the PSB, it only needs to read out the index information as needed. The average length of this index information is generally about 8 bits, and the amount of data it reads is very small. Also, when the matching module writes data to the PSB, it only needs to write the small amount of feature information (generally 32 bits) it obtains into the original first PS information. Similarly, the execution module also only needs to read and write a small part of the data as needed. Generally, the average length of the data read by the execution module is about 128 bits, and the processing result written is generally 64 bits. Neither the matching module nor the execution module needs to read and write all the PS information, which greatly reduces the amount of data read and written in memory and further reduces the power consumption of the network processor.
[0065] The black dashed line is used to represent the control plane of the pipeline node. Among them, after the matching module finishes the matching action, it only needs to send a completion notice to the execution module, and there is no need to send all the second PS information to the execution module. This message is a trigger signal with a very small amount of data. This is because the matching module and the execution module share a cache, and both can directly access the shared cache. Therefore, it is possible to avoid each module reading and writing all the PS information, thereby reducing the power consumption of the network processor.
[0066] It can be seen from the above description that if the pipeline nodes are serially connected to other pipeline nodes, each pipeline node only executes the matching - processing process once. Then, compared with the existing pipeline node architecture where both the matching module and the execution module are provided with caches, the pipeline node provided by the embodiment of the present application can save one full - volume reading and writing of the PS. That is, in the existing architecture, the cache of the matching module needs to read and write all the PS information once, and the cache of the execution module needs to read and write all the PS information once. However, for the pipeline node of the embodiment of the present application, the PSB only needs to read and write all the PS information once.
[0067] Figure 5 This is a schematic structural diagram of another pipeline node provided by the embodiment of the present application. As Figure 5 shown, the difference between this pipeline node and the pipeline node shown in Figure 4 is that the pipeline node in the embodiment of the present application is loop - back type, that is, a pipeline node can execute multiple rounds of matching - processing processes.
[0068] The process of this pipeline node executing multiple rounds of matching - processing processes is introduced as follows:
[0069] First, the first PS information of the message is written in full to the PSB. In this way, when the message to be processed is transmitted to the pipeline node, the matching module first obtains index information from the first PS information stored in the PSB according to the message to be processed, and then generates an index based on the index information. Then, the characteristic information corresponding to the message to be processed is obtained from the external I / O device based on the index, and then the characteristic information is written back to the first PS information of the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.
[0070] After the matching module completes the update of the first PS information, exemplarily, the matching module needs to send a completion notice to the execution module to notify the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts to process the second PS information in the PSB.
[0071] After the execution module starts the processing process, it also reads information as needed. Specifically, the execution module first reads some information to be processed as needed from the second PS information in the PSB, and the information to be processed includes the characteristic information of the message to be processed previously written by the matching module. Then, the execution module processes the information to be processed and writes the processing result to the second PS information in the PSB. In this way, the second PS information in the PSB is updated to the third PS information.
[0072] Next, the execution module needs to send feedback information to the matching module. After receiving the feedback information, the matching module starts to execute the next round of matching - processing process. Similarly, the matching module obtains other index information from the third PS information stored in the PSB according to the message to be processed, and then generates an index based on the index information. Then, other characteristic information corresponding to the message to be processed is obtained from the external I / O device based on the index, and then the other characteristic information is written back to the third PS information of the PSB. At this time, the third PS information stored in the PSB is updated to the fourth PS information. Then, the matching module still needs to send a completion notice to the execution module to notify the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts to process the updated fourth PS information in the PSB, and so on, which will not be elaborated here.
[0073] The following uses a specific example to illustrate the two - round matching - processing process of the pipeline node:
[0074] Suppose the pipeline node is used to modify the IPv4 routing table and query the Address Resolution Protocol (ARP). Then, the PSB of the pipeline first writes the PS information of the packet. Then, according to the IPv4 routing table program entry carried in the PS, it triggers to enter the matching module to construct an index Key. Exemplarily, the matching module reads out the destination IP of the IPv4 header (32 bits) and the VPN information (16 bits) of the packet to which the packet belongs, which are required to construct the Key, from the PS of the PSB. Then, it constructs the Key of the routing table based on the two: VPN + destination IP. Next, the matching module sends the constructed key to the I / O request interface, sets the PS storage location of the I / O return result and the corresponding routing table processing program entry (for subsequent execution modules to use), and then ends.
[0075] After obtaining the corresponding routing table by looking up the routing table through an external I / O device, it will be returned to the matching module through the I / O response interface. After receiving the query result, the matching module will write the obtained routing table to the corresponding location of the PS according to the PS storage location set in the above operation. After completion, it will send a trigger message to the execution module, and the execution module will execute the subsequent processing flow.
[0076] The execution module starts the processing program for the routing table. First, it reads the PS information required by this program (including the routing table written by the matching module, the packet header, and some other information, etc.) from the PSB, and then obtains the program through the routing table processing program entry and executes it. For example, it modifies the routing table entries of the routing table, adds new routing table entries to the routing table, etc. After the execution module finishes executing the program, it needs to send the execution result to the local CPU, and needs to set the subsequent program processing entry. At the same time, it needs to write the program processing result (such as the modified routing table) back to the PS in the PSB. Suppose the execution result of the execution module is that it needs to continue to look up the ARP. Then, it first sets the subsequent ARP table program entry, and writes the previous round of program processing result (modified routing table) back to the PS information in the PSB.
[0077] Next, it enters the ARP table program entry again, and the matching module constructs a new Key. Exemplarily, the matching module reads out the next-hop IP address (32 bits, from the routing table) and the destination egress information (egress interface index 16 bits, from the routing table), which are required to construct the Key, from the PS of the PSB, and constructs the Key for querying the ARP table based on these two index information: next-hop IP address + destination interface index. Then it sends the index to the I / O request interface, sets the PS storage location of the I / O return result and the corresponding ARP table processing program entry (for subsequent execution modules to use), and then ends.
[0078] After obtaining the corresponding ARP table (128-bit) through an external I / O device, the ARP table will be returned to the matching module through the I / O response interface. After receiving the result, the matching module will write it into the PS according to the PS storage location set in the previous step, and trigger the execution module to process it after completion.
[0079] The execution module starts the processing program for the ARP table. First, it reads the required PS information (including the ARP table, packet header, and some other information, etc.) from the PSB, and executes the program, such as constructing the MAC link layer header using the destination MAC address in the ARP table. Then it sets the entry for subsequent program processing and writes the program processing result back to the PS in the PSB. Until the entry for the subsequent program is not at this node, the PS will be read out in full from the PSB and sent to the next pipeline node.
[0080] Similarly, in Figure 5 the black solid line is used to represent the data plane of the pipeline node. It can be seen that when the full first PS information is transferred from the previous pipeline node to this pipeline node, it is first written into the PSB in full. After the matching module and the execution module complete multiple rounds of matching - processing processes, the final PS information is read out from the PSB in full. In this way, no matter how many rounds of matching - execution actions are completed by the pipeline node, the full PS information will only be read and written once in a pipeline node, without being copied back and forth between the matching module and the execution module. In this way, the number of read and write operations of the full PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.
[0081] At the same time, when the matching module reads data from the PSB, it only needs to read out the index information as needed. The average length of this index information is generally about 8 bits, and its reading volume is very small. And when the matching module writes data to the PSB, it only needs to write the small amount of feature information (generally 32 bits) obtained into the original first PS information. Similarly, the execution module also only needs to read and write a small part of the data as needed. Generally, the average length of the data read by the execution module is about 128 bits, and the processing result written is generally 64 bits. Neither the matching module nor the execution module needs to read and write the full PS information, which greatly reduces the amount of data read and written in memory and further reduces the power consumption of the network processor.
[0082] The black dashed line is used to represent the control plane of the pipeline node. Among them, after the matching module finishes the matching action, it only needs to send a completion notice to the execution module. This message is a trigger signal with a very small amount of data. Similarly, when the execution module finishes the processing action, it only needs to send feedback information to the matching module to trigger the next round of matching - processing process. This message is also a feedback signal with a very small amount of data. In this way, the matching module and the execution module do not need to copy the full - volume PS information back and forth, further reducing the power consumption of the network processor.
[0083] It can be seen from the above description that if the pipeline node is loop - back type, each pipeline node can execute the matching - processing process multiple times. Then, compared with the existing pipeline node architecture where both the matching module and the execution module set caches, the pipeline node provided by the embodiment of the present application only needs to read and write the full - volume PS once. If it loops back N times, then N - 1 times of full - volume PS reading and writing can be saved, which will greatly reduce the dynamic power consumption of the network processor.
[0084] Based on the above description, Figure 6 It is a schematic flowchart of a packet - processing method provided by an embodiment of the present application. This packet - processing method is applied to any of the above network processors, where the matching module and the execution module in the network processor share a cache. As Figure 6 shown, when the network processor executes the packet - processing method, each module executes the following steps:
[0085] 601. The cache stores the first program status information.
[0086] First, there are pipeline nodes in the network processor. The pipeline node includes a matching module and an execution module, and the matching module and the execution module share the cache PSB. When the pipeline node starts to process a packet, the PSB is first used to store the PS information of the packet, such as: the Ethernet header, IP header, TCP header of the packet, global variables and local variables during program operation, table entries returned by IO, and instruction addresses related to the program control flow, etc. Among them, the first PS information of the packet needs to be written into the PSB in full volume.
[0087] 602. The matching module updates the first program status information in the cache according to the packet to be processed, and obtains the second program status information.
[0088] Next, when the message to be processed is transmitted to the pipeline node, the matching module reads part of the information as needed, obtains the feature information of the message to be processed from the external I / O device based on the partial data, and then writes the feature information back to the first PS information in the PSB, completing the update of the first PS information. Specifically, the matching module first obtains the index information from the first PS information stored in the PSB according to the message to be processed, and then generates an index based on the index information. Then, based on the index, the feature information corresponding to the message to be processed is obtained from the external I / O device, and the feature information is written back to the first PS information in the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.
[0089] 603. The matching module sends a completion notice of the update to the execution module.
[0090] Next, after the matching module completes the update of the first PS information, the matching module needs to send a completion notice of the update to the execution module to notify the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts to process the second PS information in the PSB.
[0091] 604. The execution module responds to the completion notice of the update and starts to process the second program status information in the cache.
[0092] After the execution module starts the processing process, it also reads information as needed. Specifically, the execution module first reads part of the information to be processed as needed from the second PS information in the PSB, and the information to be processed includes the feature information of the message to be processed previously written by the matching module. Then, the execution module processes the information to be processed and writes the processing result to the second PS information in the PSB. In this way, the second PS information in the PSB is updated to the third PS information.
[0093] 605. The execution module sends feedback information to the matching module.
[0094] Next, if the matching module and the execution module only execute the matching - processing process once, after the execution module finishes the processing action, it reads out all the third PS information in the PSB in full and writes it in full into the PSB of the next pipeline node. When the matching module and the execution module execute multiple rounds of matching - processing processes, the execution module needs to send feedback information to the matching module. After receiving the feedback information, the matching module starts to execute the next round of matching - processing process. Similarly, the matching module obtains other index information from the third PS information stored in the PSB according to the packet to be processed, and then generates an index based on the index information. Then, based on the index, other feature information corresponding to the packet to be processed is obtained from the external I / O device, and then the other feature information is written back into the third PS information of the PSB. At this time, the third PS information stored in the PSB is updated to the fourth PS information. Then, the matching module still needs to send an update completion notification to the execution module to notify the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts to process the updated fourth PS information in the PSB, and so on, which will not be elaborated here.
[0095] In the embodiments of the present application, the matching module and the execution module of the network processor share a cache. In this way, when the network processor executes the processing process, all the first PS information is first written into the shared cache. Then, the matching module reads out part of the information from the first PS information as needed according to the packet to be processed, obtains the feature information of the packet to be processed based on this part of the data, and writes the feature information into the first PS information to obtain the second PS information. Then, the execution module also reads out part of the information (including the feature information) in the second PS information as needed, then processes the feature information, and finally writes the processing result into the second PS information. It is not until the matching module and the execution module complete all the processing processes that the final all - volume PS information is read out from the cache. Therefore, the number of full - volume read and write times of the PS information will be greatly reduced, thus saving the bandwidth resources of the network processor. At the same time, the matching module and the execution module only need to read out part of the data as needed, and the amount of data read or written by the two modules is also greatly reduced, which further saves the bandwidth resources of the network processor and reduces the power consumption of the network processor.
[0096] The embodiments of the present application also provide a computer - readable storage medium, in which at least one computer program instruction is stored. The computer program instruction is loaded and executed by a processor to implement the Figure 6 packet - processing method in the above - mentioned embodiments, which will not be elaborated here.
[0097] The embodiments of the present application also provide a computer program product, including computer - executable instructions. When the computer - executable instructions run on a computer, the computer is enabled to execute the Figure 6 packet - processing method in the above - mentioned embodiments.
[0098] The embodiment of the present application further provides a chip, which includes a network processor as described in Figures 3A to 3C any one of the above, and each module of the network processor is used to execute the packet processing method in the above Figure 6 embodiment, which will not be elaborated here.
[0099] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes (or functions) described in the embodiment of the present application are implemented. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc. In the embodiment of the present application, the computer may include the device described above.
[0100] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
Claims
1. A network processor, characterized in that, The network processor includes a matching module, an execution module, and a cache; Among them, the matching module and the execution module share the cache; The cache is used to store first program status information; The matching module is used to update the first program status information according to the packet to be processed to obtain second program status information; wherein, the second program status information includes the feature information of the packet to be processed; The execution module is used to process the second program status information in the cache.
2. The network processor according to claim 1, characterized in that, The matching module is further used to send an update completion notification to the execution module; the update completion notification is used to indicate that the matching module has completed the update of the first program status information; The execution module is specifically used to respond to the update completion notification and start processing the second program status information in the cache.
3. The network processor according to claim 1 or 2, characterized in that, The matching module is specifically used for: Obtaining index information from the first program status information according to the packet to be processed; the index information is used to generate an index; Obtaining the feature information according to the index and writing the feature information into the first program status information in the cache to obtain the second program status information.
4. The network processor according to claim 2 or 3, characterized in that, The execution module is specifically used for: Reading the information to be processed from the second program status information in the cache, where the information to be processed includes the feature information; Processing the information to be processed and writing the processing result into the second program status information to obtain third program status information.
5. The network processor according to claim 4, characterized in that, The execution module is further used to send feedback information to the matching module; the feedback information is used to indicate that the matching module performs the next round of update on the third program status information.
6. The network processor according to any one of claims 1 to 5, characterized in that, The network processor includes pipeline nodes; Among them, each pipeline node includes the matching module, the execution module, and the cache; In each pipeline node, all the matching modules and the execution modules share the cache.
7. A message processing method, applied to a network processor, the network processor comprising a matching module, an execution module and a cache; In the network processor, the matching module and the execution module share the cache; characterized in that, The method includes: The cache stores first program status information; The matching module updates the first program status information in the cache according to the packet to be processed to obtain second program status information; wherein, the second program status information includes the feature information of the packet to be processed; The execution module processes the second program status information in the cache.
8. The message processing method according to claim 7, characterized in that, The matching module updates the first program status information in the cache according to the packet to be processed to obtain second program status information, including: The matching module obtains index information from the first program status information according to the packet to be processed; the index information is used to generate an index; The matching module obtains the feature information according to the index and writes the feature information into the first program status information in the cache to obtain the second program status information.
9. The method according to claim 7 or 8, characterized in that, The method further includes: The matching module sends an update completion notification to the execution module; the update completion notification is used to indicate that the matching module has completed the update of the first program status information; The execution module processes the second program status information in the cache, including: The execution module responds to the update completion notification and starts to process the second program status information in the cache.
10. The message processing method according to any one of claims 7 to 9, characterized in that The execution module processes the second program status information in the cache, including: The execution module reads the information to be processed from the second program status information in the cache, and the information to be processed includes the feature information; The execution module processes the information to be processed and writes the processing result into the second program status information to obtain the third program status information.
11. The message processing method according to claim 10, characterized in that The method further includes: The execution module sends feedback information to the matching module; the feedback information is used to instruct the matching module to perform the next round of update on the third program status information.
12. A computer-readable storage medium storing computer instructions, characterized in that Including: Computer instructions, wherein when the computer instructions are executed, the computer is caused to execute the method according to any one of claims 7-10.
13. A chip, characterized in that The chip includes the network processor according to any one of claims 1 to 6, and each module of the network processor is used to execute the message processing method according to any one of claims 7 to 11.
Citation Information
Cited By
Network processor and message processing method
WO2025130820A1