Network processor and message processing method

By allowing the Match module and Action module to share a cache in the network processor, the problem of frequent full write and read out of PS information in the prior art is solved, and the effect of saving bandwidth and reducing power consumption is achieved.

WO2025130820A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139647
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-12-16
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When existing network processors process packets, program status information needs to be fully written and read back and forth between the Match module and the Action module, resulting in waste of bandwidth resources and increased power consumption.

Method used

The Match module and the Action module share a cache, and the PS information is first written into the shared cache. The Match module and the Action module only need to read out some information from the shared cache as needed, and output it in full after the processing is completed.

Benefits of technology

This greatly reduces the number of read and write times of the full PS information, saves the bandwidth resources of the network processor, and reduces the power consumption of the network processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139647_26062025_PF_FP_ABST
    Figure CN2024139647_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and provides a network processor and a message processing method. The network processor comprises a matching module, an execution module and a cache, the matching module and the execution module sharing the cache. The cache is used for storing first program state information. The matching module is used for updating the first program state information on the basis of a message to be processed, so as to obtain second program state information, the second program state information comprising feature information of said message. The execution module is used for processing the second program state information in the cache. Under the network processor architecture, PS information does not need to be subjected to full write-in and read-out back and forth between two modules, such that the number of full read / write times of PS information is greatly reduced, thus saving bandwidth resources of the network processor, and reducing power consumption of the network processor.
Need to check novelty before this filing date? Find Prior Art

Description

Network processor and message processing method

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 18, 2023, with application number 202311760350.3 and application name “A Network Processor and Message Processing Method”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a network processor and a message processing method. Background Art

[0003] A network processor is a programmable device used to handle various communications tasks, such as packet processing, protocol analysis, and route lookup. Existing network processors primarily have two architectures: an asynchronous run-to-complete (RTC) architecture and a synchronous pipeline architecture. The pipeline architecture consists of numerous pipeline nodes. When processing a message, the message flows sequentially through the pipeline nodes, with each node responsible for completing only a portion of the overall processing.

[0004] The pipeline nodes in a network processor are usually Match-Action (MA) model structures, that is, each pipeline node includes a matching Match module and an execution Action module. The Match module is responsible for matching and obtaining relevant information about the message to be processed to determine the processing action of the Action module. The Action module executes the processing action based on the matching result of the Match module. Usually, a cache structure is set up in the Match module and the Action module respectively. Among them, the cache in the Match module is used to absorb the query delay of input / output I / O, while the cache in the Action module is used to absorb the processing delay.

[0005] Based on the above structure, when a pipeline node processes a message, the corresponding program state (PS) information must first be fully written to the Match module's cache. After the Match module completes the matching process, it must be fully read out and written to the Action module's cache. Thus, throughout the entire message processing flow, the program state information PS is copied back and forth between the Match module's cache and the Action module's cache, meaning it is constantly being written and read in full. This severely wastes network processor bandwidth resources and results in significant power consumption. Summary of the Invention

[0006] The embodiment of the present application provides a new network processor and message processing method. In this network processor, the Match module and the Action module share the same cache, that is, there is no need to set up separate caches for the Match module and the Action module. In this way, when processing a message, the PS information is first written into the shared cache, and the Match module and the Action module only need to read out part of the information from the shared cache as needed. After the Match module and the Action module have successively executed the matching action and the processing action, the shared cache will output the updated PS information in full. In this way, the PS information does not need to be written and read back and forth in full between the two modules, which greatly reduces the number of times the full PS information is read and written, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.

[0007] In a first aspect, embodiments of the present application provide a network processor in which a matching module and an execution module share a cache. The shared cache initially stores first PS information. When a message to be processed is transmitted to the matching module, the matching module updates the first PS information stored in the cache based on the message to be processed. This replaces the first PS information stored in the cache with second PS information that includes characteristic information of the message to be processed. The execution module then processes the second PS information in the cache.

[0008] In the above embodiment, after the full first PS information is stored in the shared cache, the matching module first reads partial information from the first PS information as needed and then updates the first PS information based on the read partial information. In this way, the first PS information stored in the shared cache is updated to the second PS information that carries the characteristic information corresponding to the message to be processed. The execution module then performs the processing action based on the second PS information in the cache. That is, after the PS information is fully written to the shared cache, it is not fully read out until the matching module and the execution module have completed the relevant actions in sequence. This eliminates the need for PS information to be copied back and forth between the matching module and the execution module, significantly reducing the number of times the full PS information needs to be read and written, thereby conserving network processor bandwidth resources and reducing network processor power consumption.

[0009] In one optional implementation, after updating the first PS information, the matching module needs to send an update completion notification to the execution module. Upon receiving this update completion notification, the execution module initiates the processing flow and begins processing the second PS information stored in the cache. Compared to existing architectures, after completing the matching action, the matching module does not need to send the full second PS information to the execution module. Instead, it simply sends a trigger signal, the update completion notification. This is because the matching module and the execution module share a shared cache and can both directly access the shared cache. This reduces the number of times the full PS information is read, thus reducing the power consumption of the network processor.

[0010] In one optional embodiment, when performing a match, the matching module first reads partial information from the first PS information stored in the shared cache as needed, then generates an index based on the partial information. Next, based on the index, it queries an external device for feature information of the message to be processed and writes the retrieved feature information to the first PS information in the shared cache. In this way, the first PS information in the shared cache is updated to the second PS information. In this embodiment, the matching module only needs to read or write a small amount of information as needed, significantly reducing the amount of data read and written, further reducing the power consumption of the network processor.

[0011] In one optional embodiment, when executing a processing action, the execution module first reads a portion of the pending information from the second PS information in the shared cache as needed. This pending information includes the characteristic information of the pending message previously written by the matching module. The execution module then processes the pending information and writes the processing results to the second PS information in the shared cache. This updates the second PS information in the shared cache to the third PS information. Similarly, the execution module only needs to read or write a small amount of information as needed, thus reducing the amount of data read and written and further reducing the power consumption of the network processor.

[0012] In an optional embodiment, after completing this round of actions, the execution module can also send feedback information to the matching module to notify the matching module that this round of updates has been completed. After receiving this feedback information, the matching module can also perform the next round of updates to the PS information. For example, it can read other partial information as needed, generate another index, and then query other feature information of the message to be processed based on this index, and then update the PS information based on this other feature information. It can be understood that when the execution module and the matching module perform multiple rounds of updates to the PS information, the PS information only needs to be fully written and read from the shared cache once, which will greatly reduce the dynamic power consumption of the network processor.

[0013] In one optional embodiment, the network processor is composed of pipeline nodes in a pipeline architecture. Each pipeline node includes a matching module and an execution module. As a message to be processed passes through each pipeline node, each pipeline node only executes a portion of the entire processing process. Each pipeline node can have only one cache, which is shared by the matching module and execution module within the pipeline node.

[0014] In a second aspect, embodiments of the present application provide a message processing method, which is applied to a network processor. The matching module and execution module in the network processor share a cache. The message processing method includes the following steps: first, storing first PS information in the shared cache. When a message to be processed is transmitted to the matching module, the matching module updates the first PS information stored in the cache based on the message to be processed. This replaces the first PS information stored in the cache with second PS information that includes characteristic information of the message to be processed. The execution module then processes the second PS information in the cache.

[0015] In the above embodiment, after the full first PS information is stored in the shared cache, the matching module first reads partial information from the first PS information as needed and then updates the first PS information based on the read partial information. In this way, the first PS information stored in the shared cache is updated to the second PS information that carries the characteristic information corresponding to the message to be processed. The execution module then performs the processing action based on the second PS information in the cache. That is, after the PS information is fully written to the shared cache, it is not fully read out until the matching module and the execution module have completed the relevant actions in sequence. This eliminates the need for PS information to be copied back and forth between the matching module and the execution module, significantly reducing the number of times the full PS information needs to be read and written, thereby conserving network processor bandwidth resources and reducing network processor power consumption.

[0016] In one optional embodiment, when updating the first PS information in the cache, the matching module first reads a portion of the first PS information stored in the shared cache as needed, then generates an index based on the read portion. Next, based on the index, it queries an external device for feature information of the message to be processed and writes the retrieved feature information to the first PS information in the shared cache. In this way, the first PS information in the shared cache is updated to the second PS information. In this embodiment, the matching module only needs to read or write a small amount of information as needed, significantly reducing the amount of data read and written, further reducing the power consumption of the network processor.

[0017] In one optional implementation, after updating the first PS information, the matching module needs to send an update completion notification to the execution module. Upon receiving this update completion notification, the execution module initiates the processing flow and begins processing the second PS information stored in the cache. Compared to existing architectures, after completing the matching action, the matching module does not need to send the full second PS information to the execution module. Instead, it simply sends a trigger signal, the update completion notification. This is because the matching module and the execution module share a shared cache and can both directly access the shared cache. This reduces the number of times the full PS information is read, thus reducing the power consumption of the network processor.

[0018] In one optional embodiment, when the execution module begins executing the processing flow, it first reads a portion of the pending information from the second PS information in the shared cache as needed. This pending information includes the characteristic information of the pending message previously written by the matching module. The execution module then processes the pending information and writes the processing results to the second PS information in the shared cache. This updates the second PS information in the shared cache to the third PS information. Similarly, the execution module only needs to read or write a small amount of information as needed. This reduces the amount of data read and written, further reducing the power consumption of the network processor.

[0019] In an optional embodiment, upon completing a current round of actions, the execution module can also send feedback to the matching module to notify it that this round of updates has been completed. After receiving this feedback, the matching module can then perform the next round of updates to the PS information. It will be appreciated that when the execution module and the matching module perform multiple rounds of PS information updates, the PS information only needs to be fully written and read from the shared cache once, significantly reducing the dynamic power consumption of the network processor.

[0020] In a third aspect, an embodiment of the present application provides a chip comprising: a network processor as described in any one of the first aspects; each module of the network processor is used to run code instructions to execute any one of the message processing methods provided in the second aspect.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer program instruction, and the computer program instruction is loaded and executed by a computer to implement any message processing method provided in the second aspect above.

[0022] The technical effects brought about by any implementation method in the third aspect to the fourth aspect can be referred to the technical effects brought about by the corresponding implementation methods in the first aspect and the second aspect, and will not be repeated here.

[0023] In an embodiment of the present application, the matching module and execution module of the network processor share a cache. In this way, when the network processor executes the processing flow, the full amount of the first PS information is first written into the shared cache. Then, the matching module reads partial information from the first PS information as needed based on the message to be processed, obtains the characteristic information of the message to be processed based on the partial data, and writes the characteristic information into the first PS information to obtain the second PS information. Next, the execution module also reads partial information (including characteristic information) from the second PS information as needed, processes the characteristic information, and finally writes the processing result into the second PS information. Only after the matching module and the execution module have completed the processing flow is the full amount of PS information finally obtained read out of the cache. Therefore, the number of full read and write times of the PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor. At the same time, the matching module and the execution module only need to read partial data as needed, and the amount of data read or written by the two modules is also greatly reduced, which further saves the bandwidth resources of the network processor and reduces the power consumption of the network processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic diagram of the structure of a Pipeline architecture network processor according to an embodiment of the present application;

[0025] FIG2A is a schematic diagram of the structure of a pipeline node shown in an embodiment of the present application;

[0026] FIG2B is a schematic diagram of the structure of another pipeline node shown in an embodiment of the present application;

[0027] FIG3A is a schematic diagram of the structure of a new network processor provided in an embodiment of the present application;

[0028] FIG3B is a schematic diagram of the structure of another new network processor provided in an embodiment of the present application;

[0029] FIG3C is a schematic diagram of the structure of another new network processor provided in an embodiment of the present application;

[0030] FIG4 is a schematic structural diagram of a new pipeline node provided in an embodiment of the present application;

[0031] FIG5 is a schematic structural diagram of another new pipeline node provided in an embodiment of the present application;

[0032] FIG6 is a flow chart of a message processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The embodiment of the present application provides a new network processor and message processing method. In this network processor, the Match module and the Action module share the same cache, that is, there is no need to set up separate caches for the Match module and the Action module. In this way, when processing a message, the PS information is first written into the shared cache, and the Match module and the Action module only need to read out part of the information from the shared cache as needed. After the Match module and the Action module have successively executed the matching action and the processing action, the shared cache will output the updated PS information in full. In this way, the PS information does not need to be written and read back and forth in full between the two modules, which greatly reduces the number of times the full PS information is read and written, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.

[0034] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.

[0035] In the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0036] In addition, to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.

[0037] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0038] It will be understood that the “embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It will be understood that in the various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0039] It is understood that some optional features in the embodiments of the present application may, in certain scenarios, be implemented independently of other features, such as the solution on which they are currently based, to solve corresponding technical problems and achieve corresponding effects. They may also be combined with other features in certain scenarios as needed. Accordingly, the devices provided in the embodiments of the present application may also implement these features or functions accordingly, which will not be described in detail here.

[0040] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referenced to each other. In this application, unless otherwise specified and there is no logical conflict between the various embodiments, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. Different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following implementation methods of this application do not constitute a limitation on the scope of protection of this application.

[0041] To facilitate understanding of the technical solutions of the embodiments of the present application, a brief introduction to the relevant technologies of the embodiments of the present application is given below:

[0042] A network processor is a programmable device used to handle various communications tasks, such as packet processing, protocol analysis, and route lookup. Existing network processors have two main architectures: an asynchronous run-to-complete (RTC) architecture and a synchronous pipeline architecture.

[0043] The Pipeline architecture is composed of numerous pipeline nodes. When processing packets, they flow through the pipeline nodes in sequence, with each pipeline node responsible for completing only a portion of the entire processing flow. The pipeline nodes typically employ a Match-Action model structure. The editable Match-Action model enables protocol-independent network data forwarding processing, consisting of matching logic and action logic. The matching logic is implemented using a hybrid lookup table of static random access memory and ternary content-addressable memory, along with a combination of counters, traffic statistics tables, and general hash tables. The action logic is implemented using a combination of ALU standard Boolean and arithmetic units, packet header modification operations, and hash operations.

[0044] Based on the above description, Figure 1 is a schematic diagram of the structure of a pipeline-architecture network processor according to an embodiment of the present application. The network processor can be used in devices such as routers, switches, and base stations that require high-bandwidth, high-performance data communications. As shown in Figure 1, the network processor includes multiple serial pipeline nodes. During packet processing, a message passes through multiple pipeline nodes in sequence. The message first passes through the parser, which parses the data packet based on the data header specified by the programmer, generating independent packet headers and intermediate information. The parsed message then passes through each pipeline node in sequence, each of which performs a series of packet processing steps. For example, by matching the parsed header with metadata in the flow table, the message can be modified, added, or deleted. Each pipeline node processes only a portion of the message. Within each pipeline node, the Match module is responsible for matching a table in external memory to determine the processing action for the Action module. The Action module then executes the processing action based on the matching result of the Match module. Finally, the processed message passes through the Deparser again, is serialized, and then enters the queue to be output.

[0045] Figure 2A is a schematic diagram of the structure of a pipeline node in an embodiment of the present application. As shown in Figure 2A, the pipeline node includes a matching module and an execution module. The matching module and the execution module each have a buffer for storing program status information (PS) of the message. PS information is program status information, including the message's Ethernet header, IP header, TCP header, global and local variables of the program, table entries returned by I / O, and instruction addresses related to the program's control flow.

[0046] Buffer 1, set up by the matching module, is typically implemented using low-cost memory to absorb I / O latency. Buffer 2, set up by the execution module, is typically implemented using registers to absorb processing latency. Based on this structure, when processing packets, pipeline nodes first copy the entire PS information to Buffer 1 in the matching module. After the matching module completes the matching action, the PS information in Buffer 1 is fully read and copied to Buffer 2 in the execution module. The execution module then performs the execution action based on the PS information in Buffer 2.

[0047] As shown in the figure above, when the matching module and the execution module execute a match-processing flow, the full PS information needs to be copied from Buffer 1 to Buffer 2. When the pipeline node is loopback, meaning that multiple match-processing flows need to be executed, the full PS information is transferred back and forth between the two buffers within the pipeline node. As shown in Figure 2B, assume that after a message enters the pipeline node, it loops back within the node three times, meaning that the node executes the match-processing flow four times for the message. In this case, the memory read and write times for the two buffers are four, respectively, meaning that the two buffers need to read and write a total of eight times the PS. Typically, the full PS in a network processor is on the order of 512 bytes. Eight times the PS read and write means that the processor, in this architecture, needs to read and write 4096 bytes of memory data. This requires a sufficiently large bus width and chip area for the network processor. Furthermore, excessive memory reads lead to a high cache flip rate, significantly increasing the dynamic power consumption of the entire network processor and severely impacting its performance.

[0048] To solve the above problems, an embodiment of the present application provides a new network processor architecture and proposes a message processing method based on the architecture. In the new network processor, the matching module and the execution module share the same cache, that is, there is no need to set up caches for the matching module and the execution module separately. In this way, when processing a message, the PS information is first written into the shared cache, and the matching module and the execution module only need to read out part of the information from the shared cache as needed. After the matching module and the execution module have successively executed the matching action and the processing action, the shared cache will output the updated PS information in full. In this way, the PS information does not need to be fully written and read back and forth between the two modules, which greatly reduces the number of times the full PS information is read and written, thereby saving the bandwidth resources of the network processor and reducing the power consumption of the network processor.

[0049] The following is an introduction to the architecture of the network processor:

[0050] FIG3A is a schematic diagram of the structure of a new network processor provided in an embodiment of the present application. As shown in FIG3A , the network processor is composed of multiple pipeline nodes. Each pipeline node includes a matching module and an execution module, and the matching module and the execution module share a cache. As shown in FIG3A , the pipeline nodes are connected in series, and the message to be processed is input to the next pipeline node after being output from the previous pipeline node. Each pipeline node includes a cache, and all matching modules and execution modules in a pipeline node share the cache.

[0051] For example, in a network processor, multiple serial pipeline nodes can also share the same cache. As shown in Figure 3B, multiple pipeline nodes are connected in series, where each pipeline node only includes a matching module and an execution module, without a cache. A shared cache is provided outside the pipeline nodes, and the matching modules and execution modules of multiple pipeline nodes share the same cache.

[0052] Exemplarily, multiple pipeline nodes in a network processor can also be connected in parallel to form a parallel pipeline architecture. As shown in Figure 3C, multiple pipeline nodes are connected in parallel to form multiple parallel branches. In this way, multiple traffic flows can be processed simultaneously. The pipeline node on one of the branches has a loopback function. For example, node 1 can loop back the message to be processed, that is, it can perform multiple rounds of matching-processing processes on the message to be processed, and then output the processed message. In the parallel architecture, the matching module and execution module of each pipeline node can share the same cache. At the same time, the matching modules and execution modules of multiple pipeline nodes on a branch can also share the same cache, which is not limited here.

[0053] Based on the above description, the following describes in detail the processing flow of the matching module and the execution module for a pipeline node:

[0054] Figure 4 is a schematic diagram of the structure of a new pipeline node provided by an embodiment of the present application. As shown in Figure 4, the pipeline node includes a matching module and an execution module, and the matching module and the execution module share a program state buffer (PSB). The PSB is used to store the PS information of the message, such as the Ethernet header, IP header, TCP header, etc. of the message, the global variables and local variables of the program running, the table entries returned by IO, and the instruction addresses related to the program control flow.

[0055] First, the first PS information of the message must be fully written to the PSB. Then, when the message to be processed is transmitted to the pipeline node, the matching module reads partial information as needed and, based on this partial data, obtains the feature information of the message to be processed from the external I / O device. This feature information is then written back to the first PS information of the PSB, completing the update. Finally, the execution module reads partial information as needed, including the newly written feature information, and then performs data processing based on this feature information to complete the execution action.

[0056] The matching module reads information by first obtaining index information from the first PS information stored in the PSB based on the message to be processed, then generating an index based on the index information. The matching module then obtains feature information corresponding to the message to be processed from an external I / O device based on the index, and then writes the feature information back to the first PS information in the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.

[0057] When the matching module completes the update of the first PS information, the matching module needs to send an update completion notification to the execution module to inform the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts processing the second PS information in the PSB.

[0058] After the execution module begins processing, it also reads information as needed. Specifically, it first reads a portion of the pending information from the second PS information in the PSB. This pending information includes the characteristic information of the pending message previously written by the matching module. The execution module then processes the pending information and writes the processing results to the second PS information in the PSB. This updates the second PS information in the PSB to the third PS information.

[0059] It can be understood that pipeline nodes can be connected serially with other pipeline nodes. When a pipeline node performs a table lookup operation, that is, performs a round of matching-processing process, after the execution module completes the processing action, the PSB outputs the full amount of PS information stored to the next pipeline node.

[0060] The following uses a specific example to illustrate the matching and processing flow of pipeline nodes:

[0061] Assume that the pipeline node is used to modify the IPv4 routing table. Then the PSB of the pipeline first writes the PS information of the message. Then, according to the IPv4 routing table program entry carried in the PS, the matching module is triggered to construct the index key. Exemplarily, the matching module reads the IP (32 bits) of the IPv4 header required to construct the key and the VPN information (16 bits) to which the message belongs from the PS of the PSB. Then, the routing table key is constructed based on the two: VPN+destination IP. Then, the matching module sends the constructed key to the I / O request interface, and sets the PS storage location of the I / O return result and the corresponding routing table processing program entry (used by the subsequent execution module) before ending.

[0062] After obtaining the corresponding routing table through an external I / O device, the matching module returns it to the matching module via the I / O response interface. Upon receiving the query result, the matching module writes the resulting routing table to the corresponding location in the PS information according to the PS storage location set in the previous step. Upon completion, a trigger message is sent to the execution module, which then executes the subsequent processing flow.

[0063] The execution module starts the processing program for the routing table, first reads the PS information required by this program from the PSB (including the routing table, message header and other information written by the matching module), and then obtains the program through the routing table processing program entry and executes it, such as modifying the routing table entries of the routing table, adding new routing table entries to the routing table, etc. After the execution module finishes executing the program, the execution result needs to be sent to the local CPU, and the subsequent program processing entry needs to be set. At the same time, the program processing result (such as the modified routing table) needs to be written back to the PS information in the PSB. Then, if the entry of the subsequent program is not in this node, the modified PS information will be read out in full from the PSB and written into the PSB of the next pipeline node.

[0064] In Figure 4, the solid black line represents the data plane of the pipeline node. As can be seen, when the full first PS information is transferred from the previous pipeline node to the current pipeline node, it is first written to the PSB in its entirety. After the matching module and execution module complete their actions, the updated PS information is fully read out of the PSB. This way, after a pipeline node completes a round of matching and execution, the full PS information does not need to be copied back and forth between the matching and execution modules; it is only read and written once. This significantly reduces the number of reads and writes of the full PS information, thereby conserving network processor bandwidth resources and reducing network processor power consumption.

[0065] At the same time, when the matching module reads data from the PSB, it only needs to read out the index information on demand. The average length of the index information is generally about 8 bits, and the reading amount is very small. And when the matching module writes data to the PSB, it only needs to write the small amount of feature information obtained (generally 32 bits) into the original first PS information. Similarly, the execution module only needs to read and write a small amount of data on demand. Generally, the average length of the data read by the execution module is about 128 bits, and the written processing result is generally 64 bits. Neither the matching module nor the execution module needs to read and write the full amount of PS information, which greatly reduces the amount of data read and written to the memory, and further reduces the power consumption of the network processor.

[0066] The black dashed line represents the control plane of the pipeline node. After completing the matching action, the matching module only needs to send an update completion notification to the execution module, without having to send the full second PS message to the execution module. This message is a very small trigger signal. This is because the matching module and the execution module share a cache and can directly access it. This prevents each module from reading and writing the full PS information, thereby reducing the power consumption of the network processor.

[0067] As can be seen from the above description, if the pipeline node is connected in series with other pipeline nodes, each pipeline node only executes the matching-processing process once. Compared with the existing pipeline node architecture in which both the matching module and the execution module are cached, the pipeline node provided by the embodiment of the present application can save the reading and writing of the full PS. That is, in the existing architecture, the cache of the matching module needs to read and write the PS information once in full, and the cache of the execution module needs to read and write the PS information once in full. However, in the pipeline node of the embodiment of the present application, the PSB only needs to read and write the PS information once in full.

[0068] Figure 5 is a schematic diagram of the structure of another pipeline node provided in an embodiment of the present application. As shown in Figure 5, the pipeline node is different from the pipeline node shown in Figure 4 in that the pipeline node in the embodiment of the present application is loop-back, that is, a pipeline node can execute multiple rounds of matching-processing processes.

[0069] The following describes the process of executing multiple rounds of matching-processing flow in this pipeline node:

[0070] First, the first PS information of the message is fully written to the PSB. Then, when a message to be processed is transmitted to a pipeline node, the matching module first retrieves index information from the first PS information stored in the PSB based on the message to be processed, and then generates an index based on the index information. Based on this index, the matching module then obtains the feature information corresponding to the message to be processed from the external I / O device and writes the feature information back to the first PS information in the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.

[0071] When the matching module completes the update of the first PS information, the matching module needs to send an update completion notification to the execution module to inform the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts processing the second PS information in the PSB.

[0072] After the execution module begins processing, it also reads information as needed. Specifically, it first reads a portion of the pending information from the second PS information in the PSB. This pending information includes the characteristic information of the pending message previously written by the matching module. The execution module then processes the pending information and writes the processing results to the second PS information in the PSB. This updates the second PS information in the PSB to the third PS information.

[0073] Next, the execution module needs to send feedback information to the matching module. After receiving the feedback information, the matching module begins to execute the next round of matching-processing process. Similarly, the matching module obtains other index information from the third PS information stored in the PSB based on the message to be processed, and then generates an index based on the index information. Then, based on the index, it obtains other feature information corresponding to the message to be processed from the external I / O device, and then writes the other feature information back to the third PS information of the PSB. At this time, the third PS information stored in the PSB is updated to the fourth PS information. Next, the matching module still needs to send an update completion notification to the execution module to notify the execution module that the matching module's matching action has been completed. In this way, the execution module triggers the execution process and begins to process the updated fourth PS information in the PSB, and so on. I will not go into details here.

[0074] The following uses a specific example to illustrate the two-round matching-processing flow of pipeline nodes:

[0075] Assume that the pipeline node is used to modify the IPv4 routing table and query the Address Resolution Protocol (ARP). Then the PSB of the pipeline first writes the PS information of the message. Then, according to the IPv4 routing table program entry carried in the PS, it triggers the entry into the matching module to construct the index Key. Exemplarily, the matching module reads the IP (32 bits) of the IPv4 header required to construct the Key and the VPN information (16 bits) to which the message belongs from the PS of the PSB. Then, based on the two, the routing table Key is constructed: VPN+destination IP. Then, the matching module sends the constructed key to the I / O request interface, and sets the PS storage location of the I / O return result and the corresponding routing table handler entry (used by the subsequent execution module) before ending.

[0076] After obtaining the corresponding routing table through an external I / O device, the matching module returns it to the matching module via the I / O response interface. Upon receiving the query result, the matching module writes the resulting routing table to the corresponding PS storage location set in the previous step. Upon completion, a trigger message is sent to the execution module, which then executes the subsequent processing flow.

[0077] The execution module starts the processing program for the routing table, first reads the PS information required by this program from the PSB (including the routing table, message header and other information written by the matching module), and then obtains the program through the routing table processing program entry and executes it, such as modifying the routing table entries of the routing table, adding new routing table entries to the routing table, etc. After the execution module finishes executing the program, it needs to send the execution result to the local CPU, and set the subsequent program processing entry. At the same time, the program processing result (such as the modified routing table) needs to be written back to the PS in the PSB. Assuming that the execution result of the execution module is to continue searching for ARP, then first set the subsequent ARP table program entry, and write the previous round of program processing results (modified routing table) back to the PS information in the PSB.

[0078] Then, the ARP table program entry is entered again, and the matching module constructs a new key. For example, the matching module reads the next-hop IP address (32 bits, from the routing table) and the destination exit information (16 bits of the outgoing interface index, from the routing table) required to construct the key from the PS of the PSB, and constructs the key for querying the ARP table based on these two index information: the next-hop IP address + the destination interface index. The index is then sent to the I / O request interface, and the PS storage location of the I / O return result and the corresponding ARP table processing program entry (used by the subsequent execution module) are set before the process ends.

[0079] After obtaining the corresponding ARP table (128 bits) through the external I / O device, the ARP table is returned to the matching module through the I / O response interface. After receiving the result, the matching module writes it to the PS according to the PS storage location set in the previous step. After completion, the execution module is triggered to process it.

[0080] The execution module initiates the ARP table processing routine, first reading the required PS information (including the ARP table, packet header, and other information) from the PSB and executing the routine. For example, it constructs the MAC link layer header using the destination MAC address in the ARP table. It then sets the entry point for subsequent routine processing and writes the results back to the PS in the PSB. Until the entry point for the subsequent routine is no longer in this node, the PS is fully read from the PSB and sent to the next pipeline node.

[0081] Similarly, in Figure 5, the black solid line represents the data plane of the pipeline node. It can be seen that when the full first PS information is transferred from the previous pipeline node to the current pipeline node, it is first written into the PSB in full. After the matching module and the execution module have completed multiple rounds of matching-processing, the final PS information is fully read out of the PSB. In this way, no matter how many rounds of matching-execution actions a pipeline node completes, the full PS information will only be read and written once in a pipeline node, without having to be copied back and forth between the matching module and the execution module. This greatly reduces the number of read and write operations for the full PS information, thereby saving bandwidth resources and reducing the power consumption of the network processor.

[0082] At the same time, when the matching module reads data from the PSB, it only needs to read out the index information on demand. The average length of the index information is generally about 8 bits, and the reading amount is very small. And when the matching module writes data to the PSB, it only needs to write the small amount of feature information obtained (generally 32 bits) into the original first PS information. Similarly, the execution module only needs to read and write a small amount of data on demand. Generally, the average length of the data read by the execution module is about 128 bits, and the written processing result is generally 64 bits. Neither the matching module nor the execution module needs to read and write the full amount of PS information, which greatly reduces the amount of data read and written to the memory, and further reduces the power consumption of the network processor.

[0083] The black dashed line represents the control plane of the pipeline node. After completing the matching action, the matching module only needs to send an update completion notification to the execution module. This message is a very small trigger signal. Similarly, after completing the processing action, the execution module only needs to send feedback information to the matching module to trigger the next round of the matching-processing process. This message is also a very small feedback signal. This eliminates the need for the matching and execution modules to copy the full amount of PS information back and forth, further reducing the power consumption of the network processor.

[0084] As can be seen from the above description, if the pipeline node is looped, each pipeline node can execute multiple matching-processing processes. Compared to the existing pipeline node architecture where both the matching module and the execution module are cached, the pipeline node provided by the embodiment of the present application only needs to read and write the full PS once. If it loops N times, it can save N-1 full PS reads and writes, which will greatly reduce the dynamic power consumption of the network processor.

[0085] Based on the above description, FIG6 is a flow chart of a message processing method provided by an embodiment of the present application. The message processing method is applied to any of the above-mentioned network processors, wherein the matching module and the execution module in the network processor share a cache. As shown in FIG6, when the network processor executes the message processing method, each module performs the following steps:

[0086] 601. Cache and store first program state information.

[0087] First, there are pipeline nodes in the network processor. The pipeline nodes include a matching module and an execution module, and the matching module and the execution module share a cache called PSB. When the pipeline node starts processing a message, the PSB is first used to store the message's PS information, such as the message's Ethernet header, IP header, TCP header, etc., global variables and local variables of the program running, table entries returned by IO, and instruction addresses related to the program control flow. Among them, the first PS information of the message needs to be fully written to the PSB.

[0088] 602. The matching module updates the first program state information in the cache according to the message to be processed to obtain second program state information.

[0089] Next, when the message to be processed is transmitted to the pipeline node, the matching module reads partial information as needed and obtains the characteristic information of the message to be processed from the external I / O device based on the partial data. It then writes this characteristic information back to the first PS information of the PSB, completing the update of the first PS information. Specifically, the matching module first obtains index information from the first PS information stored in the PSB based on the message to be processed, and then generates an index based on the index information. Based on this index, the characteristic information corresponding to the message to be processed is obtained from the external I / O device, and the characteristic information is then written back to the first PS information of the PSB. In this way, the first PS information stored in the PSB is updated to the second PS information.

[0090] 603. The matching module sends an update completion notification to the execution module.

[0091] Then, when the matching module completes the update of the first PS information, the matching module needs to send an update completion notification to the execution module to inform the execution module that the matching action of the matching module has been completed. In this way, the execution module triggers the execution process and starts processing the second PS information in the PSB.

[0092] 604. The execution module responds to the update completion notification and starts processing the second program status information in the cache.

[0093] After the execution module begins processing, it also reads information as needed. Specifically, it first reads a portion of the pending information from the second PS information in the PSB. This pending information includes the characteristic information of the pending message previously written by the matching module. The execution module then processes the pending information and writes the processing results to the second PS information in the PSB. This updates the second PS information in the PSB to the third PS information.

[0094] 605. The execution module sends feedback information to the matching module.

[0095] Next, if the matching module and execution module only perform the matching-processing process once, the execution module, after completing the processing action, reads the entire third PS information from the PSB and writes it entirely to the PSB of the next pipeline node. However, if the matching module and execution module perform multiple rounds of the matching-processing process, the execution module needs to send feedback to the matching module. After receiving the feedback, the matching module begins the next round of the matching-processing process. Similarly, based on the message to be processed, the matching module retrieves additional index information from the third PS information stored in the PSB and then generates an index based on the index information. Based on this index, the matching module then obtains additional feature information corresponding to the message to be processed from an external I / O device and writes the additional feature information back to the third PS information in the PSB. At this point, the third PS information stored in the PSB is updated to the fourth PS information. The matching module then sends an update completion notification to the execution module to notify it that the matching module's matching action has completed. This triggers the execution process and begins processing the updated fourth PS information in the PSB. This continues in a similar fashion, and will not be further elaborated here.

[0096] In an embodiment of the present application, the matching module and execution module of the network processor share a cache. In this way, when the network processor executes the processing flow, the full amount of the first PS information is first written into the shared cache. Then, the matching module reads partial information from the first PS information as needed based on the message to be processed, obtains the characteristic information of the message to be processed based on this partial data, and writes the characteristic information into the first PS information to obtain the second PS information. Next, the execution module also reads partial information (including characteristic information) from the second PS information as needed, processes the characteristic information, and finally writes the processing result into the second PS information. The full amount of PS information finally obtained is not read out of the cache until the matching module and the execution module have completed the entire processing flow. Therefore, the number of full read and write times of the PS information will be greatly reduced, thereby saving the bandwidth resources of the network processor. At the same time, the matching module and the execution module only need to read partial data as needed, and the amount of data read or written by the two modules is also greatly reduced, which further saves the bandwidth resources of the network processor and reduces the power consumption of the network processor.

[0097] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program instruction. The computer program instruction is loaded and executed by a processor to implement the message processing method in the embodiment of Figure 6 above, which will not be described in detail here.

[0098] An embodiment of the present application further provides a computer program product, including computer execution instructions. When the computer execution instructions are executed on a computer, the computer executes the message processing method in the embodiment of FIG. 6 .

[0099] An embodiment of the present application further provides a chip, which includes a network processor as described in any one of Figures 3A to 3C. Each module of the network processor is used to execute the message processing method in the embodiment of Figure 6 above, which will not be described in detail here.

[0100] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes (or functions) described in the embodiments of the present application are implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more media that can be integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)). In the embodiment of the present application, the computer may include the aforementioned device.

[0101] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art can understand and implement other changes to the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

Claims

1. A network processor, characterized in that: The network processor includes a matching module, an execution module and a cache; Wherein, the matching module and the execution module share the cache; The cache is used to store first program state information; The matching module is used to update the first program state information according to the message to be processed to obtain the second program state information; wherein the second program state information includes the characteristic information of the message to be processed; The execution module is used to process the second program state information in the cache.

2. The network processor according to claim 1, characterized in that: The matching module is further configured to send an update completion notification to the execution module; the update completion notification is configured to indicate that the matching module has completed the update of the first program state information; The execution module is specifically configured to respond to the update completion notification and start processing the second program state information in the cache.

3. The network processor according to claim 1 or 2, characterized in that: The matching module is specifically used for: Acquire index information from the first program state information according to the message to be processed; the index information is used to generate an index; The characteristic information is acquired according to the index, and the characteristic information is written into the first program state information in the cache to obtain the second program state information.

4. The network processor according to claim 2 or 3, characterized in that: The execution module is specifically used for: Reading information to be processed from the second program state information in the cache, the information to be processed including the feature information; The information to be processed is processed, and the processing result is written into the second program state information to obtain the third program state information.

5. The network processor according to claim 4, characterized in that: The execution module is further used to send feedback information to the matching module; the feedback information is used to instruct the matching module to perform a next round of update on the third program state information.

6. The network processor according to any one of claims 1 to 5, characterized in that: The network processor includes a pipeline node; Wherein, each of the pipeline nodes includes the matching module, the execution module and the cache; In each of the pipeline nodes, all the matching modules and the execution modules share the cache.

7. A message processing method, applied to a network processor, the network processor comprising a matching module, an execution module and a cache; In the network processor, the matching module and the execution module share the cache; characterized in that, The method comprises: The cache stores first program state information; The matching module updates the first program state information in the cache according to the message to be processed to obtain the second program state information; wherein the second program state information includes the characteristic information of the message to be processed; The execution module processes the second program state information in the cache.

8. The method according to claim 7, characterized in that The matching module updates the first program state information in the cache according to the message to be processed to obtain the second program state information, including: The matching module obtains index information from the first program state information according to the message to be processed; the index information is used to generate an index; The matching module obtains the feature information according to the index, and writes the feature information into the first program state information in the cache to obtain the second program state information.

9. The method according to claim 7 or 8, characterized in that: The method further comprises: The matching module sends an update completion notification to the execution module; the update completion notification is used to indicate that the matching module has completed the update of the first program state information; The execution module processes the second program state information in the cache, including: The execution module responds to the update completion notification and starts processing the second program state information in the cache.

10. The method according to any one of claims 7 to 9, characterized in that: The execution module processes the second program state information in the cache, including: The execution module reads information to be processed from the second program state information in the cache, where the information to be processed includes the feature information; The execution module processes the information to be processed and writes the processing result into the second program state information to obtain the third program state information.

11. The message processing method according to claim 10, characterized in that: The method further comprises: The execution module sends feedback information to the matching module; the feedback information is used to instruct the matching module to perform a next round of update on the third program state information.

12. A computer-readable storage medium storing computer instructions, characterized in that: include: Computer instructions, wherein when the computer instructions are executed, the computer is caused to perform the method according to any one of claims 7-11.

13. A chip, characterized in that: The chip comprises the network processor as claimed in any one of claims 1 to 6, and each module of the network processor is used to execute the method as claimed in any one of claims 7 to 11.

Citation Information

Patent Citations

  • Network processor and message processing method

    CN120179577A

  • Data cache loading method and device

    CN113934654A

  • Message transmission method and device of network card, computer equipment and storage medium

    CN116545932A

  • Managing network forwarding configurations using algorithmic policies

    US20160050117A1

  • Packet processing method and chip

    WO2022068614A1