Flow unloading method and device based on DPU network card and medium
By adding hardware bubblers to the traffic offload component of the DPU network card, flexible adjustments to the matching fields are achieved, solving the problem that users cannot specify precise flow table field matching rules at will, improving configuration flexibility and efficiency, and reducing costs.
Patent Information
- Application Number
- CN202510400355.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, when processing network traffic in DPU network cards, users cannot arbitrarily specify the field matching rules of the precise flow table. In the P4 programming language, the code needs to be rewrite when the requirements change, resulting in high thresholds and high labor costs.
The hardware bubbler is added to the traffic unloading component of the DPU network card. Through the hardware bubbler, the full number of key fields is filtered according to the mask information configured by the user, and the fields to be matched are obtained, and they are sent to the processor for matching, so as to achieve flexible adjustments to the fields to be matched.
It improves the flexibility and efficiency of precise flow table configuration, reduces configuration thresholds and labor costs, and allows users to flexibly adjust the fields to be matched according to their needs.
Smart Images

Figure CN120151279A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a traffic offloading method, device, and medium based on a DPU network card. Background Art
[0002] With the rapid development of cloud network and data center related technologies and services, DPU network cards have become network devices that can perform flow table hardware offloading. In the process of a DPU network card processing network traffic, there are the following two processing scenarios.
[0003] In the flow table scenario under a general virtual switch, the virtual switch and the offloading module need to reach an agreement in advance before traffic processing. When the processor matches the exact flow table, it performs table look-up matching according to the negotiated field matching rules. However, in this scenario, users cannot arbitrarily specify the field matching rules of the exact flow table.
[0004] In the flow table scenario under the Programming Protocol-independent Packet Processors (P4 for short), professional developers can write a code file corresponding to the packet parsing and flow table matching rules according to user requirements, and download the compiled file to the offloading component to reflect user requirements in the application process. When user requirements change, developers need to modify or write the P4 language again based on the changed user requirements, resulting in a higher threshold for requirement changes and higher labor costs. Summary of the Invention
[0005] To solve the above technical problems, the present disclosure provides a traffic offloading method, device, and medium based on a DPU network card.
[0006] In a first aspect, the present disclosure provides a traffic offloading method based on a DPU network card. The method is applied to a traffic offloading component in the DPU network card. The traffic offloading component includes a hardware packet parser, a hardware bubble squeezer, a processor, and a traffic manager. The method includes:
[0007] The hardware packet parser extracts full keyword fields from the pre-received packets and sends the full keyword fields to the hardware bubble squeezer; the hardware bubble squeezer filters some keyword fields in the full keyword fields according to the latest mask information to obtain fields to be matched, and sends the fields to be matched to the processor. The mask information is configured by the user based on requirements, and the field length of the fields to be matched is less than the field length of the full keyword fields; the processor matches the fields to be matched with the exact flow table stored in the processor, and obtains the processing rule of the packet when the matching is successful; the processor processes the packet in accordance with the processing rule and coordinates with the preset scheduling policy in the traffic manager.
[0008] In some alternative embodiments, before the hardware message parser extracts all key fields from the pre-received message and sends the all key fields to the hardware bubble squeezer, the method further includes:
[0009] The hardware bubble squeezer receives the mask information sent by the virtual switch.
[0010] In some alternative embodiments, when the exact flow table is the first exact flow table updated based on the mask information, before the hardware message parser extracts all key fields from the pre-received message and sends the all key fields to the hardware bubble squeezer, the method further includes:
[0011] The processor receives the second exact flow table sent by the virtual switch, and updates the original exact flow table stored in the processor based on the second exact flow table to obtain the first exact flow table. The second exact flow table is the flow table obtained by the virtual switch updating the original exact flow table stored therein based on the mask information.
[0012] In some alternative embodiments, when the exact flow table is the original exact flow table before update, the method further includes:
[0013] When the field to be matched fails to match the exact flow table, match the field to be matched with the second exact flow table stored in the virtual switch.
[0014] In some alternative embodiments, the processor matches the field to be matched with the exact flow table stored in the processor, including:
[0015] The processor matches the field to be matched with the fields in the corresponding field domains of each entry in the exact flow table respectively.
[0016] In some alternative embodiments, when the field to be matched successfully matches the second exact flow table and the exact flow table has not been updated yet, the method further includes:
[0017] The processor receives the matching entry sent by the virtual switch, and updates the matching entry to the exact flow table. The matching entry is the entry in the second exact flow table that matches the field to be matched.
[0018] In some alternative embodiments, when a match is successful, obtaining the processing rule for the message includes:
[0019] When the field to be matched successfully matches the exact flow table, determine the entry in the exact flow table that matches the field to be matched as the target entry; determine the processing rule in the rule domain corresponding to the target entry as the processing rule for the message.
[0020] In a second aspect, the present disclosure provides a computer device, including:
[0021] A memory and a processor are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the traffic offloading method based on a DPU network card according to the first aspect and any of its embodiments.
[0022] In a third aspect, the present disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to execute the traffic offloading method based on a DPU network card according to the first aspect and any of its embodiments.
[0023] In a fourth aspect, the present disclosure provides a computer program product including a computer program, where when the computer program is executed by a processor, the steps of the traffic offloading method based on a DPU network card according to the first aspect and any of its embodiments are implemented.
[0024] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0025] In the traffic offloading method based on a DPU network card provided in this embodiment, first, a hardware packet parser extracts full-length key fields from a pre-received packet and sends the full-length key fields to a hardware bubble squeezer; second, the hardware bubble squeezer filters some of the key fields in the full-length key fields according to the latest mask information to obtain fields to be matched, and sends the fields to be matched to the processor. The mask information is configured by the user based on requirements, and the field length of the fields to be matched is less than the field length of the full-length key fields; third, the processor matches the fields to be matched with an exact flow table stored in the processor and obtains a processing rule for the packet when the match is successful; finally, the processor processes the packet according to the processing rule in coordination with a preset scheduling policy in a traffic manager. The above solution realizes flexible adjustment of the fields to be matched by adding a hardware bubble squeezer in the traffic offloading component, not only improves the subsequent matching efficiency by reducing the field length of the fields to be matched, but also can continuously change the fields to be matched according to user requirements, improving the flexibility of the exact flow table configuration. In addition, the mask information in the embodiments of the present disclosure is configured by the user himself, reducing the configuration threshold and labor cost. Description of the Drawings
[0026] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0027] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a schematic structural diagram of a traffic offloading component in a DPU network card in the related art;
[0029] Figure 2 It is a schematic structural diagram of the traffic offloading component provided by an embodiment of the present disclosure;
[0030] Figure 3 It is a schematic flowchart of a traffic offloading method based on a DPU network card provided by an embodiment of the present disclosure;
[0031] Figure 4 It is a structural connection diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0032] In order to be able to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0033] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.
[0034] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0035] With the rapid development of cloud network and data center related technologies and services, the scale of DPU network cards as computing nodes in cloud network servers is getting larger and larger, and the requirements of upper-layer services are also becoming more and more diverse. Considering the packet forwarding delay and throughput, various hardware acceleration technologies have emerged, and DPU has also become a network device that can perform flow table hardware offloading.
[0036] The processing process of the current message in the DPU is mainly implemented through the offloading component as shown in Figure 1 . The offloading component includes a hardware message parser 101, a processor 102, and a traffic manager 103. After receiving the message, the hardware message parser 101 first extracts the key fields from the received message according to the message parsing rules, and sends the set composed of all key fields to the processor 102. Then, the processor 102 matches the key fields in the set with the exact flow table pre-stored in the processor 102. When the key fields of the message hit the exact flow table, the processor 102 collaborates with the traffic manager 103 to send the message, thus realizing the offloading of the CPU by using the DPU network card.
[0037] However, the message parsing rules and the matching rules of the exact flow table involved in the above traffic processing process are either negotiated and agreed before application, or written by professional developers according to user requirements and obtained after being compiled by a Programming Protocol-independent Packet Processors (abbreviated as P4) compiler. The first method has the defect that users cannot change it at will, and the second method has the deficiencies of high technical threshold and high labor cost.
[0038] Therefore, the embodiments of the present disclosure provide a traffic offloading method, device, and medium based on a DPU network card. By adding a hardware bubble squeezer to the traffic offloading component, flexible adjustment of the fields to be matched is realized. Not only the efficiency of subsequent matching is improved by reducing the length of the fields to be matched, but also the fields to be matched can be continuously changed according to user requirements, improving the flexibility of the exact flow table configuration. In addition, the mask information in the embodiments of the present disclosure is configured by the user himself, reducing the configuration threshold and labor cost.
[0039] According to an embodiment of the present invention, an embodiment of a traffic offloading method based on a DPU network card is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0040] In this embodiment, a traffic offloading method based on a DPU network card is provided, which is applied to a traffic offloading component in the DPU network card. The traffic offloading component is as shown in Figure 2 . It includes a hardware message parser 201, a hardware bubble squeezer 202, a processor 203, and a traffic manager 204. Figure 3 is a flowchart of a traffic offloading method based on a DPU network card according to an embodiment of the present invention, as shown in Figure 3As shown, the process includes the following steps:
[0041] S301, the hardware message parser extracts all keyword fields from the pre-received message and sends the all-keyword fields to the hardware bubble squeezer.
[0042] Among them, the all-keyword fields are all fields carried in the message related to the message transmission and processing process. For example, the all-keyword fields include the destination MAC address (Destination MAC Address, abbreviated as DMAC), source MAC address (Source MAC Address, abbreviated as SMAC), virtual local area network tag (abbreviated as VLAN tag), source IP address (Source IP Address, abbreviated as SIP), destination IP address (Destination IP Address, abbreviated as DIP), and protocol type (Protocol Type, abbreviated as PROTOCAL), etc.
[0043] Specifically, when the traffic offloading component in the DPU network card receives a message, it first transmits the message to the hardware message parser. The hardware message parser parses the message according to the pre-configured field parsing rules to extract all keyword fields from the message and sends the extracted all-keyword fields to the hardware bubble squeezer. It should be noted that in the embodiments of the present disclosure, the field parsing rules relied on by the hardware message parser during the message parsing process are used to parse and extract all keyword fields carried in the message. Once the field parsing rules are downloaded, they will remain unchanged, regardless of how the user requirements change during the subsequent message processing process, which has nothing to do with the field parsing rules.
[0044] In some optional embodiments, before S301, the method further includes: the hardware bubble squeezer receives the mask information sent by the virtual switch.
[0045] Among them, the mask information is the mask corresponding to the all-keyword fields, and the mask information is used to mask (i.e., filter) the keyword fields that the user does not need and retain the keyword fields that the user needs.
[0046] Specifically, when the user's requirements change, the user can configure the mask information corresponding to all keyword fields based on their own requirements. After the mask information configuration is completed, the configured mask information is uploaded to the virtual switch. After receiving the mask information uploaded by the user, the virtual switch distributes the mask information to the hardware bubble squeezer. Therefore, when the hardware bubble squeezer receives the mask information distributed by the virtual switch, it obtains the mask information for subsequent determination of the fields to be matched in the packet. It should be noted that after obtaining the mask information, the original mask information in the hardware bubble squeezer can be overwritten with the mask information to ensure that only the latest version of the mask information is stored in the hardware bubble squeezer, thereby ensuring the accuracy of the fields to be matched determined subsequently.
[0047] S302. The hardware bubble squeezer filters out some keyword fields in all keyword fields according to the latest version of the mask information to obtain the fields to be matched, and sends the fields to be matched to the processor.
[0048] Among them, the latest version of the mask information is the mask information obtained by the hardware bubble squeezer from the virtual switch, and this mask information is configured by the user based on requirements. The fields to be matched are the fields that need to be matched during the packet transmission process, and the fields to be matched correspond to the user's requirements. It can be understood that the fields to be matched are the result of the hardware bubble squeezer filtering all keyword fields according to the latest version of the mask information. Therefore, the field length of the fields to be matched is less than the field length of all keyword fields.
[0049] Specifically, after receiving all keyword fields transmitted by the hardware packet parser, the hardware bubble squeezer performs mask processing on all keyword fields extracted by the hardware packet parser according to the latest version of the mask information distributed by the virtual switch, filters out the unnecessary keyword fields (i.e., some keyword fields) in all keyword fields, and determines the remaining keyword fields as the fields to be matched. After obtaining the fields to be matched, the fields to be matched are sent to the processor.
[0050] Exemplarily, in an example, all keyword fields are DMAC, SMAC, VLAN, SIP, DIP, and PROTOCAL. After mask information processing, keyword fields such as SMAC, VLAN, and SIP are masked, and the remaining keyword fields DMAC, DIP, and PROTOCAL are retained as the fields to be matched.
[0051] S303. The processor matches the fields to be matched with the exact flow table stored in the processor, and obtains the processing rule of the packet when the match is successful.
[0052] Among them, the exact flow table is the exact flow table stored in the processor. The exact flow table can be the original exact flow table stored in the processor (i.e., the exact flow table obtained after the last update operation), or the first exact flow table obtained by updating the original exact flow table based on the latest mask information. The exact flow table includes multiple entries, and each entry includes a field domain and a rule domain. The field domain includes at least one matching field, and the rule domain includes the processing rule or processing method of the packet. The exact flow table is sent by the virtual switch.
[0053] Specifically, after obtaining the fields to be matched, the processor matches the fields to be matched with the fields in the field domain of each entry in the exact flow table respectively. When the fields to be matched are exactly the same as the fields in the field domain of a certain entry in the exact flow table, it is determined that the fields to be matched match the exact flow table successfully. At this time, the entry in the exact flow table that matches the fields to be matched is determined as the target entry, and the processing rule in the rule domain corresponding to the target entry is determined as the processing rule of the packet. The processing rules include forwarding the packet to a certain port and discarding the current packet, etc.
[0054] In some optional embodiments, when the exact flow table is the first exact flow table updated based on the mask information, before S301, the method further includes: the processor receives the second exact flow table sent by the virtual switch, and updates the original exact flow table stored in the processor based on the second exact flow table to obtain the first exact flow table.
[0055] Among them, the second exact flow table is the flow table obtained by the virtual switch updating the original exact flow table stored therein based on the mask information. The original exact flow table can be understood as the exact flow table obtained after the last update operation. Both the virtual switch and the processor have the original exact flow table.
[0056] Specifically, when the exact flow table is the first exact flow table updated based on the mask information, the update operation of the original exact flow table in the processor needs to be completed before S301. Therefore, before S301, when the virtual switch receives the mask information uploaded by the user, while sending the mask information to the hardware squeezer, the virtual switch updates the original exact flow table stored in the virtual switch based on the received mask information to obtain the second exact flow table, and sends the second exact flow table to the processor. Through the above method, the embodiments of the present disclosure ensure the consistency between the first exact flow table in the processor and the second exact flow table in the virtual switch.
[0057] After receiving the second precise flow table sent by the virtual switch, the processor updates the original precise flow table stored in the processor according to the second precise flow table to obtain the first precise flow table. The update method of the first precise flow table in this embodiment is not limited. For example, the original precise flow table in the processor can be directly overwritten with the second precise flow table sent by the virtual switch to obtain the first precise flow table. Or the original precise flow table in the processor can be modified according to the second precise flow table sent by the virtual switch to obtain the first precise flow table.
[0058] In some optional embodiments, when the precise flow table is the original precise flow table before update, the method further includes: when the field to be matched fails to match the precise flow table, matching the field to be matched with the second precise flow table stored in the virtual switch.
[0059] Specifically, the second precise flow table is the flow table obtained by the virtual switch after updating the original precise flow table therein based on the mask information. Therefore, the second precise flow table is updated in real time according to the user requirements. Theoretically, the second precise flow table in the virtual switch is always consistent with the precise flow table stored in the processor. However, in the actual application process, there may be a short-term inconsistency between the two due to reasons such as network conditions, transmission delay, and data volume. For example, the virtual switch has synchronized the mask information to the hardware squeezer and has also sent the second precise flow table obtained after updating based on the mask information to the processor, but the processor has not fully received the second precise flow table. In this scenario, the precise flow table in the virtual switch is the updated second precise flow table, while the precise flow table in the processor is the original precise flow table before update. At this time, if the processor performs a matching operation on the field to be matched with the precise flow table stored therein, it is easy to encounter a matching failure. When the field to be matched fails to match the precise flow table stored in the processor, the processor transmits the field to be matched to the virtual switch, and then the virtual switch matches the field to be matched with the second precise flow table stored in the virtual switch. After a successful match, the virtual switch determines the processing rule of the packet and completes the subsequent processing of the packet according to the processing rule.
[0060] In some optional embodiments, when the field to be matched successfully matches the second precise flow table and the precise flow table has not been updated yet, the method further includes: the processor receives the matching entry sent by the virtual switch and updates the matching entry to the precise flow table.
[0061] The matching entry is the entry in the second precise flow table that matches the field to be matched.
[0062] Specifically, when the field to be matched successfully matches the second exact flow table stored in the virtual switch and the exact flow table stored in the processor has not been updated yet, the virtual switch sends the entry in the second exact flow table that successfully matches the field to be matched (i.e., the matching entry) to the processor. After the processor receives the matching entry sent by the virtual switch, it adds the matching entry to the exact flow table stored in the processor, so that before the exact flow table in the processor is updated, when encountering a field matching situation of a similar packet again, it can directly perform the matching in the exact flow table of the processor without having to perform the matching in the virtual switch, thereby improving the matching efficiency of the packet.
[0063] S304. The processor processes the packet in accordance with the processing rules and in coordination with the preset scheduling policy in the traffic manager.
[0064] Among them, the preset scheduling policy includes but is not limited to processing priority, traffic size, etc.
[0065] Specifically, after determining the processing rules of the packet, the traffic manager determines the scheduling method of the packet according to the preset scheduling policy in combination with the full set of keyword fields of the packet. Then, based on this scheduling method, the processor processes the packet in accordance with the previously determined processing rules.
[0066] In the traffic offloading method based on the DPU network card provided in this embodiment, first, the hardware packet parser extracts the full set of keyword fields from the pre-received packet and sends the full set of keyword fields to the hardware squeezer; secondly, the hardware squeezer filters out some of the keyword fields in the full set of keyword fields according to the latest mask information to obtain the field to be matched, and sends the field to be matched to the processor. The mask information is configured by the user based on requirements, and the field length of the field to be matched is less than the field length of the full set of keyword fields; thirdly, the processor matches the field to be matched with the exact flow table stored in the processor and obtains the processing rules of the packet when the match is successful; finally, the processor processes the packet in accordance with the processing rules and in coordination with the preset scheduling policy in the traffic manager; the above solution realizes the flexible adjustment of the field to be matched by adding a hardware squeezer in the traffic offloading component, not only improves the subsequent matching efficiency by reducing the length of the field to be matched, but also can continuously change the field to be matched according to the user's needs, improving the flexibility of the exact flow table configuration. In addition, the mask information in the embodiments of the present disclosure is configured by the user himself, reducing the configuration threshold and labor cost.
[0067] The embodiment of the present invention also provides a computer device, please refer to Figure 4 . Such as Figure 4As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 4 In the figure, one processor 10 is taken as an example.
[0068] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0069] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0070] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0071] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.
[0072] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0073] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0074] In addition to the above computer devices and computer-readable storage media, embodiments of the present application can also be a computer program product, which includes computer program instructions that cause the processor to execute the steps of the sound source localization method provided in any embodiment of the present application when the computer program instructions are run by the processor.
[0075] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0076] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A traffic unloading method based on a DPU network card, characterized in that: The method is applied to a traffic unloading component in a DPU network card, wherein the traffic unloading component includes a hardware message parser, a hardware bubble squeezer, a processor and a traffic manager, and the method includes: The hardware message parser extracts the full key field from the pre-received message and sends the full key field to the hardware bubble squeezer; The hardware bubble squeezer filters some key fields in the full key field according to the latest version of the mask information to obtain the to-be-matched field, and sends the to-be-matched field to the processor, wherein the mask information is configured by the user based on the requirements, and the field length of the to-be-matched field is less than the field length of the full key field; The processor matches the to-be-matched field with the precise flow table stored in the processor, and obtains a processing rule for the message when the match succeeds; The processor processes the message according to the processing rule and in coordination with the scheduling strategy preset in the traffic manager.
2. The method according to claim 1, characterized in that Before the hardware message parser extracts the full key field from the pre-received message and sends the full key field to the hardware bubble squeezer, the method further includes: The hardware bubbler receives the mask information sent by the virtual switch.
3. The method according to claim 2, characterized in that When the precise flow table is a first precise flow table updated based on the mask information, before the hardware message parser extracts the full key field from the pre-received message and sends the full key field to the hardware bubble squeezer, the method further includes: The processor receives a second precise flow table sent by the virtual switch, and updates the original precise flow table stored by the processor based on the second precise flow table to obtain the first precise flow table, where the second precise flow table is a flow table obtained after the virtual switch updates the original precise flow table stored therein based on the mask information.
4. The method according to claim 1, characterized in that: When the precise flow table is an original precise flow table before updating, the method further includes: When the to-be-matched field fails to match the precise flow table, the to-be-matched field is matched with a second precise flow table stored in the virtual switch.
5. The method according to claim 1, characterized in that The processor matches the to-be-matched field with the precise flow table stored by the processor, including: The processor matches the to-be-matched fields with fields in the field domain corresponding to each entry in the precise flow table.
6. The method according to claim 4, characterized in that When the to-be-matched field successfully matches the second precise flow table, and the precise flow table has not yet been updated, the method further includes: The processor receives a matching entry sent by the virtual switch, and updates the matching entry to the precise flow table, wherein the matching entry is an entry in the second precise flow table that matches the field to be matched.
7. The method according to claim 1, characterized in that The processing rule of the message obtained when the match is successful includes: When the to-be-matched field successfully matches the precise flow table, determining an entry in the precise flow table that matches the to-be-matched field as a target entry; The processing rule in the rule field corresponding to the target entry is determined as the processing rule of the message.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the traffic unloading method based on the DPU network card according to any one of claims 1 to 7 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the traffic unloading method based on the DPU network card according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the traffic unloading method based on the DPU network card described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Message forwarding method and device, SDN system, network equipment and storage medium
CN117201402A
Packet processing method, programmable network card device, physical server, and storage medium
WO2024067336A1