Cross-card link convergence method, device, equipment and medium

Through the aggregation device, link aggregation between multiple DPU network cards is solved, and the reliability of DPUs is improved.

CN120034483APending Publication Date: 2025-05-23CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510186991.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In multi-DPU scenarios, due to the independence of the DPU, link aggregation cannot be achieved between DPUs, thereby reducing the reliability of the DPU.

Method used

Through the aggregation device, the link aggregation between multiple DPU network cards is negotiated, and the user's aggregation operation of the target port of the DPU network card is obtained, and the aggregation configuration information is sent, and the target port of the DPU network card is associated and aggregated into an aggregation link.

Benefits of technology

It realizes cross-card link aggregation between multiple DPUs without destroying DPU independence, improving the DPU operation reliability in multi-DPU scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034483A_ABST
    Figure CN120034483A_ABST
Patent Text Reader

Abstract

The invention provides a cross-card link aggregation method and device, equipment and a medium, and relates to the technical field of data processing. The method is applied to a convergence device, the convergence device is in communication connection with a first data processing unit (DPU) network card and a second DPU network card, and the method comprises the following steps: acquiring convergence operation of a user on a first target port in the first DPU network card and a second target port in the second DPU network card, the first target port comprises all or part of physical ports of the first DPU network card, and the second target port comprises all or part of physical ports of the second DPU network card; and in response to the aggregation operation, sending aggregation configuration information to the first DPU network card and the second DPU network card, the aggregation configuration information being used for associating and aggregating the first target port and the second target port into an aggregation link. According to the method, link convergence among a plurality of DPUs is negotiated through the convergence device, so that the reliability of the DPUs can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a cross-card link aggregation method, device, equipment and medium. Background Art

[0002] With the rapid development of technologies such as cloud computing and big data, data centers are facing tremendous pressure in data processing and transmission. The complexity of network processing is gradually increasing, requiring more computing resources. Traditional network interface cards can no longer meet the growing bandwidth and latency requirements.

[0003] In addition to the network transmission functions of traditional basic network cards, smart network cards also provide rich hardware offload acceleration capabilities, which can improve the forwarding rate of network data and reduce the computing pressure of the central processing unit (CPU). As an emerging computing technology derived from smart network cards, data processing units (DPUs) have become a new choice for data center networks with their high performance, programmability and security.

[0004] DPU has an independent operating system and can run independently at the business level. This independence enhances the flexibility and processing capabilities of DPU. However, in a multi-DPU system, due to the independence of DPU, it is impossible to achieve business association between DPUs, which reduces the reliability of DPU. Summary of the invention

[0005] The present application provides a cross-card link aggregation method, device, equipment and medium. The method can improve the reliability of the DPU by negotiating link aggregation between multiple DPUs through an aggregation device.

[0006] In a first aspect, the present application provides a cross-card link aggregation method, which is applied to an aggregation device, and the aggregation device is communicatively connected with a first data processing unit DPU network card and a second DPU network card. The method includes: obtaining a user's aggregation operation on a first target port in a first DPU network card and a second target port in a second DPU network card. The first target port includes all or part of the physical ports of the first DPU network card, and the second target port includes all or part of the physical ports of the second DPU network card. In response to the aggregation operation, aggregation configuration information is sent to the first DPU network card and the second DPU network card, and the aggregation configuration information is used to associate the first target port and the second target port to aggregate into an aggregated link.

[0007] The cross-card link aggregation method provided by the present application firstly allows users to perform different aggregation operations through the aggregation device, thereby improving the user's ability to configure the network independently. Secondly, by negotiating the port communication of multiple DPU network cards through the aggregation device, cross-card aggregation of multiple DPUs can be achieved without destroying the independence of the DPU, solving the limitation of the prior art that multiple DPUs cannot perform link aggregation due to the independence of the DPU, and can improve the reliability of DPU operation in the multi-DPU scenario.

[0008] In a possible implementation manner, the aggregation configuration information includes aggregation mode information, and the aggregation mode information is used to indicate an aggregation mode for performing associated aggregation.

[0009] In another possible implementation, the aggregation mode includes an active-standby mode, in which the active-standby mode is a state where the active port in the aggregated link is in a forwarding state and the standby port in the aggregated link is in a silent state. The active port is the first target port or the second target port, and the standby port is a port other than the active port in the first target port and the second target port.

[0010] In another possible implementation, priority parameters of the first DPU network card and the second DPU network card are obtained, and based on the priority parameters of the first DPU network card and the second DPU network card, the primary port and the backup port are determined from the first target port and the second target port. The priority parameter of the network card where the primary port is located is higher than the priority parameter of the network card where the backup port is located.

[0011] In another possible implementation, the priority parameter is related to at least one of the following: a bandwidth corresponding to the DPU network card, a historical transmission delay of the DPU network card, a packet loss rate of the DPU network card, and a media access control MAC address of the DPU network card.

[0012] In another possible implementation, in response to the primary port being in the down state, the backup port is switched from the silent state to the forwarding state.

[0013] In another possible implementation, in response to the primary port recovering to the up state, the primary port is switched to the forwarding state, and the backup port is switched from the forwarding state to the silent state.

[0014] In another possible implementation, the aggregation mode includes a balanced mode. In the balanced mode, the first target port and the second target port in the aggregated link are both in a forwarding state.

[0015] Another possible implementation is to obtain a load factor, which is used to indicate the proportion of traffic carried by the DPU network card.

[0016] In a second aspect, the present application provides an inter-card link aggregation device, which includes various functional modules used in the method described in the first aspect above.

[0017] In a third aspect, the present application provides a computer program product, including: computer instructions; when the computer instructions are executed on an electronic device, the electronic device implements the method described in the first aspect above.

[0018] In a fourth aspect, the present application provides an electronic device, comprising: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device implements the method described in the first aspect above.

[0019] In a fifth aspect, the present application provides a readable storage medium, which includes: software instructions; when the software instructions are executed in an electronic device, the electronic device implements the method described in the first aspect above.

[0020] The beneficial effects of the second to fifth aspects mentioned above can be referred to the first aspect and will not be elaborated on again. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A schematic diagram of an application scenario of a cross-card link aggregation method provided in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of a flow chart of a cross-card link aggregation method provided in an embodiment of the present application;

[0024] Figure 3 A schematic diagram of a flow chart of a cross-card link aggregation method in a master-slave mode provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of a process flow of switching the fault state of a primary port in a primary-standby mode provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of a process flow of switching a backup port failure state in a primary-backup mode provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of a flow chart of a cross-card link aggregation method in a balanced mode provided in an embodiment of the present application;

[0028] Figure 7A schematic diagram of the structure of a cross-card link aggregation method execution module provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of a flow chart of a cross-card link aggregation method based on an execution module provided in an embodiment of the present application;

[0030] Fig. 9 A schematic diagram of a cross-card link aggregation device provided by the present application;

[0031] Fig.10 A schematic diagram of the composition of an electronic device provided in this application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0033] It should be noted that, in the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.

[0034] In order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical items or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.

[0035] The following is an explanation of the professional terms involved in the embodiments of the present application.

[0036] Data processing unit (DPU): is a hardware unit designed for data processing. Functionally, it can offload the data processing tasks of the CPU and focus on efficient data processing and transmission, such as taking on network and storage data processing and related complex operations, reducing the burden on the CPU to improve system performance. From an architectural perspective, the DPU has an independent operating system, including processor cores, cache, memory and dedicated hardware acceleration modules. It can work with the CPU, graphic processing unit (GPU), etc. Its internal architecture optimizes the data processing process and improves parallelism and throughput. In terms of application scenarios, DPU is widely used in data centers, cloud computing, 5G communications, intelligent networks and other fields, helping various scenarios to achieve efficient data processing and management.

[0037] Link aggregation (LAG): also known as link aggregation or port aggregation, is a technology that logically bundles multiple physical links into a link set. Through LAG, multiple physical links can work as a whole to achieve parallel transmission of data on these links, thereby increasing the total bandwidth of the link and improving data transmission efficiency. At the same time, it can also provide link redundancy backup functions to enhance network reliability and stability. It is widely used in scenarios such as enterprise networks and data centers to meet the growing network bandwidth requirements and high availability requirements.

[0038] Link aggregation control protocol (LACP): It is a protocol used to automatically negotiate, establish and manage link aggregation between network devices. By exchanging LACP protocol data units, it automatically selects available physical links to bundle into logical link aggregation groups based on device configuration and link status, thereby achieving functions such as increasing bandwidth, link redundancy backup and load balancing. LACP is a key protocol for implementing LAG, providing automatic negotiation and configuration capabilities for LAG to ensure that LAG can operate correctly. It is an indispensable part of the LAG technology system and collaborates with other parts to complete link aggregation tasks.

[0039] The above is an explanation of the professional terms involved in the embodiments of the present application.

[0040] As can be seen from the background art and related technical terms, a significant feature of the DPU is that it has an independent operating system and can perform data processing independently, reducing the computing pressure on the CPU. However, due to the independence of the DPU, in a multi-DPU scenario, the DPUs cannot achieve link aggregation like the network interfaces on traditional switches or servers, such as the Link Aggregation Control Protocol (LACP). In view of this, how to overcome the link aggregation problem caused by the independent operating system of the DPU is an urgent problem to be solved when the data center adopts the DPU technology.

[0041] Based on this, the cross-card link aggregation method provided in this application can negotiate the aggregation information configuration among multiple DPUs based on the aggregation device, acting as a business communication bridge between the DPUs to achieve link aggregation among multiple DPUs and improve the reliability of the DPU in a multi-DPU scenario.

[0042] A cross-card link aggregation method provided in an embodiment of this application can be applied to a scenario as Figure 1 shown. In this scenario, it includes: an aggregation device 110, a DPU network card 120, a DPU network card 130, and a network device 140, where:

[0043] The aggregation device 110 is communicatively connected to the DPU network card 120 and the DPU network card 130.

[0044] The DPU network card 120 and the DPU network card 130 are physically connected to the network device 140.

[0045] The aggregation device 110 is used to obtain the aggregation operation of the user on the first target port in the DPU network card 120 and the second target port in the DPU network card 130.

[0046] The aggregation device 110 is further used to send aggregation configuration information to the DPU network card 120 and the DPU network card 130 in response to the aggregation operation.

[0047] In some embodiments, the aggregation device 110 can be a CPU, which is deployed in a computing device. Relying on the computing resources and operating environment of the computing device, it can quickly respond to and parse the user's aggregation operation. At the same time, based on the communication connection between the computing device and the DPU network card, the parsed user aggregation operation can be sent to the associated DPU network card.

[0048] The DPU network card 120 and the DPU network card 130 are used to control the ports to perform network communication based on the aggregation configuration information sent by the aggregation device 110.

[0049] In some embodiments, the aggregation device may be connected to the DPU via a PCle (peripheral component interconnect express) bus, and the aggregation device may send aggregation configuration information to the first DPU network card and the second DPU network card via the PCle bus.

[0050] In some embodiments, the mainboard where the convergence device and the DPU network card are located needs to be equipped with an optical module and an optical fiber interface respectively, and connected through a single optical fiber cable. When sending the convergence configuration information, the optical module of the convergence device converts the convergence configuration information from an electrical signal to an optical signal, and sends it to the optical module of the DPU mainboard through the optical fiber. The optical module converts the optical signal to an electrical signal and sends it to the DPU network card.

[0051] The DPU network card 120 and the DPU network card 130 are also used to report priority parameters and port status to the aggregation device.

[0052] It should be noted that the aggregation device 110, the DPU network card 120, and the DPU network card 130 can be deployed in different computing devices respectively.

[0053] Specifically, the computing device may be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The embodiments of the present application do not limit the specific device form of the computing device.

[0054] It should be understood that Figure 1 The scenario shown does not constitute a limitation. In actual application, the scenario may include more or fewer devices, such as 3 or more DPU network cards, and the embodiments of the present application do not impose any limitation on this.

[0055] The cross-card link aggregation method provided in the embodiment of the present application can be used in the above-mentioned aggregation device, such as Figure 2 As shown, the method may include S101-S102:

[0056] S101. Obtain a user's aggregation operation on a first target port in a first DPU network card and a second target port in a second DPU network card.

[0057] The first target port includes all or part of the physical ports of the first DPU network card, and the second target port includes all or part of the physical ports of the second DPU network card.

[0058] Specifically, the aggregation operation may include: configuring the association between the first target port in the first DPU network card and the second target port in the second DPU network card, and configuring the aggregated link.

[0059] In some embodiments, the aggregation device can also be used to respond to the user's display instructions, obtain and return the port information and device performance parameters of multiple DPU network cards, so as to facilitate the user to understand the device information and status of the DPU network card.

[0060] In some embodiments, the aggregation operation may be an aggregation command input by a user in a command line interface. The aggregation device associates the first target port in the first DPU network card with the second target port in the second DPU network card by parsing the aggregation command.

[0061] In other embodiments, the aggregation operation may also be an interactive action between the user and the graphical interface, and the aggregation device associates the first target port in the first DPU network card with the second target port in the second DPU network card by parsing the interactive action. In this case, the user may also use a preset aggregation operation template in the graphical interface and import it into the aggregation device with one click.

[0062] Exemplarily, the user performs a convergence operation in the command line interface, that is, inputs the convergence command: dpu bond createbond1 port eth1 / 1eth2 / 1mode active-up ethtool speed 1000, wherein dpu bond create indicates the creation of a DPU aggregate link bond1, port eth1 / 1eth2 / 1 indicates the first port eth1 / 1 of the first network card DPU and the second port eth2 / 1 of the second network card DPU to which bond1 is to be associated, mode active up indicates that the convergence mode of bond1 is the primary / standby mode (for details, see below), and ethtool speed 1000 indicates setting the bandwidth of the aggregate link bond1 to 1000Mbps. Afterwards, the convergence device will parse the DPU link convergence command input by the user and perform subsequent operations. The example code provided in this embodiment is only used for understanding, and the specific code of the DPU link convergence command is not limited.

[0063] S102: In response to the aggregation operation, aggregation configuration information is sent to the first DPU network card and the second DPU network card.

[0064] The aggregation configuration information is used to associate and aggregate the first target port and the second target port into an aggregated link.

[0065] Specifically, the aggregation device can obtain the aggregation parameters by parsing the obtained user aggregation operation. The aggregation parameters are the configuration parameters for the aggregated link contained in the user aggregation operation. The aggregation parameters are checked for integrity and legality to detect whether the DPU port information in the aggregation parameters exists and whether the aggregated link information in the aggregation parameters complies with the network protocol. On the premise that the aggregation parameters are determined to be legal, the aggregation parameters are encapsulated into aggregation configuration information and sent to the target DPU indicated by the aggregation parameters.

[0066] In some embodiments, the aggregation configuration information includes aggregation mode information, and the aggregation mode information is used to indicate an aggregation mode for performing association aggregation.

[0067] Specifically, the convergence mode may include an active / standby mode and a balanced mode. For a detailed introduction to the convergence mode, please refer to the following.

[0068] In some embodiments, after receiving the aggregation configuration information, the first DPU network card and the second DPU network card perform an aggregation link configuration operation based on the information. The configuration operation may include: determining the target port based on the aggregation configuration information, configuring the target port parameters to support the LACP protocol, and setting the port bandwidth parameters of the target port based on the aggregation link configuration parameters in the aggregation configuration information.

[0069] For example, in the case of configuring the target port parameters to support the LACP protocol, first, the DPU network card configures the first target port eth1 / 1 indicated by the aggregation configuration information to support the LACP protocol. It will change the link layer protocol mode of port eth1 / 1 from the default mode to the LACP mode, and then set the port's LACP key for security verification. At the same time, the LACP short time interval timer is set to 1 second to speed up the negotiation speed when the link status changes.

[0070] For another example, for setting the port bandwidth parameters of the target port based on the aggregation link configuration parameters in the aggregation configuration information, assume that the aggregation configuration information indicates that the total bandwidth of the aggregation link composed of the port eth1 / 1 of the first DPU network card and the port eth2 / 1 of the second DPU network card is 15Gbps, and port eth1 / 1 is required to bear 7Gbps bandwidth and port eth2 / 1 is required to bear 8Gbps bandwidth. The first DPU network card will adjust the bandwidth control parameters of port eth1 / 1, setting the upper bandwidth limit of both the inbound and outbound directions to 7Gbps. The second DPU network card performs similar operations on port eth2 / 1, setting its upper bandwidth limit of both the inbound and outbound directions to 8Gbps, and ensures that the port operates stably within the specified bandwidth through speed limiting and traffic shaping mechanisms.

[0071] It can be seen from S101-S102 that the cross-card link aggregation method effectively solves the problem that multiple DPUs cannot perform link aggregation due to the independence of DPUs. By obtaining the user's aggregation operation on the target ports of the first DPU network card and the second DPU network card and sending configuration information, flexible aggregation of ports between different DPU network cards is achieved, thereby improving the reliability of DPU network cards in multi-DPU network card scenarios.

[0072] The following is a detailed introduction to the link aggregation mode when it is in active / standby mode.

[0073] In some embodiments, the aggregation configuration information in step S102 also includes an aggregation mode. When the aggregation mode is an active-standby mode, the active port of the aggregated link is in a forwarding state, and the standby port of the aggregated link is in a silent state.

[0074] The primary port is the first target port or the second target port, and the backup port is a port other than the primary port among the first target port and the second target port.

[0075] The port of the DPU corresponds to a physical link. Correspondingly, for the main port, the corresponding link can also be called the main link, and for the backup port, the corresponding link can be called the backup link.

[0076] Specifically, in the link aggregation in the active-standby mode, the links participating in the aggregation will be divided into the active link and the standby link. The active link undertakes the main data transmission task and is the main channel for data traffic. The standby link is in standby status and does not participate in data transmission when the active link is operating normally. However, it is always ready to take over the work of the active link when a failure occurs, to ensure that network communication is not interrupted.

[0077] In some embodiments, in the active-standby mode, the process of determining the active port and the standby port in the aggregated link is as follows: Figure 3 As shown, after S102, S103 and S104 may also be included:

[0078] S103: Obtain a priority parameter of the first DPU network card and a priority parameter of the second DPU network card.

[0079] In some embodiments, the priority parameter is related to at least one of the following: a bandwidth corresponding to the DPU network card, a historical transmission delay of the DPU network card, a packet loss rate of the DPU network card, and a media access control MAC address of the DPU network card.

[0080] Specifically, the explanation of the above related factors:

[0081] The bandwidth corresponding to the DPU network card refers to the data transmission rate that the DPU network card can carry. The higher the bandwidth, the larger the amount of data that can be transmitted per unit time, and the stronger the network transmission capacity. It is usually measured in Mbps (megabits per second) or Gbps (gigabits per second).

[0082] The historical transmission delay of the DPU network card, that is, the time it takes for data to be transmitted from the sender to the receiver via the DPU network card, reflects the timeliness of data transmission. The smaller the delay, the faster the data transmission. It is generally measured in milliseconds (ms).

[0083] The packet loss rate of the DPU network card refers to the proportion of data packets lost by the DPU network card during data transmission. The lower the packet loss rate, the better the integrity and reliability of data transmission. It is usually expressed as a percentage.

[0084] The DPU network card media access control (MAC) address is the physical address of the network card, which is globally unique and is used to identify the DPU network card in the network. After it is mapped to the range of 0-1 according to the numerical size, it can be used as a factor in calculating the priority parameter. The smaller the MAC address, the higher the priority.

[0085] It should be noted that the DPU network card MAC address is mapped to a 0-1 value for priority calculation because its uniqueness can distinguish priorities when multiple DPU network cards have similar performance. It can also assist in network planning and load balancing, and its randomness ensures fair resource competition. It is also compatible with the existing architecture and adapts to network expansion, so the new DPU network card MAC can naturally participate.

[0086] It should also be noted that when calculating the DPU network card priority parameters, the bandwidth is positively correlated with the priority parameters, and the historical transmission delay, packet loss rate, and MAC address comprehensive score are negatively correlated with the priority parameters.

[0087] In some embodiments, the priority parameter may be obtained by parsing a priority configuration input by a user in a user aggregation operation, or the priority parameter of the DPU network card may be obtained from feedback information of the DPU network card.

[0088] In one possible implementation, the priority parameter is obtained from the feedback of the DPU network card. In this case, when the DPU network card receives the aggregation configuration information of the aggregation device, if the information instructs the DPU network card to calculate the priority parameter, the DPU network card will calculate the priority parameter based on the above-mentioned related factors and feedback it to the aggregation device.

[0089] For example, the priority parameter is calculated based on the MAC address of the DPU network card. Since the hexadecimal range of each byte is 00-FF (the decimal range is 0-255), the value of each byte can be mapped to 0-100 to obtain the score of the byte. Assuming that the MAC address of the first DPU network card is 00:11:22:33:44:55, the score mapping calculation is performed on each byte: the score corresponding to the first byte 00 is 0÷255×100=0 points, and the score corresponding to the second byte 11 (decimal is 17) The corresponding score is 17÷255×100≈6.67 points, the third byte 22 (decimal 34) corresponds to a score of 34÷255×100=13.33 points, the fourth byte 33 (decimal 51) corresponds to a score of 51÷255×100=20 points, the fifth byte 44 (decimal 68) corresponds to a score of 68÷255×100≈26.67 points, and the sixth byte 55 (decimal 85) corresponds to a score of 85÷255×100≈33.33 points. Then find the average score of all bytes (0+6.67+13.33+20+26.67+33.33)÷6≈15 points. In the absence of other relevant factors involved in the priority parameter calculation, the priority parameter of the first network card DPU is 15. The smaller the score, the greater the priority parameter of the DPU network card.

[0090] For another example, the priority parameters of the DPU network card are calculated based on the above-mentioned related factors by weighted summation method. Assume that the bandwidth weight is 0.4, the standard score is 0-100 points (2000Mbps corresponds to 100 points), the historical transmission delay weight is 0.3, the standard score is 0-100 points (5ms corresponds to 100 points), the packet loss rate weight is 0.2, the standard score is 0-100 points (0% corresponds to 100 points), and the MAC address comprehensive score weight is 0.1. The first DPU network card has a bandwidth of 1500Mbps, a historical transmission delay of 8ms, a packet loss rate of 0.4%, and a MAC address comprehensive score of 15. Its bandwidth standard score is 1500÷2000×100=75 points, its historical transmission delay standard score is 5÷8×100=62.5 points, and its packet loss rate standard score is (1-0.004)×100=99.6 points. The priority parameter of the DPU network card is 75×0.4+62.5×0.3+99.6×0.2+15×0.1=70.47. The larger the score, the greater the priority parameter of the DPU network card.

[0091] In another possible implementation, the priority parameters of the DPU network card can also be directly configured by the user. In this case, the aggregation device will obtain the priority parameters of each DPU network card by parsing the user's aggregation operation, and instruct the DPU network card not to calculate the priority parameters in the aggregation configuration information sent to the DPU network card.

[0092] S104: Determine a primary port and a backup port from the first target port and the second target port based on the priority parameter of the first DPU network card and the priority parameter of the second DPU network card.

[0093] The priority parameter of the network card where the primary port is located is higher than the priority parameter of the network card where the backup port is located.

[0094] In some embodiments, if the first DPU network card or the second DPU network card in the aggregation configuration information includes multiple target ports, and in the link aggregation in the active-standby mode, only one master port is set, it is necessary to compare the priority parameters of multiple target ports of the multiple DPU network cards. The target port with the highest priority parameter is selected as the primary port. The other target ports are used as standby ports.

[0095] It can be seen from steps S103-S104 that when the aggregation mode is the primary-standby mode, by incorporating factors such as the processing capacity, bandwidth, forwarding rate, and delay of the network card into the calculation of the DPU network card priority parameter, the performance of the DPU network card can be reflected from the hardware level. Selecting the DPU network card port with a higher priority parameter as the primary port can give full play to the performance advantage of the DPU network card.

[0096] In some embodiments, to ensure the stability of the aggregated link, the DPU network card also needs to periodically detect the target port (i.e., the port bound to the aggregated link), and determine the monitoring status of the target port based on the comparison between the periodic detection data and the monitoring parameter threshold.

[0097] Specifically, the monitoring parameter threshold is related to at least one of the following: a bandwidth utilization threshold, a packet loss rate threshold, an error frame ratio threshold, and a port response time threshold.

[0098] Specifically, the monitoring state of the target port can be divided into a down state and an up state. The down state indicates that the periodic detection data of the target port is less than the monitoring parameter threshold, whereas the up state indicates that the periodic detection data of the target port is greater than or equal to the monitoring parameter threshold.

[0099] It should be noted that the monitoring parameter threshold may be obtained by user configuration or by the default configuration of the aggregation device or the DPU network card.

[0100] It should also be noted that when the DPU network card periodically detects the target port, if the monitoring status of the target port changes, it needs to send its monitoring status to the aggregation device. If the monitoring status of the target port does not change, it may not send its monitoring status to the aggregation device.

[0101] In some embodiments, when the aggregation mode in the aggregation configuration information is the master-slave mode, when the master port is in the forwarding state, it may be detected as the down state. In this case, Figure 4 As shown, after S104 in the method, S105 is further included:

[0102] S105 . In response to the primary port being in the down state, switch the backup port from the silent state to the forwarding state.

[0103] Specifically, when the DPU network card detects that the monitoring state of the main port of the aggregation link is down, the monitoring state of the port is sent to the aggregation device, and the aggregation device switches the backup port from the silent state to the forwarding state based on the monitoring state of the main port. At the same time, the main port is switched from the forwarding state to the silent state.

[0104] In some embodiments, when the aggregation mode in the aggregation configuration information is in the master-slave mode, when the master port is in the silent state, it can be restored to the up state within a certain period of time. When the aggregation device receives that the monitoring state of the master port has changed to the up state, it switches the state of the master port to the forwarding state again. Figure 5 As shown, after step S105, the method further includes:

[0105] S106 , in response to the primary port recovering to the up state, switching the primary port to the forwarding state, and switching the backup port from the forwarding state to the silent state.

[0106] In other embodiments, when the main port is restored to the up state, the aggregation device may not switch the main port to the forwarding state and maintain the forwarding state of the backup port. When receiving the status of the backup port as down, the backup port is switched to the silent state and the main port is switched to the forwarding state.

[0107] It can be seen from the above steps S104-S106 that through the periodic detection of the DPU network card, the aggregation device can timely perceive the changes in the status of the main port, switch the main port in the down state to the silent state, and switch the backup port to the forwarding state, thereby realizing redundant backup of multiple DPU network card links and ensuring the stability of the network.

[0108] The following is a detailed introduction to the link aggregation mode in balanced mode.

[0109] In some embodiments, the aggregation mode further includes a balanced mode. When the aggregation mode is in the balanced mode, the first target port and the second target port in the aggregated link are both in a forwarding state.

[0110] It should be noted that in the balanced mode, the aggregation device ensures the consistency of the traffic input and output through the target port, that is, the traffic input and output need to pass through the port of the same DPU network card. In this case, the aggregation device also needs to coordinate the consistency of traffic input and output based on the hash strategy.

[0111] The hash strategy is a key means to achieve load balancing. It takes the source IP, destination IP, source port, destination port, protocol number, and any combination of them as input, and sends them to the hash function for calculation to generate a hash value of fixed length. Ideally, these hash values ​​should be evenly distributed, that is, different inputs can be mapped to different hash values ​​as evenly as possible. And through modulo operations, the hash values ​​are matched with the number of network cards or links, and evenly distributed to the output path, to achieve load balancing of traffic on different network cards or links, and ensure that the input and output of traffic with the same characteristics are consistent.

[0112] Specifically, the hash strategy may include the following methods: source IP hash, destination IP hash, source IP and destination IP hash, source port hash, destination port hash, and five-tuple hash.

[0113] In some embodiments, the specific method of the hash strategy can be configured by the user, obtained and parsed by the aggregation device, and applied to the aggregation link in the balanced mode.

[0114] In some embodiments, when the aggregation mode is a balanced mode, it is necessary to determine the ratio of the traffic carried by multiple target ports. In this case, the aggregation mode needs to determine a load factor for adjusting the traffic distribution ratio of the target ports in the aggregated link, such as Figure 6 As shown, after method S102, the method further includes S103a:

[0115] S103a. Obtain a load factor.

[0116] The load factor is used to indicate the proportion of the DPU network card carrying links.

[0117] In some embodiments, the aggregation device can parse the load factor configured by the user from the user's aggregation operation. In this case, the aggregation device can directly use the load factor for the bearer link allocation of the target port in the balancing mode.

[0118] In other embodiments, the aggregation device may obtain the load factor through the performance data of the DPU network card. The performance data may be a priority parameter fed back by the DPU network card. In this case, the load factor obtained through the priority parameter may be used for the initial bearer traffic allocation of the aggregated link. The performance data may also be periodic detection data of the DPU network card. In this case, the load factor obtained through the periodic detection data may be used for the periodic adjustment of the target port bearer traffic allocation during the forwarding process of the aggregated link.

[0119] Exemplarily, in the balanced mode, the load factor of the aggregated link is obtained by the priority parameter. Assume that the priority parameter value fed back by the first DPU network card obtained by the aggregation device is 80, and the priority parameter value fed back by the second DPU network card is 85. The process of normalizing the two is: 80 / (80+60)×100%≈57%, 60 / (80+60)×100%≈43%, that is, {57%, 43%} is the load factor of each target port in the balanced mode.

[0120] It should be understood that when the aggregation mode in the aggregation configuration information is balanced mode, the aggregated link can distribute network traffic to the ports of multiple DPU network cards, making full use of the link resources of each DPU and improving network resource utilization. It can also allow data to be transmitted in parallel on multiple links of the DPU, improving network response speed and transmission efficiency, and automatically switching traffic when a link fails to ensure network stability and reliability.

[0121] Figure 7 A schematic diagram of a cross-card link aggregation method execution module provided in an embodiment of the present application. Figure 7 As shown, the architecture includes a D-LAG module 110, a D-LAG module 120, and a D-LAG agent module 130. The D-LAG module 110 is deployed in the first DPU network card, the D-LAG module 120 is deployed in the second DPU network card, and the D-LAG agent module 130 is deployed in the aggregation device.

[0122] The DPU network card runs the D-LAG module to negotiate parameters and send information with the D-LAG agent module of the aggregation device.

[0123] Specifically, the D-LAG agent module 130 may perform the following tasks:

[0124] LAG-config_db task: responsible for responding to the user's aggregation operation, parsing the aggregation operation into aggregation configuration information, and sending the aggregation configuration information to the DPU network card.

[0125] LAG-hash_control task: responsible for switching the forwarding state and silent state of the aggregation link port, and also responsible for running the hash policy to ensure the consistency of the input and output traffic of the aggregation link port.

[0126] Specifically, the D-LAG module includes at least the following tasks:

[0127] LAG-link status task: responsible for periodically detecting the status of the DPU network card port.

[0128] LAG-syn task: responsible for synchronously sending the status of the DPU network card port to the aggregation device, and receiving the aggregation configuration information sent by the aggregation device.

[0129] Figure 7 The execution module describes the operation process of the cross-card aggregation link system from a functional perspective. Furthermore, based on the above-mentioned D-LAG module and D-LAG agent module, the cross-card link aggregation method is introduced.

[0130] Figure 8 A flow chart of a cross-card link aggregation method based on an execution module provided in an embodiment of the present application is provided. In this method, a CPU is used as an aggregation device to negotiate the link aggregation of multiple DPU network cards, such as Figure 8 As shown in the figure, the process specifically includes:

[0131] S801. The CPU responds to a user aggregation operation.

[0132] The CPU starts the LAG-config_db task to parse the user aggregation operation and generate aggregation configuration information, including: the aggregation mode of the aggregated link is the active / standby mode, the aggregated link is associated with port port1 / 1 of DPU network card 1, and the aggregated link is associated with port port2 / 1 of DPU network card 2.

[0133] S802. The CPU sends aggregation configuration information.

[0134] The LAG-config_db module of the CPU is also used to send the aggregation configuration information to DPU network card 1 and DPU network card 2.

[0135] S803. DPU network card 1 and DPU network card 2 configure ports and feedback priority parameters.

[0136] Taking DPU NIC 1 as an example, DPU NIC 1 configures the parameters of port 1 / 1 based on the received aggregation configuration information, and starts the LAG-link_status task to periodically detect port 1 / 1. At the same time, based on its own performance parameters and the performance parameters of port 1 / 1, it calculates the priority parameters of the port and starts the LAG-syn task to feed back to the CPU.

[0137] For the specific operation of DPU network card 2, refer to DPU network card 1.

[0138] S804: The CPU determines a primary port and a backup port of the aggregated link based on the priority parameter.

[0139] The CPU starts the LAG-hash_control task, first compares the priority parameters fed back by DPU NIC 1 and DPU NIC 2, and then determines that port 1 / 1 of DPU NIC 1 with a larger priority parameter is used as the primary port, and port 2 / 1 of DPU NIC 2 is used as the backup port. The primary port is set to the forwarding state, and the backup port is set to the silent state.

[0140] S805: The CPU receives the down state of the primary port and switches the state of the backup port to the forwarding state.

[0141] When DPU network card 1 detects that the status of the main port port1 / 1 becomes down, it starts the LAG-syn task to send a monitoring status report to the CPU. Based on the monitoring status report of DPU network card 1, the CPU switches the main port port1 / 1 from the forwarding state to the silent state, and switches the backup port port2 / 1 of DPU network card 2 to the forwarding state.

[0142] S806: The CPU receives the up state of the primary port and maintains the state of the existing port.

[0143] At this time, although the monitoring state of the primary port is changed back to the up state, no monitoring state change information of the backup port is received, so the existing states of the primary and backup ports are maintained.

[0144] S807: The CPU receives the down state of the standby port and switches the state of the primary port to the forwarding state.

[0145] When DPU network card 2 detects that the status of the backup port port2 / 1 becomes down, it starts the LAG-syn task to send a monitoring status report to the CPU. Based on the monitoring status report of DPU network card 2, the CPU switches the primary port port1 / 1 from the silent state to the forwarding state, and switches the backup port port2 / 1 of DPU network card 2 to the silent state.

[0146] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed in this article, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0147] In an exemplary embodiment, the present application also provides a cross-card link aggregation device, such as Fig. 9 As shown, the inter-card link aggregation device includes: an acquisition module 910 and a processing module 920 .

[0148] The acquisition module 910 is used to obtain the user's aggregation operation on the first target port in the first DPU network card and the second target port in the second DPU network card. The first target port includes all or part of the physical ports of the first DPU network card, and the second target port includes all or part of the physical ports of the second DPU network card.

[0149] The processing module 920 is used to send aggregation configuration information to the first DPU network card and the second DPU network card in response to the aggregation operation, where the aggregation configuration information is used to associate the first target port and the second target port to aggregate into an aggregate link.

[0150] In a possible implementation manner, the aggregation configuration information includes aggregation mode information, and the aggregation mode information is used to indicate an aggregation mode for performing associated aggregation.

[0151] In another possible implementation, the aggregation mode includes an active-standby mode. In the active-standby mode, the primary port in the aggregated link is in a forwarding state, the backup port in the aggregated link is in a silent state, the primary port is the first target port or the second target port, and the backup port is a port other than the primary port between the first target port and the second target port.

[0152] In another possible implementation, in the active-standby mode, the processing module 920 is further used to: obtain the priority parameter of the first DPU network card and the priority parameter of the second DPU network card. Based on the priority parameter of the first DPU network card and the priority parameter of the second DPU network card, the primary port and the backup port are determined from the first target port and the second target port, and the priority parameter of the network card where the primary port is located is higher than the priority parameter of the network card where the backup port is located.

[0153] In another possible implementation, the priority parameter is related to at least one of the following: a bandwidth corresponding to the DPU network card, a historical transmission delay of the DPU network card, a packet loss rate of the DPU network card, and a media access control MAC address of the DPU network card.

[0154] In another possible implementation, in the active-standby mode, the processing module 920 is further configured to: in response to the active port being in the down state, switch the standby port from the silent state to the forwarding state.

[0155] In another possible implementation, in the active-standby mode, the processing module 920 is further configured to: in response to the active port recovering to the up state, switch the active port to the forwarding state, and switch the standby port from the forwarding state to the silent state.

[0156] In another possible implementation, the aggregation mode includes a balanced mode. In the balanced mode, the first target port and the second target port in the aggregated link are both in a forwarding state.

[0157] In another possible implementation, in the balancing mode, the processing module 920 is further used to: obtain a load factor, where the load factor is used to indicate the proportion of the traffic carried by the DPU network card.

[0158] It should be noted that Fig. 9 The division of modules in the example is schematic and is only a logical function division. There may be other division methods in actual implementation. For example, two or more functions may be integrated into one processing module. The above integrated modules may be implemented in the form of hardware or software function modules.

[0159] In an exemplary embodiment, as described above, the computing device may be a computer or a server or other electronic device having a computing function. In this case, the present application also provides an electronic device, Fig.10 The following is a schematic diagram of the composition of the electronic device provided in the embodiment of the present application. Fig.10 As shown, the electronic device includes: a processor 10 , a memory 20 , a communication line 30 , a communication interface 40 , and an input / output interface 50 .

[0160] The processor 10 , the memory 20 , the communication interface 40 and the input / output interface 50 may be connected via a communication line 30 .

[0161] The processor 10 is used to execute the instructions stored in the memory 20 to implement the cross-card link aggregation method provided in the above embodiment of the present application. The processor 10 can be a CPU, a general-purpose processor network processor (network processor, NP), a digital signal processor (digital signal processing, DSP), a microprocessor, a microcontroller (micro control unit, MCU) / single-chip microcomputer / single-chip microcomputer, a programmable logic device (programmable logic device, PLD) or any combination thereof. The processor 10 can also be any other device with processing functions, such as a circuit, a device or a software module, which is not limited in the embodiments of the present application. In one example, the processor 10 may include one or more CPUs, such as Fig.10 As an optional implementation, the electronic device may include multiple processors, for example, in addition to the processor 10, it may also include a processor 60 ( Fig.10 The dashed line is used as an example.

[0162] The memory 20 is used to store instructions. For example, the instructions may be computer programs. Optionally, the memory 20 may be a read-only memory (ROM) or other types of static storage devices that can store static information and / or instructions, or a random access memory (RAM) or other types of dynamic storage devices that can store information and / or instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage devices, etc., and the embodiments of the present application are not limited to this.

[0163] It should be noted that the memory 20 may exist independently of the processor 10, or may be integrated with the processor 10. The memory 20 may be located inside the electronic device, or may be located outside the electronic device, which is not limited in the embodiment of the present application.

[0164] The communication line 30 is used to transmit information between various components included in the electronic device.

[0165] The communication interface 40 is used to communicate with other devices or other communication networks. The other communication networks may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The communication interface 40 may be a module, a circuit, a transceiver or any device capable of achieving communication.

[0166] The input / output interface 50 is used to implement human-computer interaction between a user and an electronic device, for example, to implement action interaction or information interaction between a user and an electronic device.

[0167] Exemplarily, the input / output interface 50 may be a mouse, a keyboard, a display screen, or a touch display screen, etc. Action interaction or information interaction between a user and an electronic device may be achieved through a mouse, a keyboard, a display screen, or a touch display screen, etc.

[0168] It should be noted that Fig.10 The structure shown in the figure does not constitute a limitation on the electronic device, except Fig.10 In addition to the components shown, the electronic device may include more or fewer components than shown, or a combination of certain components, or a different arrangement of components.

[0169] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed in an electronic device, the electronic device implements the method in the aforementioned method embodiment.

[0170] In an exemplary embodiment, the present application also provides a readable storage medium, which includes software instructions. When the software instructions are executed in an electronic device, the electronic device implements the method in the aforementioned method embodiment. The computer-readable storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device.

[0171] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer-executable instructions. When loading and executing computer-executable instructions on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-executable instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer-executable instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0172] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other changes to the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0173] Although the present application has been described in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are deemed to have covered any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

[0174] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A cross-card link aggregation method, characterized in that: The method is applied to a convergence device; The convergence device is in communication connection with the first data processing unit DPU network card and the second DPU network card; the method comprises: Obtaining a user's aggregation operation on a first target port in the first DPU network card and a second target port in the second DPU network card; The first target port includes all or part of the physical ports of the first DPU network card; the second target port includes all or part of the physical ports of the second DPU network card; In response to the aggregation operation, aggregation configuration information is sent to the first DPU network card and the second DPU network card; the aggregation configuration information is used to associate the first target port and the second target port to aggregate into an aggregated link.

2. The method according to claim 1, characterized in that The aggregation configuration information includes aggregation mode information, and the aggregation mode information is used to indicate an aggregation mode for performing associated aggregation.

3. The method according to claim 2, characterized in that The aggregation mode includes an active-standby mode; in the active-standby mode, the active port in the aggregated link is in a forwarding state, and the standby port in the aggregated link is in a silent state; the active port is the first target port or the second target port, and the standby port is a port of the first target port and the second target port other than the active port.

4. The method according to claim 3, characterized in that The method further comprises: Obtaining a priority parameter of the first DPU network card and a priority parameter of the second DPU network card; Based on the priority parameter of the first DPU network card and the priority parameter of the second DPU network card, the primary port and the backup port are determined from the first target port and the second target port; the priority parameter of the network card where the primary port is located is higher than the priority parameter of the network card where the backup port is located.

5. The method according to claim 4, characterized in that The priority parameter is related to at least one of the following: a bandwidth corresponding to the DPU network card, a historical transmission delay of the DPU network card, a packet loss rate of the DPU network card, and a media access control MAC address of the DPU network card.

6. The method according to claim 3, characterized in that The method further comprises: In response to the primary port being in the down state, the backup port is switched from the silent state to the forwarding state.

7. The method according to claim 6, characterized in that The method further comprises: In response to the primary port recovering to the up state, the primary port is switched to the forwarding state, and the backup port is switched from the forwarding state to the silent state.

8. The method according to claim 2, characterized in that: The aggregation mode includes a balanced mode; in the balanced mode, the first target port and the second target port in the aggregated link are both in a forwarding state.

9. The method according to claim 8, characterized in that The method further comprises: Obtain a load factor; the load factor is used to indicate the proportion of traffic carried by the DPU network card.

10. A cross-card link aggregation device, characterized in that: The device comprises: an acquisition module and a processing module; The acquisition module is used to acquire the user's aggregation operation on the first target port in the first DPU network card and the second target port in the second DPU network card; The processing module is used to send aggregation configuration information to the first DPU network card and the second DPU network card in response to the aggregation operation.

11. An electronic device, characterized in that: include: Processor and memory; The memory stores instructions executable by the processor; When the processor is configured to execute the instructions, the electronic device implements the method according to any one of claims 1 to 9.

12. A readable storage medium, characterized in that: include: Software instructions; When the software instructions are executed in an electronic device, the electronic device implements the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that include: Computer instructions; When the computer instructions are executed in an electronic device, the electronic device is enabled to implement the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Deployment method and device for hardware link aggregation, chip, network card and equipment

    CN121334040A