Communication method, communication device, network adapter and network switching equipment
By splitting the interfaces of the network card/accelerator card and the switching equipment, the traffic of the faulty link is switched to other normal links, thereby achieving stability and efficiency in communication between accelerator cards. This solves the problems of resource waste and increased power consumption caused by interface failures in existing technologies and meets the high-efficiency communication requirements of AI accelerator cards.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-28
AI Technical Summary
In existing multi-accelerator card communication solutions, when one link of the network card or accelerator card interface fails, the entire interface will fail, affecting accelerator card communication. Furthermore, configuring redundant interfaces leads to resource waste and increased power consumption, which cannot meet the high-efficiency communication requirements of AI accelerator cards.
By splitting the interconnection interface between the network card/accelerator card and the switching device, link resources can be managed in a fine-grained manner. When a link within an interface fails, the entire interface does not need to be interrupted. Traffic is forwarded through other normal links to ensure uninterrupted communication. Furthermore, the status of the network adapter and the switching device is synchronized through fault notification messages.
Maximize the use of existing link resources, avoid resource waste and increased power consumption, improve the availability and stability of inter-card communication, and meet the high-efficiency communication requirements of AI distributed training.
Smart Images

Figure CN121940264A_ABST
Abstract
Description
Technical Field
[0001] This application relates to communication technology, and more particularly to a communication method, communication device, network adapter, and network switching equipment. Background Technology
[0002] Existing multi-accelerator card inter-card communication solutions have the following drawbacks: when a link (e.g., tx0) of a network card interface or accelerator card interface fails, it will cause the entire interface using that link to fail, affecting network card / accelerator card communication. Summary of the Invention
[0003] This application provides a communication method, a communication device, a network adapter, and a network switching device.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a communication method applied to a network adapter, the method comprising: If it is determined that the first sub-interface of the network adapter has failed, the state of the first sub-interface is set to a fault state. A fault notification message is sent to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0005] This application provides a communication method applied to a network switching device, the method comprising: Receive a fault notification message sent by a network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed; Based on the identification information and the fault information, the state of the third sub-interface of the network switching device is set to a fault state, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0006] This application provides a communication device applied to a network adapter, the device comprising: The first processing unit is configured to determine that the first sub-interface of the network adapter has failed, and set the state of the first sub-interface to a fault state. The first sending unit is configured to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0007] This application provides a communication device applied to a network switching device, the device comprising: The second receiving unit is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed. The second processing unit is used to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and the fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0008] This application provides a network adapter, including a first communication interface and a first processor; wherein, The first processor is configured to determine that the first sub-interface of the network adapter has failed, and set the state of the first sub-interface to a fault state; The first communication interface is used to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of its third sub-interface to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface. This application provides a network switching device, including a second communication interface and a second processor; wherein... The second communication interface is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed; The second processor is configured to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and the fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0009] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the communication provided in this application when executed by a processor.
[0010] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the communication provided in this application.
[0011] The embodiments of this application have the following beneficial effects: It determines that the first sub-interface of the network adapter has failed and sets the state of the first sub-interface of the network adapter to a failed state; it sends a fault notification message to the network switching device through the second sub-interface of the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the failed network adapter, and the fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a failed state, and the third sub-interface of the network switching device has a one-to-one mapping relationship with the first sub-interface of the network adapter; this application quickly switches the service traffic that originally passed through the failed sub-interface to other normal sub-interfaces, maximizing service continuity, and sends the fault notification message through the second sub-interface, i.e., the normal sub-interface, accurately carrying key fault information, ensuring that the network switching device can quickly locate the corresponding link and complete state synchronization. Attached Figure Description
[0012] Figure 1 This is a first flowchart illustrating the communication method provided in an embodiment of this application; Figure 2 This is a second flowchart illustrating the communication method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the interconnection between the accelerator card and the switching equipment provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of a communication device applied to a network adapter according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a communication device applied to a network switching equipment according to an embodiment of this application; Figure 6 This is a schematic diagram of the network adapter provided in an embodiment of this application; Figure 7 This is a schematic diagram of the network switching device provided in the embodiments of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0016] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0017] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0018] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0019] Existing multi-accelerator card communication solutions have the following drawbacks: When a link (such as tx0) of a network interface card or accelerator card interface fails, the entire interface using that link will fail, affecting normal communication between accelerator cards and ultimately hindering the progress of the distributed training tasks of the entire cluster. A common fault-handling approach is to configure redundant interfaces, switching traffic to the backup interface in case of failure. However, this method wastes interface resources and fails to fully utilize network bandwidth. Given the current limitations of interconnect bandwidth to meet the needs of Artificial Intelligence (AI) accelerator cards, this further restricts the communication efficiency of the accelerator cards. Additionally, extra interfaces occupy chip space on the AI accelerator card, increasing device power consumption. To avoid resource waste, many AI accelerator cards are configured with only a single interface. Once that interface fails, the accelerator card becomes unusable, significantly reducing the availability and stability of the cluster.
[0020] Based on this, this application manages link resources in a refined manner by splitting the interconnection interfaces of network interface cards / accelerator cards and switching equipment. When a link within an interface (such as tx0) fails, the entire interface does not need to be interrupted. Traffic is redirected to other normal links (such as tx1-tx3) for forwarding (the interface can also be slowed down), ensuring uninterrupted communication between accelerator cards. This application maximizes the use of existing link resources without adding additional interface configurations, avoiding resource waste and increased power consumption caused by redundant interfaces. At the same time, it reduces the risk of accelerator card unavailability due to single interface failure, improves the availability and operational stability of the entire computing cluster, and meets the high-efficiency communication requirements of AI distributed training.
[0021] Figure 1 This is a first flowchart illustrating the communication method provided in this application embodiment. The following will be combined with... Figure 1 The steps shown are explained as follows: Figure 1 As shown, this method is applied to a network adapter, and the method includes the following steps 101 to 102. Step 101: Determine that the first sub-interface of the network adapter has failed, and set the status of the first sub-interface to the fault state.
[0022] In this embodiment, the network adapter includes, but is not limited to, network interface cards (NICs) and accelerator cards. The network adapter senses the physical link status of the first sub-interface in real time. Once it detects anomalies such as link disconnection or negotiation failure, it preliminarily determines that the sub-interface is faulty, thus achieving real-time monitoring at the link layer. In practical applications, the network adapter can use its built-in link detection module to sense the physical link status of the first sub-interface in real time based on Ethernet link up / down detection and Point-to-Point Protocol (PPP) Link Control Protocol (LCP) negotiation.
[0023] In some embodiments, the network adapter can also incorporate a service connectivity detection mechanism, such as periodically sending ping test packets to critical nodes on the other end and sending dedicated service test messages. If packet loss exceeds a threshold or there is no response after a timeout, it can assist in confirming sub-interface failures, avoiding misjudgments caused by brief fluctuations at the link layer, and achieving service layer-assisted verification. Therefore, the network adapter of this application can accurately identify faulty sub-interfaces through dual detection at the link layer and service layer, reducing the probability of misjudgment and ensuring the accuracy of fault determination.
[0024] In this embodiment, upon confirming a fault, the network adapter marks the first sub-interface as "faulty" in its local configuration database. Simultaneously, it updates the status information of the local routing module and the service application layer via management protocols such as Network Configuration Protocol (Netconf) and Simple Network Management Protocol (SNMP), notifying relevant modules to stop forwarding data through that sub-interface. The local routing and service modules promptly adjust their strategies, quickly switching service traffic that was originally passing through the faulty sub-interface to other normal sub-interfaces, maximizing service continuity.
[0025] Step 102: Send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the status of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0026] In this embodiment, the identification information (such as sub-interface number, Media Access Control Address (MAC) address, Internet Protocol Address (IP) address, Virtual LAN Identifier (VLAN ID), etc.) and fault information (such as fault type, occurrence time, fault cause description, etc.) of the first sub-interface are extracted from the local configuration database through the second sub-interface of the network adapter. The information is then encapsulated according to a preset fault notification message format. The message typically includes fields such as source device identifier, fault sub-interface identifier, fault type, trigger time, and checksum to ensure that it can be recognized and parsed by the network switching device.
[0027] In this embodiment, the fault notification message is sent to the network switching device through the normal link of the second sub-interface. The transmission follows normal IP forwarding rules and can use either User Datagram Protocol (UDP) or Transmission Control Protocol (TCP) (TCP is more reliable and suitable for scenarios with high reliability requirements). If UDP is used, a retransmission mechanism can be configured to avoid message loss. In this way, the fault notification message accurately carries critical fault information, ensuring that the network switching device can quickly locate the corresponding link and complete state synchronization. Sending the message through the normal sub-interface avoids the impact of the faulty link, ensuring reliable transmission of fault information and preventing state inconsistencies caused by notification failures.
[0028] The communication method provided in this application determines that the first sub-interface of the network adapter has failed and sets its state to a fault state. A fault notification message is then sent to the network switching device through the second sub-interface of the network adapter. This fault notification message includes the identification information and fault information of the first sub-interface of the faulty network adapter. The fault notification message also notifies the network switching device to set its third sub-interface to a fault state. The third sub-interface of the network switching device and the first sub-interface of the network adapter have a one-to-one mapping relationship. This application quickly switches service traffic that would otherwise pass through the faulty sub-interface to other normal sub-interfaces, maximizing service continuity. Furthermore, by sending the fault notification message through the second sub-interface (i.e., the normal sub-interface), it accurately carries critical fault information, ensuring that the network switching device can quickly locate the corresponding link and complete state synchronization.
[0029] In some embodiments, the above communication method further includes: Step 201: Based on the granularity of the serializer / deserializer, the physical interface of the network adapter is split into multiple sub-interfaces.
[0030] In this embodiment, the serializer / deserializer (SerDes) hardware capability built into the network adapter is utilized to split the link bandwidth of the physical interface according to a preset granularity (e.g., 10Gbps, 25Gbps), generating multiple independent sub-interfaces. For example, a 100Gbps physical interface can be split into four 25Gbps sub-interfaces or two 50Gbps sub-interfaces, with the splitting granularity matching the specifications supported by the device hardware. This achieves fine-grained allocation of physical interface bandwidth, allowing different bandwidth granularities to be allocated to different sub-interfaces according to service requirements, improving link resource utilization; it also provides the hardware foundation for parallel processing of multiple services, enabling different sub-interfaces to carry different service flows and avoiding mutual interference between services.
[0031] In this embodiment, the physical interface view is accessed through the device command-line interface, and a split command is executed to specify the number of sub-interfaces and granularity parameters after splitting. In a feasible scenario, the portswitch command is executed in the physical interface view to enable port switching, and then the interface range command is used to create sub-interfaces in batches, configuring the sub-interface numbering and granularity mapping relationship.
[0032] In some embodiments, the sub-interface partitioning granularity may also include a logical sub-interface dimension composed of multiple SerDes.
[0033] Step 202: Split the link resources of the physical interface of the network adapter, allocate independent transmission and reception channels to each of the multiple sub-interfaces, and configure the identification information of each sub-interface.
[0034] In this embodiment, the link resource management module of the network adapter splits the transmit and receive link resources of the physical interface as needed, allocating independent transmit and receive channels to each sub-interface. The bandwidth allocation and priority of each sub-interface can be specified via configuration commands; for example, a higher-priority transmit channel can be allocated to critical service sub-interfaces to ensure low latency in service forwarding. This achieves link resource isolation between sub-interfaces, preventing interference between the transmit and receive data of different service flows and ensuring the stability of service forwarding.
[0035] In this embodiment, each sub-interface is configured with unique identification information, including sub-interface number, MAC address, and VLAN ID. The sub-interface number is used for internal device identification, the MAC address must be associated with the physical interface MAC address (usually an extension of the physical interface MAC address), and the VLAN ID is used to distinguish the service VLANs carried by different sub-interfaces. In this way, accurate identification and management of sub-interfaces are achieved through unique identification information, facilitating subsequent fault detection, status synchronization, and other operations.
[0036] In some embodiments, after a sub-interface is created, the network adapter's sub-interface automatically sends a Fast Link Pulse (FLP) message to the connected network switch. This message contains parameters such as the supported rate and duplex mode of the sub-interface. The network switch parses the parameters in the FLP message and matches them with the parameters supported by its own interface, selecting the highest rate and optimal duplex mode (full-duplex preferred) supported by both. If the network switch does not support auto-negotiation, the network adapter's sub-interface can fall back to a half-duplex mode supported by both parties through a parallel detection mechanism. After negotiation, both parties stop sending FLP messages, the link enters a stable state, and the network adapter and network switch record the negotiated rate and duplex mode of the sub-interface for subsequent forwarding policy configuration. This ensures that the sub-interface communicates with the network switch interface at the optimal rate and duplex mode, maximizing link transmission performance. It avoids link failures and packet loss caused by rate and duplex mode mismatches, ensuring the reliability of service forwarding.
[0037] In some embodiments, the above communication method further includes: Step 301: Determine that the link signal of the first sub-interface has been restored and the message sending and receiving function has been restored. Send a fault recovery message to the network switching device through the first sub-interface. The fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier.
[0038] In this embodiment, the network adapter uses its built-in link detection module to monitor the physical link signal of the first sub-interface in real time. When the link signal strength recovers to the normal threshold, the link status switches from "Down" to "Up," and the link layer protocol (such as Ethernet Link Up or PPP protocol LCP negotiation succeeds) returns to normal, it is preliminarily determined that the link signal has recovered. This provides an accurate basis for subsequent fault recovery notifications and status switching, ensuring the reliability of network status synchronization.
[0039] In some embodiments, dual verification at the link layer and service layer, combined with a service connectivity detection mechanism, can be used to periodically send ping test packets to key nodes on the other end and send dedicated service test messages. If multiple consecutive tests show no packet loss or timeouts, and the rate and latency of sending and receiving messages return to normal, the message sending and receiving functions of the first sub-interface are confirmed to have fully recovered. Once the link signal and sending / receiving functions are confirmed to have recovered, the network adapter triggers a status marking process to prepare to switch the state of the first sub-interface to normal. Here, dual verification at the link layer and service layer accurately confirms that the link signal and sending / receiving functions of the sub-interface have fully recovered, avoiding misjudgments caused by brief link fluctuations.
[0040] In this embodiment, the network adapter extracts the identification information of the first sub-interface (such as sub-interface number, MAC address, IP address, VLAN ID, etc.), the current system time as the recovery time, and the network adapter's device identifier (such as device MAC address, SN number) from the local configuration database. This information is then encapsulated according to a preset fault recovery message format. The message typically includes fields such as source device identifier, fault recovery sub-interface identifier, recovery time, and checksum to ensure it can be recognized and parsed by the network switching device. The fault recovery message is sent to the network switching device through the normal link of the first sub-interface. During transmission, normal IP forwarding rules are followed. TCP protocol can be used to ensure reliable message transmission; if UDP protocol is used, a retransmission mechanism can be set to avoid message loss. The fault recovery message in this application accurately carries key recovery information, ensuring that the network switching device can quickly locate the corresponding link and complete state synchronization. Sending the message through the recovered sub-interface directly utilizes the recovered link to complete the notification, ensuring reliable transmission of recovery information and avoiding notification failures due to anomalies in other links.
[0041] Step 302: Receive an acknowledgment message sent by the network switching device, wherein the acknowledgment message is used to inform the network adapter that the recovery notification has been received.
[0042] In this embodiment, the network switching device receives the fault recovery message through its corresponding sub-interface, parses the sub-interface identification information in the message, and locates the corresponding sub-interface according to the one-to-one mapping table previously generated through static configuration or dynamic protocols (such as Link Layer Discovery Protocol (LLDP) and Cisco Discovery Protocol (CDP)).
[0043] The network switching device marks the status of the corresponding sub-interface as "normal" in its local configuration database, and synchronously updates its own routing table and MAC address table to restore the forwarding capability through the sub-interface. At the same time, it sends an acknowledgment message to the first sub-interface of the network adapter to confirm that the status synchronization has been completed.
[0044] Step 303: Set the state of the first sub-interface to normal state, where normal state indicates that the bidirectional communication capability is met.
[0045] In this embodiment, after receiving the confirmation message, the network adapter officially marks the status of the first sub-interface as "normal" and synchronously updates the status information of the local routing module and the service application layer, notifying the relevant modules to resume forwarding data through the sub-interface.
[0046] This application achieves complete synchronization of the link status between the network adapter and the network switching device, avoiding the "state inconsistency" problem where one end is marked as normal while the other end is still marked as faulty, thus ensuring the reliability of bidirectional communication. The network switching device and the network adapter simultaneously restore the forwarding capability of their sub-interfaces, switching traffic back to the restored link, fully utilizing link resources, and improving overall network performance.
[0047] In some embodiments, the above communication method further includes: Based on the number of remaining sub-interfaces of the network adapter and the bandwidth requirements, bandwidth reduction is applied to each of the remaining sub-interfaces. The remaining sub-interfaces include those in a normal state from among the network adapter's multiple sub-interfaces. A normal state indicates a state that satisfies bidirectional communication capabilities.
[0048] In this embodiment, the network adapter's resource management module collects real-time information on the bandwidth utilization, current service type (e.g., critical services, general services), and service priority of all normally functioning sub-interfaces. Simultaneously, it acquires global bandwidth demand data, such as bandwidth control commands issued by the service management system and peak prediction results from real-time traffic monitoring. The module matches the bandwidth utilization of the remaining sub-interfaces with the bandwidth demand, calculating the required speed reduction for each sub-interface. For example, if the current global bandwidth demand decreases and the average utilization of the remaining sub-interfaces is below 50%, each sub-interface can be reduced proportionally. If only some services require speed reduction, adjustments can be made individually for the sub-interfaces carrying those services. This allows for accurate understanding of the bandwidth utilization status and service attributes of the remaining sub-interfaces, providing a scientific basis for subsequent bandwidth reduction and avoiding blind adjustments that could impact services. It enables dynamic matching of bandwidth resources, allocating them rationally according to actual needs and improving bandwidth resource utilization efficiency.
[0049] In some embodiments, the bandwidth of each remaining sub-interface is finely configured through the hardware control module of the network adapter. Based on SerDes hardware capabilities, the bandwidth granularity of the sub-interface is adjusted; for example, reducing the speed of a sub-interface from 25Gbps to 10Gbps, or limiting the upper limit of the sub-interface's transmit and receive bandwidth by configuring traffic shaping parameters. The corresponding sub-interface view is accessed in the device command-line interface, and the bandwidth reduction command is executed. Taking Huawei devices as an example, the `bandwidth` command is executed in the sub-interface view to specify a new bandwidth value; the `traffic shaping` command is executed to configure traffic shaping rules and limit bandwidth usage. For multiple remaining sub-interfaces, a batch configuration method can be used, completing the bandwidth reduction settings at once through scripts or batch commands, improving configuration efficiency. This allows for rapid completion of the bandwidth reduction operation for remaining sub-interfaces, ensuring accurate execution of bandwidth adjustment commands and meeting bandwidth demand control requirements. Hardware-level bandwidth limiting ensures the stability of the reduction, avoiding bandwidth anomalies caused by software configuration fluctuations.
[0050] In some embodiments, after the speed reduction is completed, the bandwidth utilization, packet loss rate, latency, and other indicators of the remaining sub-interfaces, as well as the overall network bandwidth load, are continuously monitored. If packet loss or excessive latency is detected in a critical service carried by a sub-interface, a bandwidth recovery mechanism is automatically triggered to restore the bandwidth of that sub-interface to its original value or appropriately increase it. If bandwidth demand changes again, the bandwidth demand statistics and speed reduction configuration process is re-executed. In this way, potential service anomalies after speed reduction can be detected in a timely manner, ensuring the normal operation of critical services and avoiding the impact of bandwidth reduction on service experience. This enables dynamic and elastic adjustment of bandwidth resources, flexibly optimizing bandwidth allocation based on real-time network status and improving the network's adaptive capabilities.
[0051] Figure 2This is a second flowchart illustrating the communication method provided in the embodiments of this application. The following will be combined with... Figure 2 The steps shown are explained as follows: Figure 2 As shown, this method is applied to a network switching device, and the method includes steps 401 to 404. Step 401: Receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed.
[0052] In this embodiment, the network switching device continuously monitors the link through its normal sub-interfaces. When it receives a fault notification message from the network adapter, it identifies the message as a fault notification message based on information such as the protocol identifier and source device identifier in the message header. The message content is parsed to extract the identification information (such as sub-interface number, MAC address, VLAN ID, etc.) and fault information (such as fault type, fault occurrence time, etc.) of the first sub-interface of the faulty network adapter. The extracted identification information and fault information are verified to confirm the legality of the message source and the completeness and accuracy of the information, preventing erroneous operations due to message tampering or errors. In this way, fault notifications from the network adapter are accurately captured, and key fault-related information is quickly obtained, providing accurate basis for subsequent fault handling. Legality verification ensures the authenticity and reliability of the fault information, preventing misconfiguration of the switching device caused by erroneous information.
[0053] Step 402: Based on the identification information and fault information, set the status of the third sub-interface of the network switching device to a fault state, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0054] In this embodiment, the network switching device queries a pre-configured sub-interface mapping table or a sub-interface mapping table negotiated via dynamic protocols (such as LLDP or CDP). Based on the first sub-interface identifier extracted from the fault notification message, it accurately locates the corresponding third sub-interface. The device's configuration management module sets the matched third sub-interface to a fault state. The corresponding sub-interface view can be accessed via the command-line interface, and the shutdown command can be executed to close the interface; alternatively, the interface status can be modified directly in the graphical interface through the device's management platform. The routing table and MAC address table of the network switching device are updated synchronously, marking the routing entries and MAC address mappings related to the faulty sub-interface as invalid to prevent subsequent traffic from being forwarded to the faulty link. This quickly completes the status marking of the faulty sub-interface, synchronizing the link status between the network adapter and the network switching device, avoiding the "state inconsistency" problem where one end is marked as faulty while the other end continues to forward normally. By synchronously updating the routing and MAC tables, traffic forwarding on the faulty link is blocked, preventing packet loss, data anomalies, and other problems caused by the faulty link, thus ensuring the reliability of overall network communication.
[0055] In some embodiments, the above communication method further includes: The resource links of the physical interface of the network switching device are split into sub-interfaces corresponding to each sub-interface of the network adapter, and the identification information of each sub-interface is configured.
[0056] In this embodiment, the total bandwidth and port type (such as 10Gbps optical port and 25Gbps electrical port) of the physical interface of the network switching device are first evaluated. Then, the number of sub-interfaces of the network adapter and the service requirements of each sub-interface (such as the critical service sub-interfaces requiring higher bandwidth) are combined to plan the bandwidth allocation ratio and port type adaptation scheme for each split sub-interface.
[0057] In some embodiments, the physical interface is split through the hardware management module of the network switching device. Taking a network switching device that supports sub-interface splitting as an example, the physical interface view is accessed in the command-line interface, and a sub-interface creation command is executed (for example, executing `interface range GigabitEthernet 0 / 0 / 1.1 to GigabitEthernet 0 / 0 / 1.10` creates 10 sub-interfaces based on this physical interface); at the same time, the bandwidth parameters of each sub-interface are configured, such as `bandwidth1000` setting the sub-interface bandwidth to 1Mbps.
[0058] In some embodiments, if the physical interface hardware does not support direct splitting of sub-interfaces, virtualization technology can be used, such as through the Virtual Extensible LAN (VXLAN) protocol, to virtualize multiple logical sub-interfaces on the physical interface. Each logical sub-interface corresponds to an independent VLAN, thereby realizing the logical splitting of resource links.
[0059] This application breaks down the resources of a single physical interface into multiple independent sub-interface resources, meeting the interconnection requirements of multiple sub-interfaces in a network adapter and achieving fine-grained allocation of physical resources. Through reasonable planning of bandwidth and port types, it ensures the resource requirements of different service sub-interfaces and avoids performance degradation caused by resource contention.
[0060] In this embodiment, unique identification information is planned for each split sub-interface, including sub-interface number (e.g., GigabitEthernet 0 / 0 / 1.1), MAC address (which can be a combination of physical interface MAC address and sub-interface number, or a globally unique MAC address can be assigned separately), VLAN ID, and IP address (if Layer 3 communication is required). It is also ensured that the identification information corresponds one-to-one with the identification information of the corresponding sub-interface on the network adapter. The corresponding sub-interface view is accessed in the network switching device command line, and each identification information is configured sequentially.
[0061] In some embodiments, for multiple sub-interfaces, the identification information can be configured in batches by script. After the configuration is completed, a verification command is executed, such as display interface, to view the sub-interface configuration information, to ensure that all identification information is configured accurately and without duplication or conflict.
[0062] This application assigns a unique and accurate identifier to each sub-interface, enabling precise identification and management of the sub-interfaces and providing a foundation for subsequent mapping and connection with network adapter sub-interfaces. By ensuring a one-to-one correspondence between the identifier and the network adapter sub-interface identifier, accurate connection between the two ends of the link is guaranteed, avoiding communication failures caused by identifier mismatches.
[0063] In some embodiments, the above communication method further includes: Step 501: Receive a fault recovery message sent by the network adapter, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier.
[0064] In this embodiment, the corresponding sub-interface of the network switching device continuously monitors the link. When it receives a fault recovery message from the network adapter, it identifies the message as a fault recovery message based on information such as the protocol identifier and source device identifier in the message header. The message content is parsed to extract the identification information of the first sub-interface of the network adapter where the fault recovery occurred (such as sub-interface number, MAC address, VLAN ID, etc.), recovery time, and device identifier. The extracted identification information, recovery time, and device identifier are verified to confirm the legality of the message source and the completeness and accuracy of the information, preventing erroneous operations due to message tampering or errors.
[0065] This application accurately captures network adapter fault recovery notifications, quickly obtains key recovery-related information, and provides an accurate basis for subsequent recovery processing. Through legality verification, it ensures the authenticity and reliability of the recovery information, preventing misconfiguration of network switching equipment due to erroneous information.
[0066] Step 502: Send an acknowledgment message to the network adapter, wherein the acknowledgment message is used to indicate to the network adapter that the recovery notification has been received.
[0067] In this embodiment, the network switching device generates an acknowledgment message containing the corresponding first sub-interface identifier and device identifier based on the key information in the received fault recovery message. An acknowledgment timestamp can be added to the message for the network adapter to verify the timeliness of the acknowledgment. The acknowledgment message is sent to the network adapter through a third sub-interface mapped to the first sub-interface, ensuring accurate delivery. After sending, the device's link status monitoring module can confirm whether the acknowledgment message was successfully sent. If sending fails, a retransmission mechanism is triggered to ensure the network adapter receives the recovery acknowledgment feedback.
[0068] This application sends a notification to the network adapter that it has received the recovery notification, thus synchronizing the fault recovery status of both ends of the device and avoiding the "inconsistent status" problem where one end has recovered while the other end is still in the fault handling process. Through the exchange of confirmation messages, the fault recovery process is ensured to be closed-loop, allowing the network adapter to know that the network switching device has completed the status update and can resume normal service interaction.
[0069] Figure 3 This is a schematic diagram of the interconnection between the accelerator card and the switching device provided in the embodiments of this application, such as... Figure 3 As shown, by splitting the interfaces of the network interface card (NIC / accelerator card) and the switching equipment into smaller, more granular sub-interfaces (divided according to SerDes granularity), the physical interface of the NIC / accelerator card is divided into four sub-interfaces, each connected to the switch via fiber optic or copper cable. In high-speed network interfaces (such as 400GE), a single physical interface consists of multiple parallel SerDes channels, each containing a pair of differential signal lines used for transmission (tx) and reception (rx), respectively.
[0070] For example, the initial wire pair tx0 and rx0 can be set as sub-interface 0, tx1 and rx1 as sub-interface 1, tx2 and rx2 as sub-interface 2, and tx3 and rx3 as sub-interface 3. The number of sub-interfaces corresponds to the number of SerDes pairs on the physical interface. Thus, if the network card / accelerator card bandwidth is 400GE, then each sub-interface will have a bandwidth of 100GE, and the total bandwidth of the four sub-interfaces will be 4 × 100GE.
[0071] Of course, logical interfaces can also be configured, such as setting tx0 and tx1 as sub-interface 1, and tx2 and tx3 as sub-interface 2. In this way, if the network card / accelerator card bandwidth is 400GE, then each sub-interface will have a bandwidth of 200GE, and the total bandwidth of the two sub-interfaces will be 2×200GE.
[0072] In this embodiment, the fault handling mechanism is as follows: According to the above interconnection scheme, configure the interfaces of the network card / acceleration card and the switching equipment respectively, create the corresponding sub-interfaces, and negotiate the corresponding interface speeds; The network interface card / accelerator card and the switching equipment monitor the status of the TX and RX devices on the local interface respectively. When device and link abnormalities are detected, such as when the transmitted and received optical power is lower than a certain threshold, it is determined that the sub-interface has a link or device failure. When a sub-interface fault is detected, the sub-interface is set to a fault state. The faulty end sends a fault notification message to the peer device through other normal sub-interfaces, such as tx1. After receiving a fault message, the peer device sets the corresponding sub-interface to a fault state, and simultaneously sends and receives the message through other normal interfaces. The bandwidth of the entire physical interface is reduced to... Figure 3 For example, if a link fails, the speed drops to 3 / 4. Synchronously trigger log alarms and report sub-interface failures to the operation and maintenance management platform.
[0073] In this embodiment, the fault recovery mechanism is as follows: When the faulty end detects that the sub-interface has returned to normal, it sends a fault recovery message to the peer device through the sub-interface. After receiving the fault recovery message, the peer device replies with an acknowledgment message, sets the faulty sub-interface to the normal state, and forwards normal messages. After receiving the acknowledgment message, the fault recovery end sets the faulty sub-interface to the normal state and forwards normal messages.
[0074] Simultaneously report fault recovery logs to the operation and maintenance management platform.
[0075] This application provides a method for failover of communication interface links of an accelerator card, ensuring that when one link on the interface fails, the entire interface isolates the failed link while other links continue data forwarding normally, maintaining service continuity. By breaking down the physical interfaces of the network card / accelerator card and switching equipment into finer-grained sub-interfaces, a fine-grained fault isolation mechanism is implemented. This maximizes the utilization of interface bandwidth while minimizing the impact of sub-interface link failures on the entire interface, ensuring the stability of the entire computing cluster. After a sub-interface failure and recovery, a fault notification and recovery notification mechanism is implemented through other normal sub-interfaces, achieving a fine-grained fault isolation scheme for the physical interface.
[0076] This application also provides a communication device applied to a network adapter, such as... Figure 4 As shown, the communication device includes: The first processing unit 601 is used to determine that the first sub-interface of the network adapter has failed and set the state of the first sub-interface to a fault state. The first sending unit 602 is used to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0077] In some embodiments, the first processing unit 601 is configured to split the physical interface of the network adapter into multiple sub-interfaces based on the granularity of the serializer / deserializer; split the link resources of the physical interface of the network adapter; allocate independent transmission channels and reception channels to each of the multiple sub-interfaces; and configure the identification information of each sub-interface.
[0078] In some embodiments, the first processing unit 601 is used to determine that the link signal of the first sub-interface has been restored and the message sending and receiving function has been restored; The first sending unit 602 is used to send a fault recovery message to the network switching device through the first sub-interface, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier; Figure 4 The communication device also includes a first receiving unit 603, which is used to receive an acknowledgment message sent by the network switching device, wherein the acknowledgment message is used to send feedback to the network adapter that a recovery notification has been received; The first processing unit 601 is used to set the state of the first sub-interface to a normal state, wherein the normal state indicates a state that satisfies the bidirectional communication capability.
[0079] This application also provides a communication device applied to network switching equipment, such as... Figure 5 As shown, the communication device includes: The second receiving unit 701 is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed. The second processing unit 702 is used to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0080] In some embodiments, the second processing unit 702 is used to split the resource links of the physical interface of the network switching device into sub-interfaces corresponding to each sub-interface of the network adapter, and configure the identification information of each sub-interface.
[0081] The second receiving unit 701 is used to receive a fault recovery message sent by the network adapter, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier; Figure 5 The communication device also includes a second sending unit 703, which is used to send an acknowledgment message to the network adapter, wherein the acknowledgment message is used to inform the network adapter that a recovery notification has been received.
[0082] This application also provides a network adapter, such as... Figure 6 As shown, the network adapter 800 includes: a first communication interface 801 and a first processor 802; wherein, The first communication interface 801 is capable of exchanging information with network switching equipment; The first processor 802 is connected to the first communication interface 801 to enable information interaction with the network switching device and to execute the methods provided by one or more technical solutions on the network adapter side when running a computer program. The first memory 803 stores computer programs that can run on the first processor 802.
[0083] The first processor 802 is configured to determine that the first sub-interface of the network adapter has failed and set the state of the first sub-interface to a fault state. The first communication interface 801 is used to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
[0084] In some embodiments, the first processor 802 is configured to split the physical interface of the network adapter into multiple sub-interfaces based on the granularity of the serializer / deserializer; split the link resources of the physical interface of the network adapter; allocate independent transmission and reception channels to each of the multiple sub-interfaces; and configure identification information for each sub-interface.
[0085] In some embodiments, the first processor 802 is configured to determine that the link signal of the first sub-interface has been restored and the message sending and receiving function has been restored; The first communication interface 801 is used to send a fault recovery message to the network switching device through the first sub-interface, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier; and to receive an acknowledgment message sent by the network switching device, wherein the acknowledgment message is used to feedback to the network adapter that a recovery notification has been received. The first processor 802 is configured to set the state of the first sub-interface to a normal state, wherein the normal state indicates a state that satisfies bidirectional communication capability.
[0086] This application also provides a network switching device, such as... Figure 7 As shown, the network switching device 900 includes: a second communication interface 901 and a second processor 902; wherein, The second communication interface 901 is capable of exchanging information with the network adapter; The second processor 902 is connected to the second communication interface 901 to enable information interaction with the network adapter and to execute the methods provided by one or more technical solutions on the network switching device side when running computer programs. The second memory 903 stores computer programs that can run on the second processor 902.
[0087] The second communication interface 901 is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed. The second processor 902 is configured to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and the fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface. In some embodiments, the second processor 902 is configured to split the resource links of the physical interface of the network switching device into sub-interfaces corresponding to each sub-interface of the network adapter, and configure identification information for each sub-interface.
[0088] In some embodiments, the second communication interface 901 is used to receive a fault recovery message sent by the network adapter, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier; and to send an acknowledgment message to the network adapter, wherein the acknowledgment message is used to indicate to the network adapter that the recovery notification has been received.
[0089] It should be noted that the explanations of the steps in this embodiment that are the same as those in the above embodiments can be found in the descriptions in the above embodiments, and will not be repeated here.
[0090] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the communication method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0091] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of a network adapter / network switch reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the network adapter / network switch to perform the communication method described in this application.
[0092] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the communication method provided in this application, for example... Figure 1 or Figure 2 The communication method is shown.
[0093] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0094] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0095] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0096] As an example, computer-executable instructions can be deployed to execute on a single device, or on multiple devices located in one location, or on multiple devices distributed across multiple locations and interconnected via a communication network.
[0097] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A communication method, characterized in that, Applied to a network adapter, the method includes: If it is determined that the first sub-interface of the network adapter has failed, the state of the first sub-interface is set to a fault state. A fault notification message is sent to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
2. The method according to claim 1, characterized in that, The method further includes: Based on the granularity of the serializer / deserializer, the physical interface of the network adapter is split into multiple sub-interfaces; The link resources of the physical interface of the network adapter are split, and an independent transmit channel and receive channel are allocated to each of the multiple sub-interfaces, and the identification information of each sub-interface is configured.
3. The method according to claim 1, characterized in that, The method further includes: Once it is determined that the link signal of the first sub-interface has been restored and the message sending and receiving function has been restored, a fault recovery message is sent to the network switching device through the first sub-interface. The fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier. Receive an acknowledgment message sent by the network switching device, wherein the acknowledgment message is used to inform the network adapter that a recovery notification has been received; The state of the first sub-interface is set to normal, wherein the normal state indicates a state that satisfies bidirectional communication capability.
4. A communication method, characterized in that, Applied to network switching equipment, the method includes: Receive a fault notification message sent by a network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed; Based on the identification information and the fault information, the state of the third sub-interface of the network switching device is set to a fault state, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
5. The method according to claim 4, characterized in that, The method further includes: The resource links of the physical interface of the network switching device are split into sub-interfaces corresponding to each sub-interface of the network adapter, and the identification information of each sub-interface is configured.
6. The method according to claim 4, characterized in that, The method further includes: Receive a fault recovery message sent by the network adapter, wherein the fault recovery message includes the identification information of the first sub-interface, the recovery time, and the device identifier; Send an acknowledgment message to the network adapter, wherein the acknowledgment message is used to indicate to the network adapter that the recovery notification has been received.
7. A communication device, characterized in that, Applied to a network adapter, the device includes: The first processing unit is configured to determine that the first sub-interface of the network adapter has failed, and set the state of the first sub-interface to a fault state. The first sending unit is configured to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
8. A communication device, characterized in that, Applied to network switching equipment, the device includes: The second receiving unit is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed. The second processing unit is configured to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and the fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
9. A network adapter, characterized in that, Includes a first communication interface and a first processor; wherein, The first processor is configured to determine that the first sub-interface of the network adapter has failed, and set the state of the first sub-interface to a fault state; The first communication interface is used to send a fault notification message to the network switching device through the second sub-interface of the network adapter. The fault notification message includes the identification information and fault information of the first sub-interface. The fault notification message is used to notify the network switching device to set the state of the third sub-interface of the network switching device to a fault state. The third sub-interface has a one-to-one mapping relationship with the first sub-interface.
10. A network switching device, characterized in that, Includes a second communication interface and a second processor; wherein, The second communication interface is used to receive a fault notification message sent by the network adapter, wherein the fault notification message includes the identification information and fault information of the first sub-interface of the network adapter that has failed; The second processor is configured to set the state of the third sub-interface of the network switching device to a fault state based on the identification information and the fault information, wherein the third sub-interface has a one-to-one mapping relationship with the first sub-interface.
Citation Information
Patent Citations
Method, device and system for processing PCIe link failure
CN104170322A
Link switching method, device, network equipment and network system
CN109218107A
Method and device for keeping butt joint of optical module channels
CN120750425A
Smart serdes
US20230336198A1