Method and apparatus for handling failure of a link, computing device, and storage medium

CN122845495APending Publication Date: 2026-09-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365064.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]在上述过程中,上述两个设备之间以链路为粒度来感知故障,在确定链路发生故障之后,停止使用该链路进行数据传输,并在通过链路宽度协商来停止使用链路上发生故障的通道之后,再进行数据传输,故障处理效率较低,导致业务停顿时间较长,数据传输可靠性较低

Benefits of technology

[0017]第二方面,提供了一种链路的故障处理方法,该方法中发送端周期性通过链路上的通道,向接收端发送第一码流,接收端在根据第一码流的接收状况确定出链路上发生故障的通道的情况下,向发送端发送第一控制信息,发送端接收第一控制信息,基于该第一控制信息,对链路上处于使用状态的通道的数量进行调整,以停止使用链路上发生故障的通道,从而实现以通道为粒度来感知和处理链路上的故障,与在以链路为粒度来感知和处理故障的情况下,通过暂时停止使用链路来进行故障处理相比,能够在故障处理过程中保持数据的传输,从而避免业务停顿,提高数据传输的可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845495A_ABST
    Figure CN122845495A_ABST
Patent Text Reader

Abstract

The application provides a link fault processing method and device, a computing device and a storage medium, and belongs to the technical field of communication. In the link fault processing method provided by the application, a sending end transmits a first code stream to a receiving end through a channel on a link, the receiving end determines whether the channel has a fault according to a state of receiving the first code stream through the channel, and in the case that the channel has a fault, the receiving end sends first control information to the sending end, so that the sending end can adjust the number of channels in a use state on the link according to the first control information, to stop using the channel on the link that has a fault, thereby realizing sensing and processing of the fault on the link in the granularity of the channel. Compared with the case of sensing and processing the fault in the granularity of the link, the transmission of data can be maintained during the fault processing, thereby avoiding service interruption and improving the reliability of data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, computing device, and storage medium for handling link faults. Background Technology

[0002] In the field of communication technology, any two devices within a data center are connected via a link. Each link includes at least one lane, which is used for data transmission between the two ports. At the physical level, these links are implemented as cables or optical fibers. Failure of the cable or fiber will cause the link or the lane on the link to malfunction, leading to abnormal data transmission.

[0003] In the fault handling method for related links, the sending end inside the data center sends data and corresponding checksums to the receiving end through multiple channels on the link. The receiving end generates a new checksum based on the data received from the multiple channels. If the new checksum differs from the checksum received by the receiving end, the receiving end determines that the link has failed, stops data transmission with the sending end, and instead renegotiates the link width with the sending end through multiple interactions, thereby ceasing the use of the failed channel on the link.

[0004] In the above process, the two devices detect faults at the link level. After determining that a link has failed, they stop using the link for data transmission. Data transmission is then resumed only after the faulty channel on the link is stopped through link width negotiation. This results in low fault handling efficiency, long service interruption time, and low data transmission reliability. Summary of the Invention

[0005] This application provides a method, apparatus, computing device, and storage medium for handling link failures, which can improve the reliability of data transmission. The technical solution is as follows.

[0006] Firstly, a link fault handling method is provided. In this method, the receiving end receives a first code stream transmitted through a channel on the link. Based on the status of receiving the first code stream through the channel on the link, it determines whether a channel has failed, thus realizing fault detection at the channel level. If it is determined that at least one channel on the link has failed, the receiving end sends first control information to the sending end. This first control information instructs the number of channels in use on the link to be adjusted, so that the sending end can adjust the number of channels in use on the link according to the first control information to stop using the failed channel. This achieves link fault handling at the channel level. Compared with fault handling at the link level, which involves temporarily stopping the use of the link, this method can maintain data transmission during fault handling, thereby avoiding service interruption and improving the reliability of data transmission.

[0007] Optionally, during the process of the receiving end sending the first control information if it is determined that at least one channel on the link has failed, if it is determined that some channels on the link have failed, the receiving end sends the first control information, which carries a channel reduction request. Since the channel reduction request indicates a reduction in the number of channels in use on the link, after receiving the channel reduction request, the sending end can stop using the channels that have failed on the link by reducing the number of channels in use on the link, thereby improving the reliability of data transmission.

[0008] Optionally, the aforementioned channel reduction request includes channel information and a quantity. The channel information indicates the first channel that failed in the sequential arrangement of channels on the link, and the quantity refers to the number of channels that continue to be used on the link. After receiving the channel reduction request, the sending end can deduce the number of channels to be reduced as indicated by the channel reduction request based on the channel information and quantity carried in the request.

[0009] Optionally, the aforementioned channel downgrade request carries the identifier of the channel that has failed on the link, so that the sending end can accurately locate the failed channel based on the channel downgrade request, and then process the failed channel to improve the accuracy of fault handling.

[0010] Optionally, the above method further includes: the receiving end, in response to receiving second control information from the sending end, stopping the use of the channel indicated by the second control information. The second control information is sent by the sending end to the receiving end after the sending end has downgraded the channel according to the first control information. This allows the receiving end to downgrade the channel only after the sending end has downgraded it, avoiding data reception errors caused by premature downgrading and improving the reliability of data transmission.

[0011] Optionally, the above method further includes: the receiving end responding to a decrease in the number of channels in use on the link, starting a timer; if the timer reaches a first preset duration, sending a first channel upgrade request to the sending end; the first channel upgrade request instructs to increase the number of channels in use on the link, so as to increase the data transmission bandwidth and thus increase the data transmission speed after a period of channel downgrading by increasing the number of channels in use on the link.

[0012] Optionally, during the process of sending the first control information if at least one channel on the link is determined to be faulty, if all channels on the link are determined to be faulty, the receiving end sends the first control information to the sending end. The first control information carries a link establishment request, which instructs the physical layer to re-establish the link, so that the link between the sending end and the receiving end can be restored automatically after the fault occurs. Furthermore, during the restoration process, the link training state machine on the physical layer is kept in a connected state to prevent the application layer and other upper layers from perceiving the fault that occurred on the link.

[0013] Optionally, the above method also provides various conditions for stopping the use of the link in case of link failure. For example, if it is determined that all channels on the link have failed, the receiving end stops using the link; or, if it is determined that all channels on the link have failed, the receiving end sends a link disconnection request to stop using the link. The link disconnection request indicates that the link should be stopped, which improves the applicability of this method.

[0014] Optionally, the above method further includes: the receiving end responding to the successful re-establishment of the link starts timing; if the timing reaches a second preset duration, the receiving end sends a second channel upgrade request to the sending end, the second channel upgrade request indicating to increase the number of channels in use on the re-established link in order to increase bandwidth and improve data transmission speed.

[0015] Optionally, determining that at least one channel on the link has failed includes: for each channel on the link, whenever a first bitstream is received from that channel, a timer is started; if the timer exceeds a preset detection duration and no next first bitstream is received from that channel, the channel is determined to have failed. This enables fault detection at the channel level based on the reception status of the first bitstream, refining the granularity of fault detection and improving the sensitivity of fault detection.

[0016] Optionally, the above method is applied to a computing device, wherein the link training state machine on the physical layer of the computing device is always in a connected state. When the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link, thereby preventing faults on the link from affecting upper-layer services.

[0017] Secondly, a link fault handling method is provided. In this method, the transmitting end periodically sends a first code stream to the receiving end through a channel on the link. When the receiving end determines the channel on the link that has failed based on the reception status of the first code stream, it sends first control information to the transmitting end. The transmitting end receives the first control information and, based on the first control information, adjusts the number of channels in use on the link to stop using the channel that has failed. This allows for the perception and handling of link faults at the channel level. Compared with the method of temporarily stopping the use of the link to handle faults at the link level, this method can maintain data transmission during the fault handling process, thereby avoiding service interruption and improving the reliability of data transmission.

[0018] Optionally, the first control information carries a channel reduction request, which indicates a reduction in the number of channels in use on the link. During the process of adjusting the number of channels in use on the link based on the first control information, the sending end reduces the number of channels in use on the link based on the channel reduction request carried by the first control information, so as to stop using the channels that have failed on the link and improve the reliability of data transmission.

[0019] Optionally, the aforementioned channel reduction request includes channel information and a quantity. The channel information indicates the first channel that failed in the sequential arrangement of channels on the link, and the quantity refers to the number of channels that continue to be used on the link. After receiving the channel reduction request, the sending end can deduce the number of channels to be reduced as indicated by the channel reduction request based on the channel information and quantity carried in the request.

[0020] Optionally, the aforementioned channel downgrade request carries the identifier of the channel that has failed on the link, so that the sending end can accurately locate the failed channel based on the channel downgrade request, and then process the failed channel to improve the accuracy of fault handling.

[0021] Optionally, the above method further includes: the transmitting end sending second control information to the receiving end, the second control information indicating to stop using the reduced channels on the link. By sending the second control information to the receiving end after channel reduction, the transmitting end controls the receiving end to perform channel reduction, which can avoid data reception errors caused by the receiving end prematurely reducing channels, thereby improving the reliability of data transmission.

[0022] Optionally, during the process of reducing the number of channels in use on the link based on the channel reduction request carried by the first control information, the sending end reduces the number of channels in use on the link based on the channel reduction request carried by the first control information, and restores the previously reduced channels on the link to increase the number of channels in use on the link, increase the bandwidth of data transmission, and thus increase the data transmission speed.

[0023] Optionally, the above method further includes: the sending end responding to a decrease in the number of channels in use on the link starts timing; if the timing reaches a first preset duration, the sending end sends a first channel upgrade request to the receiving end; the first channel upgrade request indicates an increase in the number of channels in use on the link, so as to increase the data transmission bandwidth and thus increase the data transmission speed after a period of channel downgrading by increasing the number of channels in use on the link.

[0024] Optionally, the first control information carries a link establishment request, which instructs the physical layer to re-establish the link. During the process of adjusting the number of channels in use on the link based on the first control information, the transmitting end re-establishes the link with the receiving end based on the link establishment request carried by the first control information, so that the link between the transmitting end and the receiving end can recover automatically after a failure. Furthermore, during the recovery process, the link training state machine on the physical layer is always in a connected state to prevent the application layer and other upper layers from perceiving the failure that occurred on the link.

[0025] Optionally, the above method also provides various conditions for stopping the use of the link in case of link failure. For example, the sending end stops using the link in response to the receiving end stopping the use of the link; or, the sending end stops using the link in response to receiving a link disconnection request, wherein the link disconnection request indicates that the link should be stopped, thereby improving the applicability of the method.

[0026] Optionally, the above method further includes: the sending end initiating the process of reconstructing the link, that is, when the link is no longer in use, the sending end sends the first control information to the receiving end.

[0027] Optionally, the above method further includes: the sending end responding to the successful re-establishment of the link starts timing; if the timing reaches a second preset duration, the sending end sends a second channel upgrade request to the receiving end, the second channel upgrade request indicating to increase the number of channels in use on the re-established link in order to increase bandwidth and speed up data transmission.

[0028] Optionally, the above method is applied to a computing device, wherein the link training state machine on the physical layer of the computing device is always in a connected state. When the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link, thereby preventing faults on the link from affecting upper-layer services.

[0029] Thirdly, a link fault handling apparatus is provided for executing the aforementioned link fault handling method. Specifically, the link fault handling apparatus includes a functional module for executing the link fault handling method provided in the first aspect or any optional embodiment of the first aspect.

[0030] Fourthly, a link fault handling apparatus is provided for executing the aforementioned link fault handling method. Specifically, the link fault handling apparatus includes a functional module for executing the link fault handling method provided in the second aspect or any alternative method of the second aspect.

[0031] Fifthly, a computing device or cluster of computing devices is provided, the computing device including a processor for executing program code, causing the computing device or cluster of computing devices to perform methods as provided in any of the first and second aspects above or in various alternative implementations of such aspects.

[0032] A sixth aspect provides a computer-readable storage medium storing at least one piece of program code that is read by a processor to cause a computing device to perform a method as provided in any of the first and second aspects above or in various alternative implementations of such aspects.

[0033] In a seventh aspect, a computer program product or computer program is provided, the computer program product or computer program including program code stored in a computer-readable storage medium, a processor of a computing device reading the program code from the computer-readable storage medium, the processor executing the program code, causing the computing device to perform the method provided in any of the first and second aspects or various alternative implementations of any of the aspects.

[0034] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the implementation environment of a link fault handling method provided in an embodiment of this application;

[0036] Figure 2 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0037] Figure 3 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0038] Figure 4 This is a schematic diagram of a layered structure provided in an embodiment of this application;

[0039] Figure 5 This is a schematic diagram of a code stream for transmitting first control information provided in an embodiment of this application;

[0040] Figure 6 This is a data interaction diagram of a link fault handling method provided in an embodiment of this application;

[0041] Figure 7 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application;

[0042] Figure 8 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application;

[0043] Figure 9 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application;

[0044] Figure 10 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application;

[0045] Figure 11 This is a schematic diagram of a physical layer structure provided in an embodiment of this application;

[0046] Figure 12 This is a schematic diagram of the structure of a link fault handling device provided in an embodiment of this application;

[0047] Figure 13 This is a schematic diagram of another link fault handling device provided in an embodiment of this application. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0049] The implementation environment of the embodiments of this application is described below.

[0050] The technical solution provided in this application can be applied to any two devices that transmit data via cables or optical fibers to detect whether the link used for data transmission between the two devices has failed, and to maintain data transmission during the fault handling process, thereby avoiding service interruption and improving the reliability of data transmission.

[0051] refer to Figure 1 , Figure 1 This is a schematic diagram illustrating the implementation environment of a link fault handling method provided in an embodiment of this application, such as... Figure 1 As shown, the implementation environment includes a first device 101 and a second device 102, which are connected by a cable.

[0052] The first device 101 can be configured as a terminal or server, or as a network device such as a switch or router; this embodiment does not limit its configuration. When the first device 101 is configured as a terminal, it can be a laptop, desktop computer, or smartphone, etc., while the second device 102 can be configured as another terminal, or as a server or network device; this embodiment does not limit its configuration. Data transmission occurs between the first device 101 and the second device 102 via a cable. For example, the first device 101 is configured as a terminal, and the second device 102 is configured as a server connected to the terminal via a cable. An application runs on the first device 101, and the second device 102 provides background services for the application. Accordingly, the first device 101 transmits data to the second device 102 through the application, or the first device 101 receives data from the second device 102 through the application.

[0053] When the first device 101 is configured as a server, it can be an independent physical server, virtual machine, or node within a data center, while the second device 102 can be configured as a terminal, server, or network device, etc. This application embodiment does not limit this. For example, the first device 101 may be configured as a node within a data center, and the second device 102 may be configured as a switch or router connected to the first device 101. The first device 101 sends data to the second device 102 via a cable, or the first device 101 receives data from the second device 102 via a cable. As another example, both the first device 101 and the second device 102 may be configured as nodes within a data center. The data center includes multiple computing nodes and multiple storage nodes. The computing nodes perform calculations based on data in the storage nodes and store the results in the storage nodes. The storage nodes store the data. Taking a scenario where the first device 101 is configured as any computing node in a data center, and the second device 102 is configured as a storage node connected to the first device 101 within the data center, during data transmission, the first device 101 sends a data processing request to the second device 102. The second device 102 then sends the data indicated by the data processing request back to the first device 101. The first device 101 performs calculations based on the data sent by the second device 102 and sends the calculation results back to the second device 102 to store the results. It should be noted that when both the first device 101 and the second device 102 are configured as nodes within a data center, they may be nodes within the same data center, or they may be nodes within different data centers. That is, the first device 101 and the second device 102 can perform data transmission across data centers. This embodiment does not limit this.

[0054] When the first device 101 is configured as a network device, the first device may be a switch or router in a data center, while the second device 102 can be configured as a terminal, server, or network device, etc., and this application embodiment does not limit this. For example, the first device 101 and the second device 102 may be configured as two interconnected switches in a data center, with the first device 101 sending data from other devices to the second device 102, or the first device 101 receiving data from the second device 102, and this application embodiment does not limit this.

[0055] In some embodiments, when the first device 101 and the second device 102 are connected by an optical fiber, an optical module is installed on the first device 101 and the second device 102. The optical module is used to convert the data to be transmitted into an optical signal and transmit the optical signal through the optical fiber.

[0056] The link used for data transmission between the first device 101 and the second device 102 described below is introduced below.

[0057] As explained above, there are two service directions between the first device 101 and the second device 102: a first service direction and a second service direction. The first service direction is the transmission of data from the first device 101 to the second device 102, and the second service direction is the transmission of data from the second device 102 to the first device 101. For the first device 101, the first service direction is the transmitting direction, and the second service direction is the receiving direction. For the second device 102, the second service direction is the transmitting direction, and the first service direction is the receiving direction.

[0058] The first device 101 and the second device 102 transmit data via a link, which includes at least one channel for data transmission. In some embodiments, the link is a bidirectional link, meaning it includes at least one channel in a first service direction and at least one channel in a second service direction. Specifically, for both the first device 101 and the second device 102, the link includes at least one channel in a transmitting direction and at least one channel in a receiving direction, enabling the link to transmit data from the first device 101 to the second device 102 and vice versa. In some embodiments, the link is a unidirectional link, meaning it includes at least one channel in the first service direction, allowing the link to transmit data only from the first device 101 to the second device 102. Alternatively, the link includes at least one channel in the second service direction, allowing the link to transmit data only from the second device 102 to the first device 101. When the link is unidirectional, the first device 101 and the second device 102 can use multiple links to transmit data. These multiple links include at least one link in the first service direction and at least one link in the second service direction, enabling the multiple links to realize data transmission in different service directions.

[0059] During the process of the first device 101 sending data to the second device 102, if the link includes multiple channels in the first service direction, the first device 101 splits the data into multiple sub-data segments and sends these sub-data segments to the second device 102 through the aforementioned multiple channels in the first service direction. The second device 102 receives these sub-data segments through the aforementioned multiple channels in the first service direction and reconstructs the original data based on these sub-data segments. For example, if the link includes four channels in the first service direction, during data transmission, the first device 101 splits the data into four sub-data segments and sends these four sub-data segments to the second device 102 through the aforementioned four channels in the first service direction. After receiving these four sub-data segments, the second device 102 reconstructs the original data based on these four sub-data segments. The process of the second device 102 sending data to the first device 101 is the same as the above process, and will not be described again in this embodiment.

[0060] In some embodiments, the link width indicates the number of channels in use (open channels) on the link. This width can be represented by XN, where N is the number of channels in use on the link, and N is an integer greater than 0. For example, X16, X8, X4, and X2 represent the number of channels in use on the link as 16, 8, 4, and 2, respectively. TXN (transmitter XN) represents the link width in the transmission direction, i.e., the number of transmission-direction channels in use on the link. RXN (receiver XN) represents the link width in the reception direction, i.e., the number of reception-direction channels in use on the link. For example, for the first device 101, a link width of TX4RX2 indicates that for the first device 101, the number of transmission-direction channels in use on the link is 4, and the number of reception-direction channels in use on the link is 2.

[0061] Both the first device 101 and the second device 102 described above can be implemented as follows: Figure 2 The computing device shown below, in conjunction with Figure 2 The hardware structure of the computing device will be described.

[0062] Figure 2 This is a schematic diagram of a computing device provided in an embodiment of this application. It should be understood that the computing device described below can implement any function of any of the methods described below. Typically, the computing device 200 includes a processor 201 and a memory 202.

[0063] Processor 201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 201 may be implemented using at least one hardware form selected from digital signal processing (DSP), field-programmable gate array (FPG A), and programmable logic array (PLA). Processor 201 may also include a main processor and a coprocessor. The main processor, also known as the central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 201 may integrate a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 201 may also include an artificial intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0064] The memory 202 may include one or more computer-readable storage media, which may be non-transitory. The memory 202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 202 are used to store at least one piece of program code, which is executed by the processor 201 to implement the link fault handling method provided in the method embodiments of this application.

[0065] In some embodiments, the computing device 200 may also optionally include a peripheral device interface 203 and at least one peripheral device. The processor 201, memory 202, and peripheral device interface 203 can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 203 via a bus, signal lines, or a circuit board.

[0066] In some embodiments, the computing device 200 may be a portable mobile terminal, such as a smartphone, tablet computer, Moving Picture Experts Group Audio Layer III (MP3) player, Moving Picture Experts Group Audio Layer IV (MP4) player, laptop computer, or desktop computer. The computing device 200 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0067] In some embodiments, the computing device 200 may be a stand-alone physical server, or, implemented as such Figure 3 The computing device cluster shown refers to a server cluster composed of multiple physical servers or a distributed file system, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Taking computing devices as cloud servers as an example, computing devices can also be called a cloud platform (short for cloud computing platform), which refers to services based on hardware and software resources that provide computing, network, and storage capabilities. Through the network "cloud," massive amounts of data are processed and analyzed remotely before being returned to the user, featuring large scale, distributed nature, virtualization, high availability, scalability, on-demand service, and security. Cloud platforms can achieve rapid deployment and release of configurable computing resources with relatively low management costs or low interaction complexity between users and service providers.

[0068] In some embodiments, the computing device 200 can be implemented as a supernode. A supernode refers to a high-performance cluster formed by interconnecting multiple nodes through a high-bandwidth, low-latency inter-chip interconnect bus and switches. It can act as a node with high computing and storage capabilities in a network, processing large-scale data and complex tasks. A supernode includes multiple nodes. These multiple nodes can be processors, servers, desktop computers, controllers of storage arrays, or memory, etc. The processor can be a central processing unit (CPU), graphics processing unit (GPU), data processing unit (DPU), neural processing unit (NPU), or an embedded neural network processing unit (NPU) for data processing, such as an XPU.

[0069] The above implementation environment is illustrated with the example that the first device 101 and the second device 102 are both computing devices. In some embodiments, the first device 101 and the second device 102 are respectively modules, chips, chipsets, circuit boards or components equipped with chips or chipsets disposed in computing devices. This application embodiment does not limit this.

[0070] The first device 101 and the second device 102 described above employ the following methods during data transmission: Figure 4 The layered structure shown is used to process data, and each layer in this structure can be implemented in hardware. Figure 4 This is a schematic diagram of a layered structure provided in an embodiment of this application. See also... Figure 4 This layered architecture, from bottom to top, includes a serializer / deserializer (SerDes), physical layer (PHY), data link layer (DLL), network layer (NL), transport layer (TL), session layer (SL), presentation layer (PL), and application layer (AL). Figure 4 Only SerDes, PHY, and DLL are shown in the image.

[0071] During the process of the first device 101 sending data to the second device 102, the data is generated in the AL of the first device 101, and then passes through the PL, SL, TL, NL, DLL, PHY, and SerDes of the first device 101 before finally being sent out through the link. The data passes through the SerDes, PHY, DLL, NL, TL, SL, and PL of the second device 102 before finally arriving at the AL of the second device 102, thus realizing the data transmission from the first device 101 to the second device 102. The process of the second device 102 sending data to the first device 101 is similar to the above process, and will not be described again in this embodiment.

[0072] The SerDes of the first device 101 includes a transmit interface (TX) and a receive interface (RX). The SerDes of the second device 102 also includes RX and TX. The TX of the first device 101 and the RX of the second device 102 are connected via a channel in a first service direction, and the RX of the first device 101 and the TX of the second device 102 are connected via a channel in a second service direction. For example, as... Figure 4 As shown, the SerDes of the first device 101 includes 4 TXs (TX0 to TX3) and 4 RXs (RX0 to RX3), and the SerDes of the second device 102 includes 4 TXs (TX4 to TX7) and 4 RXs (RX4 to RX7). The 4 TXs of the first device 101 and the 4 RXs of the second device 102 are connected through 4 channels in the first service direction (lane0 to lane3), and the 4 TXs of the second device 102 and the 4 RXs of the first device 101 are connected through 4 channels in the second service direction (lane4 to lane7). The above-mentioned 4 channels in the first service direction and 4 channels in the second service direction can be implemented as a single link or as multiple links, for example, as... Figure 4 As shown, lane0 to lane3 and lane4 to lane5 are implemented as one link, which is link 1, and lane6 to lane7 are implemented as another link, which is link 2.

[0073] In some embodiments, the ports of the first device 101 and the second device 102 are high-speed serial computer expansion bus standard (PCIe) ports, or the ports of the first device 101 and the second device 102 are other serial ports, such as industry standard architecture (ISA) ports, etc., and this application embodiment does not limit them.

[0074] In related technologies, the sending end within the data center transmits data and corresponding checksums to the receiving end through multiple channels on the link. The receiving end generates a new checksum based on the data received from these channels. The receiving end detects faults at the link level by comparing the received checksum with the generated new checksum. However, this method has a coarse-grained fault detection granularity and low sensitivity. Furthermore, after determining that a link failure has occurred, the receiving end stops data transmission with the sending end and instead renegotiates the link width through multiple interactions. Data transmission resumes only after the failed channel on the link is successfully terminated through negotiation. This results in low fault handling efficiency, prolonged service interruptions, and low data transmission reliability.

[0075] To address the aforementioned issue of low data transmission reliability, the link fault handling method provided in this application involves the sending end periodically sending a first code stream to the receiving end through a channel on the link. The receiving end determines whether a channel on the link has failed based on the status of receiving the first code stream through the channel. This allows for fault detection at the channel level, refining the granularity of fault perception and improving its sensitivity. Furthermore, when a channel failure is determined, the receiving end sends first control information to the sending end, enabling the sending end to adjust the number of channels in use on the link according to this first control information, thereby stopping the use of the failed channel. During this process, the sending end and receiving end can continue data transmission through channels on the link that are not experiencing failures, based on the first control information, avoiding service interruptions and ensuring data transmission reliability. The sending end and receiving end can be implemented as described above. Figure 1 The first device 101 and the second device 102 in the implementation environment shown are not limited in this application embodiment.

[0076] The above process involves two types of fault scenarios: one is that some channels on the link fail while others can still transmit data normally; the other is that all channels on the link fail. For these two different fault scenarios, the first control information sent by the receiving end to the sending end carries different requests. In the case of partial channel failure, the first control information carries a channel reduction request, which instructs to reduce the number of channels in use on the link, including the channels that have failed. In the case of all channels failing, the first control information carries a link establishment request, which instructs the physical layer to re-establish the link between the sending and receiving ends.

[0077] The above-mentioned downgrade request has multiple implementation forms, which are described below.

[0078] In one implementation, a channel reduction request indicates a reduction in the number of channels in use on the link. The channels to be reduced include channels that have failed on the link. The channel reduction request can indicate the channels to be reduced in various ways. One way is that the channel reduction request carries the identifier of the failed channel on the link. The device receiving the channel reduction request can stop using the failed channel on the link according to the request. The channel reduction request can carry the identifiers of all failed channels on the link, or it can carry the identifiers of some failed channels on the link; this embodiment does not limit this. When the channel reduction request carries the identifiers of some failed channels on the link, the device can generate and send multiple first control messages carrying different channel reduction requests, with different channel reduction requests carrying different identifiers of failed channels. By carrying the identifiers of failed channels, the channel reduction request enables the device receiving the request to accurately locate the failed channel on the link, and then process the failed channel, improving the accuracy of fault handling.

[0079] One approach is to send a channel downgrade request carrying first channel information and a first quantity. The first channel information indicates the first channel in the sequentially arranged channels on the link to fail, and the first quantity refers to the number of channels still in use on the link. Correspondingly, the channel downgrade request instructs the discontinuation of a second quantity of channels following the channel indicated by the channel information. This second quantity of channels includes all failed channels on the link. The second quantity is obtained by subtracting the first quantity from the total number of channels on the link, representing the number of channels no longer in use on the first link. The total number of channels is the total number of channels on the link whose service direction is the same as the failed channel's service direction. For example, if the link includes four channels, lanes 1 to 4, and lanes 3 and 4 fail, then the first channel information indicates lane 3, the first quantity is 2, the second quantity is 2, and the channel downgrade request instructs the discontinuation of the two channels following lane 3, i.e., discontinuing the use of lanes 3 and 4. For example, if lanes e2 and lane4 on the above link fail, the first channel information indicates lane2, the first quantity is 1, the second quantity is 3, and the channel drop request indicates to stop using the three channels from lane2 onwards, that is, stop using lanes 2 to lane4.

[0080] One approach is to include a channel downgrade request carrying the aforementioned first channel information and a second quantity, where the second quantity refers to the number of channels on the link that are no longer in use. Accordingly, the channel downgrade request instructs the discontinuation of a second quantity of channels following the channel indicated in the channel information. This second quantity of channels includes all failed channels on the link. For example, if the link includes four channels, lanes 1 to 4, and lanes 3 and 4 are failed, then the first channel information indicates lane 3, the second quantity is 2, and the channel downgrade request instructs the discontinuation of the two channels following lane 3, i.e., discontinuing lanes 3 and 4. As another example, if lanes 2 and 4 on the same link are failed, then the first channel information indicates lane 2, the second quantity is 3, and the channel downgrade request instructs the discontinuation of the three channels following lane 2, i.e., discontinuing lanes 2 to 4.

[0081] One approach is to send a channel downgrade request carrying second channel information and a first quantity. The second channel information indicates the first normal channel (i.e., the channel without failure) in the sequentially arranged channels on the link. Accordingly, the channel downgrade request instructs continued use of the first quantity of channels following the channel indicated by the second channel information, while discontinuing the use of the remaining channels on the link. The discontinued channels include all failed channels on the link. It should be noted that both the continued-use and discontinued channels are those whose service direction is the same as that of the failed channels. For example, if the link includes four channels, lanes 1 to 4, and lanes 3 and 4 fail, then the second channel information indicates lane 1, the first quantity is 2, and the channel downgrade request instructs continued use of the two channels following lane 1, while discontinuing the use of the remaining channels on the link. That is, lanes 1 and 2 are used, while lanes 3 and 4 are discontinued. For example, if lanes 2 and 4 on the aforementioned link fail, the second channel information indicates lane 1. The first quantity is 1, and the channel reduction request indicates that the channel following lane 1 should be discontinued, and the remaining channels on the link should be discontinued. That is, lane 1 should continue to be used, while lanes 2 through 4 should be discontinued. This method, by indicating the first normal channel on the link, allows the sender to reduce the number of channels in use on the link according to the channel reduction request, ensuring that the channels still in use are all fault-free channels, thereby improving the reliability of data transmission.

[0082] Another approach is to send a downchannel request carrying the aforementioned second channel information and the aforementioned second quantity, which will not be elaborated upon here in the embodiments of this application.

[0083] One approach is to include a channel downgrade request carrying third channel information and the aforementioned first quantity. The third channel information indicates the first faulty channel following the normal channels on the link, assuming multiple channels are arranged sequentially. For example, if the link has four channels (lane1 to lane4), and lane3 and lane4 fail, the third channel information indicates lane3. Similarly, if lane2 and lane4 fail, the third channel information indicates lane2 and lane4. The device receiving the channel downgrade request can deduce the channels to be deactivated based on the third channel information and the first quantity. For instance, if the link has four channels (lane1 to lane4), the third channel information indicates lane2 and lane4, and the first quantity is 2, then the device receiving the channel downgrade request can determine that lane2 and lane4 are the channels to be deactivated on the first link. This method, by indicating the first faulty channel following the normal channel on the link, enables the sender to, in the event of intermittent failures of multiple channels on the link, to stop using the faulty channels while continuing to use as many unfailed channels as possible. This improves data transmission reliability while maximizing the bandwidth of the link after the channel downgrade, ensuring data transmission speed. Here, "intermittent failures of multiple channels on the link" refers to a situation where there is an unfailed channel between two faulty channels.

[0084] The aforementioned channel reduction request can also indicate the channel to be deactivated in various other ways. For example, the channel reduction request may include the aforementioned third channel information and the aforementioned second quantity. Alternatively, the channel reduction request may include fourth channel information and the aforementioned second quantity, where the fourth channel information indicates the first normal channel following the faulty channel on the link, in the case of multiple channels arranged sequentially on the link. Another example is that the channel reduction request may include the aforementioned fourth channel information and the aforementioned first quantity. Channel reduction requests implemented in these various ways indicate that the number of channels in use on the link is reduced to a preset number, that is, that the link width is reduced to a specified width, in order to deactivate the faulty channel and improve the reliability of data transmission. These embodiments will not be elaborated upon further here. It should be noted that in some embodiments, the aforementioned channel reduction request also includes functional information, which is used to control the transmitting interface corresponding to the channel on the link to be in a closed state (i.e., indicating TX to be disabled).

[0085] The above describes the implementation of a channel reduction request that indicates a decrease in the number of channels in use on the link, including channels that have failed. In another implementation, the channel reduction request indicates a channel to be deactivated. Correspondingly, each channel on the link corresponds to one or more other channels. The receiving end sends a channel reduction request to the sending end based on the channel corresponding to the failed channel. The sending end determines the failed channel based on the channel to which the channel reduction request is received. For example, the link includes eight channels, lanes 1 to 8. Lanes 1 to 4 are used for data transmission from the sending end to the receiving end, and lanes 5 to 8 are used for data transmission from the receiving end to the sending end. Lanes 1 to 4 correspond to lanes 5 to 8, respectively. If lanes 2 and 4 fail, the receiving end sends a channel reduction request to the sending end via lanes 6 and 8. The sending end determines that lanes 2 and 4 have failed based on the channel to which the channel reduction request is received.

[0086] The implementation of the first control information used to carry the aforementioned landing request is described below.

[0087] During the transmission of the aforementioned first control information, the first control information can be transmitted in the form of a bitstream. This bitstream can be implemented in various structures. One structure includes a start field and a control field. The start field identifies the beginning of the bitstream, and the control field, also known as the payload, carries the aforementioned first control information (physical layer control information) to control the link's state. The length of the control field is preset. The receiving device can determine the start position of the bitstream based on the start field and the end position based on the lengths of the start and control fields. Another structure includes a start field, a control field, and an end field. This structure does not limit the length of the control field, and the end field identifies the end position of the bitstream. The receiving device can determine the start position of the bitstream based on the start field and the end position based on the end field, and then extract the control field from the bitstream based on these start and end positions. In the above structure, the end field is located at the end of the bitstream. Another structure includes a start field, an end field, and a control field, with the end field located between the start field and the control field. In this structure, the length of the control field is preset. The device receiving the bitstream can determine the start position of the bitstream based on the start field, and determine the end position based on the lengths of the end field and the control field. Then, based on the position of the end field and the end position of the bitstream, the control field can be extracted from the bitstream. This application does not limit the structure of the bitstream. By including a start field, or a start field and an end field, in the bitstream, the physical layer of the device can determine and lock the frame boundaries of the received bitstream, that is, accurately identify the start and end positions of the bitstream, so as to distinguish the bitstream carrying the first control information from the bitstream carrying other data in subsequent processing.

[0088] The following Figure 5 This is a schematic diagram of a bitstream for transmitting first control information provided in an embodiment of this application. This bitstream can be referred to as an alignment marker (AM) bitstream, such as... Figure 5As shown, the AM bitstream includes a start field (AM_Identify / AM_ID0-AM_IDN), an end field (AM_END), and a control field (AM_CTRL), where N is an integer greater than 0. The control field further includes a channel field (AM_CTRL_LaneID), a function field (AM_CTRL_Type), and a quantity field (AM_CTRL_Detail). The channel field, function field, and quantity field are used to carry the channel information, function information, and quantity included in the aforementioned downscaling request, respectively. The length of the start field can vary within a preset range as N changes.

[0089] The encoding process of the AM code stream described above will be explained below.

[0090] In some embodiments, the encoding process employs an error-correcting mechanism, such as using error-correcting code encoding to encode the above-mentioned... Figure 5 The various fields in the bitstream shown, such as Hamming code encoding, forward error correction encoding, and eBCH (extended Bose Ray-Chaudhuri Hocquen ghem) encoding, give the encoded bitstream strong error detection and correction capabilities, enabling accurate and reliable information transmission over the physical link between two devices at a preset bit error rate. Taking eBCH encoding as an example, Table 1 below shows an eBCH encoding set, which includes 32 codes, CW0 to CW31. Each code is constructed as eBCH-16, indicating that each code includes 16 bits, Bit0 to Bit15. These 16 bits include BCH(16,5) and 1 bit even parity, where 16 (bits) is the length of the code, 5 (bits) is the length of the encoded payload, and 1 (bit) is the length of the parity code. Each code in Table 1 above is generated by the pre-set generator polynomial shown in the following formula (1):

[0091] g(x) = x 10 +x 8 +x 5 +x 4 +x 2 +x+1 (1)

[0092] Here, x is used as an abstract symbol to define the generator polynomial. The Hamming distance between any two codes in this eBCH encoding set is 8, meaning that the number of different bits between any two codes is 8, which makes the accuracy of error detection and correction based on the codes in the above eBCH encoding set relatively high.

[0093] Table 1

[0094]

[0095] In some embodiments, among the 32 codes shown in Table 1 above, some codes are less affected by the transmission line during transmission. Using these codes to transmit data can improve the accuracy of data transmission. For example, Table 2 below shows some codes that are less affected by the transmission line when transmitting data through a line that meets DC equalization or through modulation lines such as NRZ or PAM4.

[0096] Table 2

[0097]

[0098]

[0099] As shown in Table 2, the codes CW8, CW28, and CW3 in Table 1 are less affected by lines that meet DC equalization, NRZ, and PAM4 lines. Taking CW8 in Table 1 as an example, the code of CW8 is 01000111_10101100. When CW8 is transmitted through NRZ lines, the code is 00100111_11011000. When CW8 is transmitted through PAM4 lines, the code is 0312_2130.

[0100] Table 3 below shows a specific example of an AM bitstream, as shown in Table 3:

[0101] Table 3

[0102]

[0103] In Table 3, bits 0 to (4*N-1) of the AM code stream represent the positions of the start field. In this example, N ranges from 1 to 5. The AM code stream uses CW21 and CW28 to encode the start field. Taking N as 3 as an example, CW21 and CW28 are concatenated together to form an AM_ID. This start field includes three repeated AM_IDs. Constructing the start field using the above encoding method enables it to have a certain fault tolerance. The device receiving the AM code stream can perform error detection and correction on the received data according to the error detection and correction principles of the corresponding encoding method, ensuring that the AM code stream can be accurately identified from the received data. Bits 4*N to (4*N+3) of the AM code stream represent the positions of the end field. The AM code stream uses CW22 and CW8 to encode the end field. The end field is similar to the start field described above, and will not be repeated here in this embodiment. In the AM bitstream, bits (4*N+4) to (4*N+27) are the locations of the control fields. Bits (4*N+4) to (4*N+11) are the locations of the channel fields, bits (4*N+12) to (4*N+19) are the locations of the function fields, and bits (4*N+20) to (4*N+27) are the locations of the quantity fields. The length of the channel fields, function fields, and quantity fields is 64 bits.

[0104] At different positions in the AM code stream, different codes correspond to different fields, and different fields carry different information. For example, as shown in Table 4 below, at the position of the function field, the code with bits 0 to 15 being CW23, bits 16 to 31 being CW22, bits 32 to 47 being CW23, and bits 48 to 63 being CW22 is the function field. The function information carried by this function field indicates that the transmitting interface corresponding to the control channel is in a closed state (i.e., Remote TX Link Width switch indicator, requesting the other end to switch the link width). For example, in the encoding, bits 0 to 15 are CW28, bits 16 to 31 are CW23, bits 32 to 47 are CW28, and bits 48 to 63 are CW23. This encoding is another field, and the information carried by this field indicates that the receiving interface corresponding to the control channel is in a closed state (i.e., Remote RX Link Width switch indicator, requesting the other end RX to switch the link width).

[0105] Table 4

[0106]

[0107] For example, as shown in Table 5 below, different codes correspond to different quantity fields at the location of the quantity field. Different quantity fields carry different quantities (link width information). For example, the code where bits 0 to 15 are CW9, bits 16 to 31 are CW10, bits 32 to 47 are CW21, and bits 48 to 63 are CW22 is a quantity field. This quantity field carries a quantity of 2, indicating that the link width should be adjusted to X2.

[0108] Table 5

[0109]

[0110] It should be noted that, in some embodiments, the above-mentioned codes and the information indicated by each code are pre-stored in the device so that the device can identify the code from the data received through the channel and determine the information indicated by the code.

[0111] Next, combine Figures 6 to 8 The text provides an illustrative example of a scenario where some channels on the link fail, combined with... Figure 9 An illustrative example is provided for a scenario where all channels on the link fail.

[0112] The following Figure 6 This is a data interaction diagram of a link fault handling method provided in an embodiment of this application. The following is in conjunction with... Figure 6 The link fault handling method provided in the embodiments of this application will be described by way of example. In the following... Figure 6 In the example shown, data transmission between two devices, NodeA and NodeB, is conducted via link 1. Link 1 includes four channels, from laneA1 to laneA4. This example illustrates how laneA3 and laneA4 fail during data transmission, causing NodeA and NodeB to stop using laneA3 and laneA4. Figure 6 As shown, the method includes the following steps.

[0113] 601. Node A transmits data to Node B through lane A1 to lane A4 on link 1. The transmitted data periodically contains a first bit stream. Link 1 includes lane A1 to lane A4.

[0114] NodeA and NodeB are implemented as described above. Figure 1In the illustrated implementation environment, the first device 101 and the second device 102 are shown, with Node A as the transmitter and Node B as the receiver. The channels lane A1 to lane A4 included in link 1 are channels in the first service direction, meaning lane A1 to lane A4 are used to transmit data from Node A to Node B. The data transmitted from Node A to Node B includes service data generated in the application layer of Node A, such as service data corresponding to applications installed on Node A. The periodic inclusion of a first bitstream in the transmitted data means that during the transmission of the aforementioned service data, Node A sends a first bitstream every time it sends service data of a first preset data size. The first bitstream is pre-encoded in the physical layer of Node A, and this first bitstream can be implemented as described above. Figure 5 The AM bitstream shown is not limited to the embodiments of this application.

[0115] In this embodiment of the application, NodeA splits the service data into multiple data segments and transmits these multiple data segments to NodeB through lanes A1 to A4 on link1. During the transmission process, for each channel on link1, NodeA sends a first bitstream for each service data segment of a first preset data size sent through that channel.

[0116] In some embodiments, during the above process, the application layer of Node A generates service data. Node A transmits the service data to its physical layer through the presentation layer, session layer, transport layer, network layer, and data link layer. Upon receiving the service data, the physical layer of Node A generates a first data stream and sends it to a serializer / deserializer. The serializer / deserializer splits the received service data into multiple segments and sends these segments to lanes A1 to A4 on link1 via multiple transmission interfaces. These multiple data segments are then transmitted to Node B via lanes A1 to A4. During the data transmission to the channels, for each channel on link1, the physical layer sends one first data stream for each instance of service data of a first preset data size transmitted through that channel.

[0117] Step 601 above is a possible implementation of the sending end periodically sending the first code stream through the channel on the link. This possible implementation is illustrated by taking the sending end periodically sending the first code stream during the transmission of service data as an example. In some embodiments, the sending end only periodically sends the first code stream through the channel on the link. This application embodiment does not limit this.

[0118] 602. NodeB receives data through lanes A1 to A4 on link1. Based on the status of receiving the first bitstream through lanes A1 to A4 on link1, it determines that lanes A3 and A4 are faulty.

[0119] Specifically, for each channel from laneA1 to laneA4 on link1, the status of receiving the first bitstream through that channel refers to the time elapsed between receiving one first bitstream through that channel and receiving the next first bitstream through that channel. If this time elapsed is greater than or equal to the preset detection time, then that channel has failed.

[0120] In this embodiment, for each channel on the link, the NodeB starts timing whenever it receives a first bitstream from that channel. If the timing exceeds a preset detection time and no next first bitstream is received from that channel, the NodeB determines that the channel has failed. Correspondingly, for each channel on link1, the NodeB receives data transmitted through that channel, identifies the data according to the encoding corresponding to the first bitstream, and starts timing each time a first bitstream is identified from the data. If the timing exceeds a preset detection time and no next first bitstream is received through that channel, the NodeB determines that the channel has failed. In the above process, for each channel from lane A1 to lane A4 on link1, the NodeB starts timing for that channel whenever it receives a first bitstream through that channel. For lane A3 and lane A4, if the NodeB's timing exceeds a preset detection time and no next first bitstream is received through these two channels, it determines that lane A3 and lane A4 have failed.

[0121] In some embodiments, during the above process, for each receive interface on the NodeB's serializer / deserializer, the NodeB receives data transmitted through the channel connected to that receive interface. The NodeB's physical layer identifies the sub-data according to the encoding corresponding to the first bitstream. For each first bitstream identified from the sub-data, the first bitstream is extracted from the sub-data, thereby obtaining data that does not contain the first bitstream. The NodeB's serializer / deserializer reconstructs the aforementioned service data based on multiple segments of data from lanes A1 to A4 that do not contain the first bitstream.

[0122] In some embodiments, the physical layer of the NodeB performs the above-described process of determining faults in laneA3 and laneA4, which is not limited in this application embodiment.

[0123] Step 602 above is a possible implementation of the receiving end receiving the first bitstream transmitted through the channel on the link and determining whether the channel has failed based on the status of receiving the first bitstream through the channel. This possible implementation is for each channel on the link. By judging whether the first bitstream can be received periodically through the channel, it is determined whether the channel has failed. It can detect faults at the channel level and quickly locate the faulty channel on the link. Compared with detecting faults at the link level, the sensitivity of fault detection is higher.

[0124] It should be noted that in steps 601 and 602 above, the sending end periodically sends the first code stream to the receiving end, and the receiving end determines whether the channel has failed by judging whether it can periodically receive the first code stream. These two periods may be the same or different periods, and this application embodiment does not limit this.

[0125] 603. NodeB transmits data to NodeA. The transmitted data periodically includes first control information. The first control information carries a lane downgrade request, which indicates that lane A3 and lane A4 should be stopped.

[0126] The data transmitted from NodeB to NodeA includes service data generated in the application layer of NodeB, such as service data corresponding to applications installed on NodeB. The periodic inclusion of first control information in the transmitted data means that during the transmission of the aforementioned service data, NodeB sends one piece of first control information every time it sends a second preset amount of service data. The channel drop request carried by this first control information can instruct the cessation of use of lanes A3 and A4 through any of the aforementioned implementation methods. In this embodiment, the channel drop request carries the aforementioned first channel information and a first quantity as an example. The first channel information carried by the channel drop request indicates lane A3, and the first quantity is 2. Correspondingly, the channel drop request instructs the cessation of use of lanes A3 and A4.

[0127] In this embodiment, the process of NodeB transmitting data to NodeA is the same as step 601 described above, and will not be repeated here. It should be noted that the period involved in step 603 may be the same as or different from the period involved in step 601, and may also be the same as or different from the period involved in step 602. This embodiment does not limit this.

[0128] In some embodiments, during the transmission of data from NodeB to NodeA, NodeB transmits data to NodeA through a channel in the second service direction on link1, or NodeB transmits data to NodeA through a channel in the second service direction on other links. This application embodiment does not limit this. For example, as... Figure 4As shown, NodeB transmits data to NodeA via lanes 4 and 5 on link 1, or NodeB transmits data to NodeA via lanes 6 and 7 on link 2.

[0129] In some embodiments, NodeB sends first control information to NodeA through multiple channels on the link. Correspondingly, the data transmitted through these multiple channels periodically includes the first control information, which can ensure the reliability of the transmission of the first control information and thus improve the accuracy of adjusting the link based on the first control information.

[0130] In some embodiments, NodeB sends first control information to NodeA through a preset number of preset channels among multiple channels on the link. Correspondingly, the data transmitted through the preset number of preset channels periodically includes the first control information, which can reduce the bandwidth occupation of the first control information and improve bandwidth utilization efficiency.

[0131] In some embodiments, the NodeB sends the same first control information through different channels on the link. Correspondingly, the data transmitted through different channels on the link periodically includes the same first control information. For example, if the first control information carries a channel downgrade request that includes the identifier of the faulty channel on the link, the NodeB generates first control information carrying a channel downgrade request based on the faulty lanes A3 and A4 on link 1. This channel downgrade request carries the identifiers of all faulty channels on link 1, i.e., the identifiers of lanes A3 and A4. The NodeB sends this first control information carrying the channel downgrade request to the NodeA through two different channels. When the NodeB sends the same first control information through different channels on the link, the reliability of the first control information transmission can be improved, thereby improving the accuracy of link adjustments based on this first control information.

[0132] In some embodiments, the NodeB sends different first control information through different channels on the link. Correspondingly, the data transmitted through different channels periodically includes different first control information. For example, if the first control information carries a channel downscaling request that includes the identifier of the faulty channel on the link, the NodeB generates two first control messages carrying different channel downscaling requests based on the faulty lanes A3 and A4 on link 1. One downscaling request carries the identifier of lane A3, and the other carries the identifier of lane A4. The NodeB sends these two first control messages with different downscaling requests to the NodeA through two different channels. When the NodeB sends the same first control information through different channels on the link, the reliability of the first control information transmission can be improved, thereby improving the accuracy of link adjustments based on the first control information. When the NodeB sends different first control information through different channels on the link, more first control information can be transmitted to the NodeA, improving data transmission efficiency.

[0133] In some embodiments, the NodeB sequentially sends different first control information through the same channel on the link, and this application embodiment does not limit this.

[0134] In some embodiments, the physical layer of the NodeB performs the above-described process of generating and sending the first control information, and this application embodiment does not limit this process.

[0135] Steps 602 to 603 above represent a possible implementation of the receiving end sending first control information if at least one channel on the link is determined to be faulty. The first control information indicates an adjustment to the number of channels in use on the link, and the fault is determined based on the status of receiving the first bitstream through a channel on the link. In this possible implementation, the receiving end detects faults at the channel level by detecting whether it can periodically receive the first bitstream through a channel. Upon detecting a fault, the receiving end promptly informs the sending end of the fault by sending the first control information, enabling the sending end to handle the fault in a timely manner, resulting in a fast fault response.

[0136] 604. Node A receives data and, based on the downlink request carried in the first control information of the received data, stops using lane A3 and lane A4.

[0137] In this embodiment, Node A receives data from Node B through a corresponding channel. Based on the encoding corresponding to the first control information, it identifies the received data and extracts the aforementioned first control information. Based on the channel reduction request carried by the first control information, Node A stops using lanes A3 and A4 on link 1. This process is a possible implementation where the transmitting end receives the first control information and adjusts the number of channels in use on the link based on the first control information, which is determined based on the reception status of the first bitstream. In this possible implementation, the transmitting end reduces the number of channels in use on the link based on the channel reduction request carried by the first control information, including channels that have failed on the link. This avoids using failed channels for data transmission, improving the reliability of data transmission.

[0138] In the above process, taking a bitstream used to transmit the first control information, which includes a start field, an end field, and a control field as an example, Node A receives data from Node B through the corresponding channel. Based on the encoding corresponding to the start field in the bitstream, Node A identifies the received data to determine the start position of the bitstream. Then, based on the encoding corresponding to the end field, Node A identifies the data after the start field to determine the end position of the bitstream. Then, based on the start and end positions, Node A extracts the control field carrying the first control information from the data, identifies the encoding of the control field to obtain the first control information, and then, based on the channel drop request carried by the first control information, stops using lanes A3 and A4 on link 1.

[0139] Step 604 above is illustrated using the example of NodeA stopping the use of the channel indicated by the channel downgrade request carried in the first control information. In some embodiments, after receiving the channel downgrade request for the first time, NodeA temporarily does not process the channel indicated by the request. Starting from the first received channel downgrade request, NodeA counts the received channel downgrade requests. If the number of received channel downgrade requests reaches a preset number, NodeA stops using the channel indicated by the request. If the number of received channel downgrade requests does not reach the preset number, NodeA continues to use the channel indicated by the request. By counting the channel downgrade requests received by NodeA, the above process allows time for automatic recovery of the channel indicated by the request, avoiding additional fault handling for the channel indicated by the request and improving fault handling efficiency.

[0140] 605. NodeA transmits data to NodeB through laneA1 and laneA2, which are in use on link1. The transmitted data periodically includes second control information, which indicates that the reduced channels on link1 should be discontinued.

[0141] The second control information is similar to the first control information described above, and will not be repeated here in the embodiments of this application.

[0142] In this embodiment, NodeA transmits data to NodeB through laneA1 and laneA2, which are in use on link1. This data transmission process is one possible implementation of the sending end sending second control information, which instructs the discontinuation of the reduced channels on the link. This process is similar to step 601 described above, and will not be repeated here.

[0143] 606. The NodeB receives data and, based on the second control information in the data, stops using the channel indicated by the second control information.

[0144] This process is one possible implementation whereby the receiving end stops using the channel indicated by the second control information in response to receiving the second control information. This process is similar to step 604 above, and will not be described again in the embodiments of this application.

[0145] In the above process, NodeB performs channel drop-down only after receiving the second control information from NodeA, which can avoid NodeB data reception errors caused by premature channel drop-down and improve the reliability of data transmission.

[0146] In some embodiments, after NodeB stops using lanes A3 and A4 on link1, it starts timing. After the timing reaches a first preset duration, NodeB sends a first channel upgrade request to NodeA. This first channel upgrade request indicates that lanes A3 and A4 be resumed, thereby increasing the number of channels in use on link1 to four. The implementation and transmission process of this first channel upgrade request are similar to those of the channel downgrade request described above. For example, the first channel upgrade request carries the first channel in use on the link and the number of channels in use on the link, or it carries the identifier of the channel in use on the link. This will not be elaborated further in this embodiment. The above process is a possible implementation where the receiving end starts timing in response to a decrease in the number of channels in use on the link. If the timing reaches a first preset duration, it sends a first channel upgrade request. This first channel upgrade request indicates an increase in the number of channels in use on the link, which can increase the number of channels in use on the link after a period of channel downgrade, thereby improving data transmission bandwidth and accelerating data transmission speed.

[0147] In some embodiments, after receiving the first channel upgrade request, NodeA resumes the use of lanes A3 and A4, which were previously discontinued, and transmits data to NodeB via lanes A1 to A4 on link 1. A third channel upgrade request is periodically inserted into the transmitted data, indicating the resumption of use of lanes A3 and A4, thereby increasing the number of channels in use on link 1 to four. After receiving the third channel upgrade request from NodeA, NodeB resumes the use of lanes A3 and A4, which were previously discontinued, and receives data transmitted by NodeA via lanes A1 to A4 on link 1.

[0148] In some embodiments, after Node A stops using lanes A3 and A4 on link 1, it starts timing. After the timing reaches a first preset duration, Node A sends a first channel upgrade request to Node B. During the process of Node A sending the first channel upgrade request to Node B, in some embodiments, Node A directly resumes using lanes A3 and A4, sending the first channel upgrade request to Node B through any one or more channels on link 1. In other embodiments, Node A first determines that lanes A3 and A4 have returned to normal before sending the first channel upgrade request to Node B to ensure data transmission reliability. Accordingly, Node A periodically sends the aforementioned first bitstream to Node B through lanes A3 and A4 on link 1. Upon periodically receiving the first bitstream through lanes A3 and A4, Node B determines that lanes A3 and A4 have returned to normal and sends third control information to Node A, indicating the resumption of use of lanes A3 and A4 on link 1. NodeA receives the third control information, resumes the use of lanes A3 and A4, and sends a first uplink request to NodeB. This application embodiment does not limit this aspect.

[0149] In the link fault handling method provided in this application embodiment, the sending end transmits a first code stream to the receiving end through a channel on the link. The receiving end determines whether the channel has failed based on the status of receiving the first code stream through the channel on the link. If the channel fails, the receiving end sends first control information to the sending end so that the sending end can adjust the number of channels in use on the link according to the first control information to stop using the failed channel on the link. This achieves link fault detection and handling at the channel level. Compared with fault handling by temporarily stopping the use of the link when detecting and handling faults at the link level, this method can maintain data transmission during fault handling, thereby avoiding service interruption, improving the reliability of data transmission, and the fault handling speed is faster.

[0150] The following is combined Figure 7 Regarding the above Figure 6 The fault handling process of the illustrated link is explained exemplarily. For example... Figure 7 As shown, NodeA sends service data (data flit) to NodeB through the four channels on link1, periodically inserting a first bitstream into the service data (this first bitstream can be implemented as described above). Figure 5(The AM stream shown is shown). When lanes A3 and A4 on link 1 fail, NodeB detects a period of no reception of the first stream via lanes A3 and A4, confirming a failure. NodeB sends the AM stream (TX AM) corresponding to the first control information to NodeA. This AM stream carries a lane reduction request, requesting NodeA to stop using lanes A3 and A4, thus reducing TX to X2. NodeA, upon receiving the lane reduction request, shuts down TX3 and TX4 to stop using lanes A3 and A4, continuing data transmission via lanes A1 and A2, i.e., transmitting data according to the link width of X2. NodeA periodically inserts the AM stream (RX AM) corresponding to the second control information into the transmitted data to request NodeB to reduce RX to X2. NodeB, upon receiving the second control information, shuts down RX3 and RX4, continuing data transmission via lanes A1 and A2, i.e., receiving data according to the link width of X2.

[0151] The above Figure 6 The example shown illustrates how the sender directly stops using the channel indicated by the channel downgrade request. The following... Figure 8 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application. Figure 8 In the example shown, data transmission between the two devices, NodeC and NodeD, is via link 2. Link 2 includes six channels, namely channels B1 to B6. In this example, B1 to B4 are in use, while B5 and B6 are out of use. Data transmission between NodeC and NodeD is via B1 to B4. If B3 and B4 fail, NodeC and NodeD will stop using B3 and B4 and resume using B5 and B6. This ensures data transmission reliability while maximizing bandwidth to accelerate data transmission. Figure 8 As shown, the method includes the following steps.

[0152] 801. NodeC transmits data to NodeD through laneB1 to laneB4 on link2. The transmitted data periodically includes the first bit stream. Link2 includes laneB1 to laneB6, where laneB1 to laneB4 are in use, and laneB5 and laneB6 are out of use.

[0153] 802. NodeD receives data through laneB1 to laneB4 on link2. Based on the status of receiving the first bitstream through laneB1 to laneB4 on link2, it determines that laneB3 and laneB4 are faulty.

[0154] 803. NodeD transmits data to NodeC. The transmitted data periodically includes first control information. The first control information carries a channel drop request, which indicates that laneB3 and laneB4 should be stopped.

[0155] Steps 801 to 803 are the same as steps 601 to 603, and will not be repeated here in the embodiments of this application.

[0156] 804. NodeC receives data and, based on the downlink request carried in the first control information of the received data, stops using lane B3 and lane B4 on link2 and resumes using lane B5 and lane B6 on link2.

[0157] In this embodiment, NodeC receives data from NodeD through a corresponding channel. Based on the encoding corresponding to the first control information, it identifies the received data and extracts the aforementioned first control information. According to the channel reduction request carried in the first control information, NodeC stops using lanes B3 and B4 on link2 and resumes using lanes B5 and B6 on link2. This process is one possible implementation of stopping the use of the channel indicated by the channel reduction request carried in the first control information. In this possible implementation, the transmitting end not only reduces the number of channels in use on the link but also resumes the use of previously reduced channels, which can improve the bandwidth of data transmission while ensuring data transmission reliability, thereby accelerating the data transmission speed.

[0158] 805. NodeC transmits data to NodeD via laneB1, laneB2, laneB5 and laneB6 on link2. The transmitted data periodically includes second control information, which indicates that the reduced channels on link2 should be discontinued.

[0159] The second control information is similar to the first control information described above, and will not be repeated here in the embodiments of this application.

[0160] This process is one possible implementation of NodeC sending a second control message, which instructs NodeC to stop using the reduced channel. In some embodiments, if lane B5 or lane B6 on link2 has not recovered, taking lane B5 as an example, Node D determines that lane B5 is faulty based on the status of the second control message received through lane B5, and requests NodeC to stop using lane B5 to avoid restoring the use of the faulty channel and improve the reliability of data transmission.

[0161] 806. NodeD receives data and, based on the second control information in the data, stops using lane B3 and lane B4 on link2 and resumes using lane B5 and lane B6 on link2.

[0162] Steps 805 to 806 described above are similar to steps 605 to 606 described above, and will not be repeated here in this embodiment. It should be noted that in this embodiment, while NodeD stops using laneB3 and laneB4 according to the second control information, it also resumes the use of previously stopped channels laneB5 and laneB6 to improve bandwidth and speed up data transmission.

[0163] 807. NodeC resumes using lanes B3 and B4 on link2 and transmits data to NodeD via lanes B1 to B6 on link2. The transmitted data periodically includes fourth control information, which indicates the resumption of use of lanes B3 and B4.

[0164] The fourth control information is similar to the first control information described above, and will not be repeated here in this embodiment. The above process is similar to step 601 described above, and will not be repeated here in this embodiment.

[0165] 808. NodeD receives data and, based on the fourth control information in the data, resumes the use of lane B3 and lane B4 on link2.

[0166] The process is the same as step 605 above, and will not be described again in the embodiments of this application.

[0167] It should be noted that steps 807 to 808 above are optional steps. These optional steps are one possible implementation of the sending end automatically restoring the channel indicated by the downlink request. Since most channels that have failed on the link have automatic recovery capabilities, and the automatic recovery speed is relatively fast, the sending end can improve the data transmission bandwidth and speed as quickly as possible while ensuring data transmission reliability by restoring the use of the channel that failed on the link. Of course, the above-mentioned process of automatically restoring the use of the previously failed channel can also be initiated by the receiving end, and this application embodiment does not limit this.

[0168] The above Figure 8 In the link fault handling method shown, the sending end transmits a first code stream to the receiving end through the channel on the link. The receiving end determines whether the channel has failed based on the status of receiving the first code stream through the channel on the link. If the channel fails, the receiving end sends a first control message to the sending end, causing the sending end to stop using the failed channel on the link according to the first control message. At the same time, it restores the use of the previously stopped channel on the link. This not only maintains data transmission during fault handling, avoids service interruption, and improves the reliability of data transmission, but also increases the bandwidth of data transmission and speeds up data transmission.

[0169] The above Figures 6 to 8 The data interaction process shown is illustrated using the example of a partial channel failure on the first link. The following... Figure 9 This is a data interaction diagram of another link fault handling method provided in the embodiments of this application. The following is in conjunction with... Figure 9 This section describes the scenario where all channels on the first link fail. (The following...) Figure 9 In the example shown, data transmission between two devices, NodeE and NodeF, is conducted via link 3. Link 3 includes four channels, namely channel C1 to channel C4. This example illustrates how, after a failure occurs between channel C1 and channel C4 during data transmission, NodeA and NodeB stop using link 3 and rebuild the link between NodeA and NodeB. Figure 9 As shown, the method includes the following steps.

[0170] 901. NodeE transmits data to NodeF through laneC1 to laneC4 on link3. The transmitted data periodically contains the first bit stream. Link3 includes laneC1 to laneC4.

[0171] 902. NodeF receives data through laneC1 to laneC4 on link3. Based on the status of receiving the first bitstream through laneC1 to laneC4 on link3, it determines that laneC1 to laneC4 have all failed.

[0172] Steps 901 to 902 are the same as steps 601 to 502, and will not be repeated here in the embodiments of this application.

[0173] 903. After NodeF sends a link break request to NodeE, it stops using link3. The link break request indicates that link3 should be stopped.

[0174] In some embodiments, the implementation of the linkdown request is similar to that of the channel reduction request described above. For example, the linkdown request may indicate reducing the number of channels in use on link3 to 0, or indicate stopping the use of channels that have failed on link3, etc. These details will not be elaborated further in this application. In some embodiments, the linkdown request is transmitted in the form of an RTSB message or the AM stream described above, with the RTSB message carrying the linkdown request. This application does not limit the scope of this implementation.

[0175] In this embodiment of the application, after NodeF sends a link disconnection request to NodeE through the channel of the second service direction on link3, it stops using link3. Alternatively, after NodeF sends a link disconnection request to NodeE through a channel other than link3, it stops using link3.

[0176] In some embodiments, NodeF periodically sends a link termination request to NodeE. After continuously sending the request for a preset duration, NodeF stops sending the link termination request to NodeE and stops using link3. By continuously sending the link termination request to NodeE for a period of time, it is ensured that NodeE successfully receives the link termination request, avoiding the situation where the link termination request is lost and NodeE fails to stop using link3 in time.

[0177] Steps 902 to 903 above describe one possible implementation whereby, if it is determined that all channels on the link have failed, the receiving end sends a link disconnection request to stop using the link. This link disconnection request indicates that the link should be stopped. This possible implementation is illustrated by the receiving end sending a link disconnection request to the sending end. In some embodiments, if it is determined that all channels on the link have failed, the receiving end directly stops using the link. This application does not limit this implementation.

[0178] 904. NodeE receives a link break request and, in response, stops using link3.

[0179] The process is the same as step 504 above, and will not be described again in the embodiments of this application.

[0180] In some embodiments, if NodeF directly stops using link3, NodeE responds to NodeF (the receiver) by stopping using link3 and also stops using link3. Accordingly, NodeE starts timing from the last time it received a response from NodeF via link3. If the timing exceeds a third preset time without receiving a response from NodeF via link3, it determines that link3 has failed and stops using link3.

[0181] It should be noted that in some embodiments, after the physical layers of NodeE and NodeF stop using link3, they shield the upper layers from the connection status of link3, that is, they do not report the disconnection of link3 to the upper layers. In other words, the link training state machine on the physical layer of the computing device is always in a connected state, and the link training state machine shields the upper layers from faults that occur on the link while it is in a connected state. Taking NodeE as an example, since the upper layer (such as the application layer) is unaware that link3 has been disconnected, the upper layer of NodeE will still send data to be transmitted to the physical layer. NodeE temporarily stores this data to be transmitted in the upper layer's buffer, such as the data link layer's buffer. After the link between NodeE and NodeF is re-established, the data in the buffer is sent to NodeF through the newly established link.

[0182] In some embodiments, if the buffer at any layer is full, the data to be transmitted is stored in the buffer of the layer above. For example, if the data link layer buffer is full, the data to be transmitted is stored in the network layer buffer.

[0183] 905. NodeF sends the first control information to NodeE. The first control information carries a link establishment request, which instructs the physical layer to re-establish the link.

[0184] The link establishment request carries information such as the NodeF device identifier, the port identifier on the NodeF, and the number of channels supported by the NodeF. Steps 902 to 905 above represent a possible implementation where, if it is determined that all channels on the link have failed, the receiving end sends first control information carrying a link establishment request, which instructs the physical layer to re-establish the link. This possible implementation, by instructing the physical layer to re-establish the link, enables the link between the sending and receiving ends to recover automatically after a failure, ensuring that the link training state machine on the physical layer remains connected, and preventing upper layers such as the application layer from perceiving the link failure.

[0185] The above process is illustrated using the receiver initiating the link establishment process as an example. In some embodiments, the sender initiates the link establishment process. Accordingly, when the sender stops using the link, it sends the first control information to the receiver. This application does not limit this aspect.

[0186] 906. NodeE receives the first control information and, based on the link establishment request carried in the first control information, re-establishes the link with NodeF, i.e., link4.

[0187] In this application, link4 and link3 may be the same link or different links. This application does not limit this.

[0188] In this embodiment, NodeE receives first control information and, based on the link establishment request carried in the first control information, records the identifier of NodeF and the identifier of the port on NodeF. If the number of channels supported by NodeF in the link establishment request matches the number of channels supported by NodeE, then NodeE sends a link establishment success message to NodeF, indicating that the link between NodeE and NodeF has been successfully established. This link establishment success message carries information such as the identifier of NodeE, the identifier of the port allocated to NodeF on NodeE, the identifier of link4, and the identifier of the channel on link4. If the number of channels supported by NodeF in the link establishment request does not match the number of channels supported by NodeE, then NodeE sends a link establishment failure message to NodeF, carrying the number of channels supported by NodeE. NodeF adjusts the number of channels carried in the link establishment request according to the link establishment failure message and sends the adjusted link establishment request to NodeE to re-establish the link between NodeF and NodeE.

[0189] 907. NodeE sends data to NodeF through multiple channels on link4, and the transmitted data periodically includes the first bit stream.

[0190] The process is the same as step 601 above, and will not be described again in the embodiments of this application.

[0191] In some embodiments, NodeE starts timing in response to successful link re-establishment. If the timing reaches a second preset duration, it sends a second channel upgrade request to NodeF. This second channel upgrade request instructs an increase in the number of channels in use on the re-established link, i.e., increasing the link width of link4 to improve bandwidth and speed up data transmission. The process of sending the second channel upgrade request can also be performed by NodeF. Accordingly, NodeF starts timing in response to successful link re-establishment, and if the timing reaches a second preset duration, NodeF sends a second channel upgrade request to NodeE. This embodiment of the application does not limit this process.

[0192] In the link fault handling method provided in this application embodiment, the sending end transmits a first bitstream to the receiving end through a channel on the link. The receiving end determines whether a channel has failed based on the status of receiving the first bitstream through the channel on the link. If all channels on the link fail, the receiving end sends first control information to the sending end so that the sending end can stop using the link according to the first control information. Afterwards, the link is rebuilt between the physical layers of the sending and receiving ends, and data transmission continues after the link is rebuilt. This effectively shields the upper layer from the link disconnection. Compared with handling faults by temporarily stopping the use of the link, this method allows the upper layer to be unaware of the link disconnection. By temporarily storing data from the upper layer in a buffer and transmitting it after the link is rebuilt, packet loss can be avoided, thus preventing impact on upper-layer services, improving the reliability of data transmission, and enhancing the reliability and availability of cluster interconnection.

[0193] The following is combined with Figure 10 Regarding the above Figure 9 The fault handling process of the illustrated link is explained by way of example. Figure 10 As shown, NodeE sends data (data flit) to NodeF through the four channels on link3, periodically inserting a first bitstream into the data (this first bitstream can be implemented as described above). Figure 5(The AM stream is shown). When lanes C1 to C4 on link 3 fail, NodeF detects a period of no first stream received through lanes C1 to C4, confirming a failure across all lanes. NodeF continuously sends RTSB messages to NodeE for a period before ceasing use of link 3, entering a linkdown state. Simultaneously, NodeF disables TX. At this time, NodeF blocks reporting linkdown information to higher layers to prevent the link down information from being transmitted to upper-layer services. NodeE enters the linkdown state based on received RTSB messages, or automatically enters the linkdown state if it hasn't received an AM stream through link 3 for a period. In this case, NodeE also blocks reporting linkdown information to higher layers to prevent the link down information from being transmitted to upper-layer services. Afterwards, NodeE and NodeF re-establish the link. Once the link is successfully established, NodeE retransmits the buffered data to NodeF. Upon receiving the data, NodeF returns a successful retransmission message to NodeE.

[0194] The following Figure 11 This is a schematic diagram of a physical layer structure provided in an embodiment of this application. Figure 11 The serializer / deserializer, physical layer, and data link layer are shown, such as Figure 11 As shown, the physical layer includes an information generation module (TX_AM_GEN), a detection module (RX_AM_DET), and a control module (LINK_CTRL). Taking the first code stream, first control information, second control information, link break request, and fourth control information all implemented as AM code streams as an example, the information generation module generates the AM code stream, carries a lane reduction request in the AM code stream, and sends the AM code stream to the link. The detection module detects the AM code stream received through the link, parses the lane reduction request and other information carried in the AM code stream, and determines whether a channel on the link has failed based on the AM code stream reception status. The control module controls the number of channels on the link and the link re-establishment process based on the fault information and lane reduction request information detected by the detection module.

[0195] It should be noted that when AM code streams are transmitted cyclically at a certain period, this period can be dynamically adjusted, and this embodiment of the application does not limit this. This method automatically identifies faults occurring on the link by recognizing the periodically transmitted AM code streams, and automatically selects either lane reduction or link re-establishment to recover from the fault based on the fault type (partial channel fault or complete channel fault), thus achieving self-recovery. Furthermore, the above fault recovery process is performed at the physical layer; the link training state machine does not need to enter the link training state during detection and fault recovery, resulting in fast fault recovery without affecting upper-layer services.

[0196] Still Figure 11 As shown, the data link layer includes a buffer (Retry_buffer) and a data control module (Retry_Ctrl). The buffer is used to temporarily store data to be transmitted when the link is disconnected, and the data control module is used to send the data in the buffer to the physical layer after the link is re-established.

[0197] The link fault handling method provided in this application embodiment can also be applied to high-speed serial interfaces other than PCIe interfaces, and this application embodiment does not limit it in this regard.

[0198] The methods of the embodiments of this application have been described above; the apparatus of the embodiments of this application will be described below. It should be understood that the apparatus described below has any of the functions of the computing device in the above methods. (The above is in conjunction with...) Figures 6 to 11 The fault handling method for the link provided according to the embodiments of this application is described in detail. Based on the same inventive concept, the following will be combined with Figure 12 and Figure 13 This application describes a link fault handling apparatus provided according to embodiments of the present application. It should be understood that the technical features described in the method embodiments are also applicable to the following apparatus embodiments.

[0199] See Figure 12 This application provides a link fault handling device, which includes:

[0200] Receiver module 1201 is used to receive the first code stream transmitted through the channel on the link;

[0201] The sending module 1202 is used to send first control information if it is determined that at least one channel on the link has failed. The first control information indicates that the number of channels in use on the link should be adjusted. The failure is determined based on the status of receiving a first bit stream through a channel on the link.

[0202] In some embodiments, the sending module 1202 is used for:

[0203] If it is determined that some channels on the link have failed, a first control message is sent, which carries a channel reduction request. The channel reduction request indicates that the number of channels in use on the link be reduced.

[0204] In some embodiments, the aforementioned channel downgrade request includes channel information and quantity, wherein the channel information indicates the first channel to fail among the sequentially arranged channels on the link, and the quantity refers to the number of channels that continue to be used on the link.

[0205] In some embodiments, the aforementioned channel downgrade request carries an identifier of the channel that has failed on the link.

[0206] In some embodiments, the above-described apparatus further includes:

[0207] The control module is used to stop using the channel indicated by the second control information in response to receiving the second control information.

[0208] In some embodiments, the above-described apparatus further includes:

[0209] The first timing module is used to start timing in response to a decrease in the number of channels in use on the link;

[0210] The sending module 1202 is also used to send a first channel upgrade request if the timing reaches a first preset duration, the first channel upgrade request indicating to increase the number of channels in use on the link.

[0211] In some embodiments, the sending module 1202 is further configured to:

[0212] If it is determined that all channels on the link have failed, a first control message is sent, which carries a link establishment request. The link establishment request instructs the physical layer to re-establish the link.

[0213] In some embodiments, the control module is further configured to stop using the link if it is determined that all channels on the link have failed; the sending module 1202 is further configured to send a link disconnection request to stop using the link if it is determined that all channels on the link have failed, and the link disconnection request indicates that the link should be stopped.

[0214] In some embodiments, the above-described apparatus further includes:

[0215] The second timing module is used to start timing in response to a successful link re-establishment.

[0216] The sending module 1202 is also used to send a second channel upgrade request if the timing reaches a second preset duration. The second channel upgrade request indicates an increase in the number of channels in use on the re-established link.

[0217] In some embodiments, the sending module 1202 is used for:

[0218] For each channel on the link, timing begins whenever the first bitstream is received from that channel;

[0219] If the preset detection time exceeds the limit and no next first bitstream is received from the channel, the channel is considered to have malfunctioned.

[0220] In some embodiments, the link training state machine on the physical layer of the computing device is always in a connected state, and when the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link.

[0221] It should be understood that the fault handling device of the above link corresponds to the second device in the above method embodiment. The modules in the device and the other operations and / or functions described above are respectively for implementing the various steps and methods implemented by the second device in the method embodiment. For specific details, please refer to the above method embodiment. For the sake of brevity, they will not be repeated here.

[0222] See Figure 13 This application provides another link fault handling device, which includes:

[0223] The transmitting module 1301 is used to periodically transmit the first code stream through the channel on the link;

[0224] The receiving module 1302 is used to receive first control information, which indicates that the number of channels in use on the link should be adjusted. The first control information is determined based on the receiving status of the first bit stream.

[0225] The adjustment module 1303 is used to adjust the number of channels in use on the link based on the first control information.

[0226] In some embodiments, the first control information carries a channel reduction request, which indicates a reduction in the number of channels in use on the link. The adjustment module 1303 is used to:

[0227] Based on the channel reduction request carried by the first control information, the number of channels in use on the link is reduced.

[0228] In some embodiments, the aforementioned channel downgrade request includes channel information and quantity, wherein the channel information indicates the first channel to fail among the sequentially arranged channels on the link, and the quantity refers to the number of channels that continue to be used on the link.

[0229] In some embodiments, the aforementioned channel downgrade request carries an identifier of the channel that has failed on the link.

[0230] In some embodiments, the sending module 1301 is further configured to:

[0231] Send a second control message, which instructs to stop using the reduced channels on the link.

[0232] In some embodiments, the adjustment module 1303 is used for:

[0233] Based on the channel reduction request carried by the first control information, the number of channels in use on the link is reduced, and the channels previously reduced on the link are restored to use.

[0234] In some embodiments, the above-described apparatus further includes:

[0235] The first timing module is used to start timing in response to a decrease in the number of channels in use on the link;

[0236] The aforementioned sending module 1301 is also used to send a first channel upgrade request if the timing reaches a first preset duration, the first channel upgrade request indicating an increase in the number of channels in use on the link.

[0237] In some embodiments, the first control information carries a link establishment request, which instructs the physical layer to re-establish the link. The adjustment module 1303 is further configured to:

[0238] Based on the link establishment request carried by the first control information, the link is re-established.

[0239] In some embodiments, the adjustment module 1303 is further configured to:

[0240] In response to the receiving end ceasing to use the link, the link is stopped.

[0241] In response to receiving a link termination request, the above link is stopped. The link termination request indicates that the above link should be stopped.

[0242] In some embodiments, the sending module 1301 is further configured to:

[0243] If the aforementioned link is discontinued, the aforementioned first control information is sent.

[0244] In some embodiments, the above-described apparatus further includes:

[0245] The second timing module is used to start timing in response to a successful link re-establishment.

[0246] The sending module 1301 is also configured to send a second channel upgrade request if the timing reaches a second preset duration, the second channel upgrade request indicating an increase in the number of channels in use on the re-established link.

[0247] In some embodiments, the link training state machine on the physical layer of the computing device is always in a connected state, and when the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link.

[0248] It should be understood that the fault handling device of the above-mentioned link corresponds to the first device in the above-mentioned method embodiment. The modules in the device and the other operations and / or functions mentioned above are respectively for implementing the various steps and methods implemented by the first device in the method embodiment. For specific details, please refer to the above-mentioned method embodiment. For the sake of brevity, they will not be repeated here.

[0249] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a computing device to perform the link fault handling method in the above embodiments. For example, the computer-readable storage medium is a non-transitory computer-readable storage medium, such as read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.

[0250] This application also provides a computer program product or computer program, which includes program code and computer instructions stored in a computer-readable storage medium. A processor in a computing device reads the program code from the computer-readable storage medium and executes the program code, causing the computing device to perform the fault handling method of the above-mentioned link.

[0251] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the link fault handling method in the above method embodiments.

[0252] In this embodiment, the apparatus, device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0253] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the link fault handling method embodiment provided in the above embodiments belongs to the same concept, and its specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0254] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0255] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0256] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0257] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0258] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0259] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0260] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sensitive words involved in this application were obtained with full authorization.

[0261] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0262] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for handling link failures, characterized in that, The method includes: Receive the first bitstream transmitted through the channel on the link; If it is determined that at least one channel on the link has failed, a first control message is sent, the first control message instructing an adjustment to the number of channels in use on the link, the failure being determined based on the status of receiving the first bitstream through the channels on the link.

2. The method according to claim 1, characterized in that, If it is determined that at least one channel on the link has failed, sending the first control information includes: If it is determined that some channels on the link have failed, the first control information is sent, which carries a channel reduction request, indicating that the number of channels in use on the link be reduced.

3. The method according to claim 2, characterized in that, The channel downgrade request includes channel information and a number, wherein the channel information indicates the first channel that has failed in the sequential arrangement of channels on the link, and the number refers to the number of channels that continue to be used on the link.

4. The method according to claim 2, characterized in that, The channel drop request carries the identifier of the channel that has failed on the link.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to receiving the second control information, the use of the channel indicated by the second control information is stopped.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Timing begins in response to a decrease in the number of channels in use on the link; If the timer reaches the first preset duration, a first channel upgrade request is sent, which indicates an increase in the number of channels in use on the link.

7. The method according to claim 1, characterized in that, If it is determined that at least one channel on the link has failed, sending the first control information includes: If it is determined that all channels on the link have failed, the first control information is sent, which carries a link establishment request, and the link establishment request instructs the physical layer to re-establish the link.

8. The method according to claim 7, characterized in that, The method further includes at least one of the following: If it is determined that all channels on the link are faulty, stop using the link; or, If it is determined that all channels on the link have failed, a link disconnection request is sent to stop using the link. The link disconnection request indicates that the link should be stopped.

9. The method according to claim 7 or 8, characterized in that, The method further includes: Timer starts upon successful link re-establishment. If the timer reaches the second preset duration, a second channel upgrade request is sent, which instructs to increase the number of channels in use on the re-established link.

10. The method according to any one of claims 1-9, characterized in that, Determining that at least one channel on the link has failed includes: For each channel on the link, timing begins whenever a first bitstream is received from the channel; If the timer exceeds the preset detection duration and no next first bitstream is received from the channel, it is determined that the channel has malfunctioned.

11. The method according to any one of claims 1-10, characterized in that, The method is applied to a computing device, where the link training state machine on the physical layer of the computing device is always in a connected state. When the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link.

12. A method for handling link failures, characterized in that, The method includes: The first bitstream is periodically sent through the channel on the link; Receive first control information, the first control information indicating to adjust the number of channels in use on the link, the first control information being determined based on the reception status of the first bitstream; Based on the first control information, the number of channels in use on the link is adjusted.

13. The method according to claim 12, characterized in that, The first control information carries a channel reduction request, which indicates a reduction in the number of channels in use on the link. Adjusting the number of channels in use on the link based on the first control information includes: Based on the channel reduction request carried by the first control information, the number of channels in use on the link is reduced.

14. The method according to claim 13, characterized in that, The channel downgrade request includes channel information and a number, wherein the channel information indicates the first channel that has failed in the sequential arrangement of channels on the link, and the number refers to the number of channels that continue to be used on the link.

15. The method according to claim 14, characterized in that, The channel drop request carries the identifier of the channel that has failed on the link.

16. The method according to any one of claims 12-13, characterized in that, The method further includes: Send a second control message, which instructs to stop using the reduced channels on the link.

17. The method according to claim 13, characterized in that, The reduction of channels in use on the link based on the channel reduction request carried by the first control information includes: Based on the channel reduction request carried by the first control information, the number of channels in use on the link is reduced, and the channels previously reduced on the link are restored to use.

18. The method according to any one of claims 12-17, characterized in that, The method further includes: Timing begins in response to a decrease in the number of channels in use on the link; If the timer reaches the first preset duration, a first channel upgrade request is sent, which indicates an increase in the number of channels in use on the link.

19. The method according to claim 12, characterized in that, The first control information carries a link establishment request, which instructs the physical layer to re-establish the link. Adjusting the number of channels in use on the link based on the first control information includes: Based on the link establishment request carried by the first control information, the link is re-established.

20. The method according to claim 19, characterized in that, The method further includes at least one of the following: In response to the receiving end ceasing to use the link, stop using the link; or, In response to receiving a link termination request, the link is stopped from being used, the link termination request indicating that the link should be stopped.

21. The method according to claim 20, characterized in that, The method further includes: If the link is stopped, the first control information is sent.

22. The method according to any one of claims 19-21, characterized in that, The method further includes: Timer starts upon successful link re-establishment. If the timer reaches the second preset duration, a second channel upgrade request is sent, which instructs to increase the number of channels in use on the re-established link.

23. The method according to any one of claims 12-22, characterized in that, The method is applied to a computing device, where the link training state machine on the physical layer of the computing device is always in a connected state. When the link training state machine is in a connected state, it shields the upper layer from faults that occur on the link.

24. A link fault handling device, characterized in that, The device includes: The receiving module is used to receive the first bit stream transmitted through the channel on the link; The sending module is configured to send first control information if it is determined that at least one channel on the link has failed, the first control information indicating an adjustment to the number of channels in use on the link, the failure being determined based on the status of receiving the first bitstream through the channels on the link.

25. A fault handling device for a link, characterized in that, The device includes: The transmitting module is used to periodically transmit the first bit stream through the channel on the link; A receiving module is configured to receive first control information, the first control information indicating an adjustment to the number of channels in use on the link, the first control information being determined based on the receiving status of the first bitstream; The adjustment module is used to adjust the number of channels in use on the link based on the first control information.

26. A computing device, characterized in that, The computing device includes a processor and a memory, the processor being configured to execute at least one piece of program code stored in the memory to enable the computing device to perform the method as described in any one of claims 1 to 23.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which, when executed by a computing device, causes the computing device to perform the method as described in any one of claims 1 to 23.

28. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device performs the method as described in any one of claims 1 to 23.

29. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 23.