Optical interconnect link component, training method of optical interconnect link, and accelerated cluster system

By utilizing optical interconnect link components and training methods, and leveraging the collaborative work of switching components and optoelectronic converters, the problem of PCIe optical interconnect links being unable to detect peer devices has been solved. This enables low-loss, long-distance, and efficient data transmission, improving the success rate of link establishment and system stability.

CN120416707BActive Publication Date: 2026-04-28INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2025-06-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional copper cable interconnect solutions suffer from severe attenuation and distance limitations under PCIe Gen4/5 high-frequency signals, making them unable to support new architectures such as GPU clusters. Furthermore, PCIe optical interconnect links cannot directly detect peer devices, leading to link training failures.

Method used

The optical interconnect link component is used. Through the interaction between the first and second switching components, the interaction of the detection character sequence and the response character sequence is used to ensure that the optical interconnect link enters the training process after the bidirectional communication link is successfully established. This includes the coordinated work of the photoelectric converter and the switching component, and signal processing in accordance with the electrical specifications of the optical interconnect protocol.

Benefits of technology

It enables low-loss, long-distance data transmission, improves the success rate and reliability of link establishment, provides an efficient and stable PCIe optical interconnect data transmission link, and solves the limitations of traditional copper cable interconnect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416707B_ABST
    Figure CN120416707B_ABST
Patent Text Reader

Abstract

The application discloses an optical interconnection link component, an optical interconnection link training method and an acceleration cluster system, relates to the technical field of optical interconnection, and comprises the following steps: when a detection character sequence sent by an opposite end is successfully received by a switching component, the optical interconnection link is allowed to enter a training process, so as to ensure normal communication of the optical interconnection link, and the success rate and reliability of link establishment are improved. Only when the bidirectional communication link of the optical interconnection link is successfully established, actual data transmission function can be realized through the optical interconnection link. The application not only solves the limitations of traditional copper cable interconnection, but also provides an efficient and stable data transmission link based on PCIe optical interconnection, and realizes low-loss and long-distance transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of optical interconnect technology, and in particular to optical interconnect link components, training methods for optical interconnect links, and accelerated cluster systems. Background Technology

[0002] With the development of artificial intelligence, 5G, and cloud computing, data centers are facing extreme challenges in bandwidth, latency, and energy efficiency. Traditional copper cables exhibit severe attenuation and distance limitations (≤3 meters) under high-frequency signals (above 16 GT / s) in PCIe Gen4 / 5, making it difficult to support new architectures such as GPU clusters and compute-in-memory separation. Traditional copper cable interconnect solutions limit the design of PCIe interconnect systems to zero-acquisition and make it difficult to continue to meet the application requirements of GPU cluster scale-up.

[0003] In related technologies, the implementation principle of PCIe (Peripheral Component Interconnect Express) electrical interconnect detection device (controlled end detection) identification logic is to determine the presence of the controlled end by detecting the DC common-mode input impedance of the receiving logic of the peer device. However, due to the presence of optoelectronic conversion devices, the switching components at both ends of the PCIe optical interconnect cannot directly detect the presence of the controlled end device. The DC common-mode input impedance of the link is much greater than the normal value, making it impossible to directly detect the peer device on the optical interconnect link, resulting in the failure of PCIe link training and the inability to carry out subsequent PCIe link training processes. Summary of the Invention

[0004] This application provides an optical interconnect link component, an optical interconnect link training method, an accelerated cluster system, an optical interconnect computing unit, an optical interconnect switching unit, an optical interconnect acceleration unit, media, and program products, to at least solve the problem in related technologies where the photoelectric conversion process changes the form of the signal, causing the DC common-mode input impedance of the link to be much greater than the normal value. Therefore, the optical interconnect link cannot directly detect the peer device like an electrical interconnect, resulting in link training failure.

[0005] This application provides an optical interconnect link component, including: a first photoelectric converter, a second photoelectric converter, a first switching component, and a second switching component, wherein the first photoelectric converter and the second photoelectric converter form an optical interconnect link, and the first switching component and the second switching component interact with each other through the optical interconnect link; the first switching component is used to continuously send a first detection character sequence and send a first response character sequence when it receives a second detection character sequence; the second switching component is used to continuously send a second detection character sequence and send a second response character sequence when it receives the first detection character sequence; after the interaction between the first switching component and the second switching component is completed, the optical interconnect link enters the training process.

[0006] This application provides a training method for an optical interconnect link. The method is applied to a first switching component of the optical interconnect link component described in the above embodiments. The method includes: controlling a first photoelectric converter to continuously send a first detection character sequence, identifying a first electrical signal output by the first photoelectric converter, and determining whether a second response character sequence is received within a preset time period. The second switching component outputs a second response character sequence after receiving the first detection character sequence. After determining that the second response character sequence is received within the preset time period, the method controls the internal state machine of the first switching component to be set to an active state. The active state indicates that the optical interconnect link is allowed to enter the training process.

[0007] This application provides a training method for an optical interconnect link. The method is applied to a second switching component of the optical interconnect link component described in the above embodiments. The method includes: controlling a second photoelectric converter to continuously send a second detection character sequence, identifying a second electrical signal output by the second photoelectric converter, and determining whether a first response character sequence is received within a preset time period. The first switching component outputs a first response character sequence after receiving the second detection character sequence. After determining that the first response character sequence is received within the preset time period, the method controls the internal state machine of the second switching component to be set to an active state. The active state indicates that the optical interconnect link is allowed to enter the training process.

[0008] This application provides an accelerated cluster system, which is composed of optical interconnect link components based on the above embodiments. The system includes: at least one optical interconnect computing unit and at least one optical interconnect acceleration unit, wherein the optical interconnect computing unit provides computing resources and the optical interconnect acceleration unit provides acceleration resources; at least one optical interconnect switching unit, wherein the optical interconnect switching unit is connected to at least one of the optical interconnect computing unit and the optical interconnect acceleration unit via an optical interconnect link, and / or the optical interconnect acceleration unit is connected to at least one optical interconnect computing unit via an optical interconnect link.

[0009] This application provides an optical interconnect computing unit, which is composed of optical interconnect link components in the above embodiments of the optical interconnect computing unit, including: at least one central processing unit and at least one switching component, wherein the central processing unit is connected to at least one switching component, the switching component is connected to at least one photoelectric converter, and the photoelectric converter inside the computing unit and the photoelectric converter outside the computing unit form an optical interconnect link.

[0010] This application provides an optical interconnect switching unit, which is composed of an optical interconnect link component in the above-described embodiment of the optical interconnect computing unit. It includes at least one optical interconnect switching board, wherein the optical interconnect switching board includes a substrate connector and at least one switching component. The substrate connector is connected to at least one switching component. The switching component is connected to at least one high-speed connector and at least one photoelectric converter. The optical interconnect switching board communicates with other optical interconnect switching boards via the high-speed connector. An optical interconnect link is formed between the photoelectric converter inside the optical interconnect switching board and the photoelectric converter outside the optical interconnect switching board.

[0011] This application also provides an optical interconnect acceleration unit, which is composed of optical interconnect link components in the above embodiments of the optical interconnect computing unit, including: at least one graphics processor and at least one switching component, wherein the switching component is connected to at least one graphics processor and at least one photoelectric converter, the photoelectric converter inside the optical interconnect acceleration unit and the photoelectric converter outside the optical interconnect acceleration unit form an optical interconnect link, and at least one graphics processor provides acceleration resources through the optical interconnect link.

[0012] This application also provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described training method for optical interconnect links.

[0013] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described training method for optical interconnect links.

[0014] This application allows the optical interconnect link to enter the training process when the switching components successfully receive the detection character sequence sent by the other party, so as to ensure normal communication of the optical interconnect link, improve the success rate and reliability of link establishment. Only when the bidirectional communication link of the optical interconnect link is successfully established can the actual data transmission function be performed through the optical interconnect link. This not only solves the limitations of traditional copper cable interconnection, but also provides an efficient and stable data transmission link based on PCIe optical interconnection, realizing low loss and long-distance transmission. Attached Figure Description

[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a block diagram of an optical interconnect link component according to an embodiment of this application;

[0017] Figure 2 This application provides a flowchart of a PCIe optical interconnect link training RX detection process;

[0018] Figure 3 A flowchart illustrating a training method for an optical interconnect link according to one embodiment of this application;

[0019] Figure 4 A flowchart illustrating a training method for an optical interconnect link, provided as another embodiment of this application;

[0020] Figure 5 This is a block diagram of the accelerated cluster system according to an embodiment of this application;

[0021] Figure 6 This is a topology diagram of the PCIe optical interconnect switching unit of a 1000-card GPU-accelerated cluster according to an embodiment of this application.

[0022] Figure 7 This is a connection topology diagram of a single PCIe optical interconnect switching unit according to an embodiment of this application;

[0023] Figure 8 This application provides a schematic diagram of a GPU acceleration rack connection topology.

[0024] Figure 9 A PCIe optical interconnect computing unit connection topology diagram provided for embodiments of this application;

[0025] Figure 10 An explanatory diagram of a PCIe optical interconnect switching board provided for embodiments of this application;

[0026] Figure 11 A PCIe optical interconnect acceleration unit connection topology diagram provided for embodiments of this application;

[0027] Figure 12 A kilocalorie GPU-accelerated cluster topology diagram provided in this application embodiment;

[0028] Figure 13 A block diagram illustrating an optical interconnect computing unit provided in an embodiment of this application;

[0029] Figure 14 A block diagram of an optical interconnect switching unit provided in an embodiment of this application;

[0030] Figure 15 A block diagram of an optical interconnect acceleration unit provided in an embodiment of this application;

[0031] Figure 16 This is a GPU acceleration rack connection topology diagram provided for an embodiment of this application.

[0032] Reference numerals: Optical interconnect link component 10, first optoelectronic converter 110, second optoelectronic converter 120, first switching component 130, second switching component 140, accelerated cluster system 20, optical interconnect computing unit 210, optical interconnect acceleration unit 220, and optical interconnect switching unit 230. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0034] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0035] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0036] Due to design flaws in copper cables, the industry is actively promoting research and implementation of PCIe optical interconnect technology. The core of PCIe optical interconnect technology lies in converting PCIe electrical signals into PCIe optical signals. After conversion, these signals can be transmitted via optical fiber. Optical fiber has a low attenuation coefficient; typical attenuation of single-mode fiber at 1310nm wavelength is 0.35dB / km, far lower than the 10dB / m@10GHz of copper cables. Furthermore, optical signals propagate via total internal reflection in optical fiber, avoiding the skin effect, dielectric loss, and electromagnetic interference experienced during electrical signal transmission. Based on these characteristics, PCIe optical signals can achieve low-loss, long-distance transmission in optical fiber.

[0037] Existing optical interconnect technologies are mostly used for data transmission in network switches. PCIe optical interconnect solutions are still in their early stages, with current applications primarily focused on establishing data transmission for single PCIe optoelectronic conversion devices, single PCIe optical interconnect links, and functional verification of single boards or systems supporting optical interconnects. There are no applications of PCIe optical interconnect solutions for cross-resource pool interconnects. In cross-resource pool interconnect applications, copper cable interconnects or network interconnects are still the most common technologies used.

[0038] To address the limitations of copper cable interconnection schemes on PCIe transmission distance and network interconnection schemes on high communication latency due to increased data hop count, this invention relates to an accelerated cluster system based on PCIe optical interconnection technology. This accelerated cluster system includes an optical interconnection computing resource pool, optical interconnection switching units, and optical interconnection acceleration units. The resource pools are connected via switching components that support optical interconnection links. These switching components include a Retimer chip and a Switch chip. The switching components themselves can be used as USP devices or DSP devices. That is, when both ends of the optical interconnection link are switching components, the switching components can bypass the process of monitoring USP devices or DSP devices and directly send device signals to the optoelectronic converter. The optoelectronic converter and optical fiber form multiple PCIe optical interconnection links, supporting low-power, long-distance, and low-latency PCIe data transmission within / between resource pools in a rack. This enables flexible expansion of the GPU accelerated cluster scale based on PCIe communication.

[0039] The optical interconnect computing resource pool, optical interconnect switching unit, and optical interconnect acceleration unit are all composed of optical interconnect link components. Specifically, the PCIe optical interconnect solution mainly consists of switching components supporting the optical interconnect link and LPO (Low Power Optical) optoelectronic conversion devices. In the PCIe optical interconnect link, the PCIe signal master control end requires optoelectronic conversion devices to convert PCIe electrical signals to PCIe optical signals, and the PCIe signal controlled end also requires optoelectronic conversion devices to convert PCIe optical signals to PCIe electrical signals. The optoelectronic conversion process involves the conversion of PCIe signal transmission media, signal properties, signal modulation methods, and signal compensation techniques. These different conversion methods also determine that optoelectronic conversion devices used in network transmission cannot be directly applied to PCIe optical interconnect transmission. For example, optoelectronic conversion devices used in 25G and 100G network switches often have built-in DSP chips that use NRZ-25G modulation. Although PCIe Gen5 and below also use NRZ modulation, unlike NRZ-25G modulation technology which is only optimized for 25G speeds, PCIe NRZ modulation technology needs to optimize signals for different communication speeds of PCIe Gen1-Gen5. In other words, optoelectronic conversion devices used in PCIe optical interconnect applications need to adjust their internal configuration information according to the characteristics of PCIe signals.

[0040] To avoid excessive involvement of the internal DSP chip in the PCIe optical signal modulation process of traditional optoelectronic converters, this solution uses an optoelectronic converter and an external switching component that supports the optical interconnect link to jointly form a PCIe optical interconnect link.

[0041] The switching components supporting optical interconnect links also participate in the PCIe optical interconnect link training process. The difference between optical interconnect link training and electrical interconnect link training lies in the PCIe link training Receiver detection stage: The PCIe electrical interconnect detection device identification logic is implemented by detecting the DC common-mode input impedance of the receiving logic of the peer device to determine whether the controlled end exists. However, due to the presence of optoelectronic conversion devices, the switching components at both ends of the PCIe optical interconnect cannot directly detect the existence of the Receiver device. The link DC common-mode input impedance is much greater than the normal value, making it impossible to directly detect the peer device on the optical interconnect link, resulting in PCIe link training failure and the inability to proceed with the subsequent PCIe link training process.

[0042] To address this issue, embodiments of this application provide an optical interconnect link component, and the system is described in detail below in conjunction with the operation of the optical interconnect link component.

[0043] Specifically, Figure 1 This is a block diagram of an optical interconnect link component provided in an embodiment of this application.

[0044] like Figure 1 As shown, the optical interconnect link component 10 includes: a first photoelectric converter 110, a second photoelectric converter 120, a first switching component 130, and a second switching component 140.

[0045] In this configuration, the first photoelectric converter 110 and the second photoelectric converter 120 form an optical interconnect link, and the first switching component 130 and the second switching component 140 interact through the optical interconnect link. The first switching component 130 is used to continuously send a first detection character sequence and send a first response character sequence when it receives a second detection character sequence. The second switching component 140 is used to continuously send a second detection character sequence and send a second response character sequence when it receives a first detection character sequence. After the interaction between the first switching component 130 and the second switching component 140 is completed, the optical interconnect link enters the training process.

[0046] It is understood that in this embodiment of the application, when the switching components successfully receive the detection character sequence sent by the other party, the optical interconnect link is allowed to enter the training process to ensure normal communication of the optical interconnect link, thereby improving the success rate and reliability of link establishment. Only when the bidirectional communication link of the optical interconnect link is successfully established can the actual data transmission function be performed through the optical interconnect link. This not only solves the limitations of traditional copper cable interconnection, but also provides an efficient and stable data transmission link based on PCIe optical interconnection, realizing low loss and long-distance transmission.

[0047] It should be noted that the first photoelectric converter is located near the first switching component and is used to convert electrical signals from the first device into optical signals for transmission to the other end via optical fiber, and to convert optical signals received from the other end back into electrical signals for processing by the first switching component. The second photoelectric converter is located near the second switching component and is used to convert electrical signals from the second device into optical signals for transmission to the other end via optical fiber, and to convert optical signals received from the other end back into electrical signals for processing by the second switching component. The first switching component is responsible for managing and processing signals from the first device and performing photoelectric signal conversion and transmission through photoelectric conversion devices. The second switching component is responsible for managing and processing signals from the second device and performing photoelectric signal conversion and transmission through photoelectric conversion devices. The first device is the master control device, which is connected to the first switching component, and the second device is the controlled device, which is connected to the second switching component, without specific limitations.

[0048] In this embodiment, the first photoelectric converter 110 and the second photoelectric converter 120 have the same structure. The structure of the first photoelectric converter 110 and the second photoelectric converter 120 includes a linear transimpedance amplifier and a linear driver. The linear transimpedance amplifier is used to convert the optical signal output by the photodetector into an electrical signal with a target linearity. The linear driver is used to convert the electrical signal received by the photodetector into a continuously modulated current to drive the laser to work in the linear region to obtain an optical signal.

[0049] It is understood that the photoelectric converter structures in the embodiments of this application are consistent, making the signal processing paths of the optical interconnect link completely symmetrical in both directions, which improves the system's compatibility and maintainability. The transimpedance amplifier converts the weak current signal output by the photodetector into a voltage signal and amplifies it, thereby improving the sensitivity of the controlled end, enhancing the detection capability of low optical power signals, and improving the data reception quality. The driver converts the electrical signal into an optical signal for transmission, ensuring that the optical signal at the transmitting end has sufficient strength and stability. The drive current can be adjusted according to the link requirements to optimize the balance between power consumption and performance, thus realizing optical signal transmission and reception that conforms to the electrical characteristics of the PCIe protocol and ensuring the smooth operation of link training and data communication.

[0050] In this embodiment, the first switching component 130 processes the first electrical signal output by the first opto-converter 110 based on the electrical specifications of the optical interconnect protocol; the second switching component 140 processes the second electrical signal output by the second opto-converter 120 based on the electrical specifications of the optical interconnect protocol.

[0051] It is understood that, by following the electrical specifications of the optical interconnect protocol, the first and second switching components in this application can ensure that the electrical signals received from the photoelectric converter conform to the standard. This includes functions such as signal dispersion compensation, error correction, and equalization, so that the signal can maintain high quality even after long-distance transmission, reduce the bit error rate and improve the reliability of data transmission. This not only improves the signal quality and transmission efficiency, but also enhances the system's compatibility, flexibility and stability.

[0052] In this embodiment, both the first photoelectric converter 110 and the second photoelectric converter 120 are connected to the substrate management controller. When the substrate management controller detects that the optical interconnect link meets the link establishment conditions, it establishes the optical interconnect link and controls the optical interconnect link to enter the training process.

[0053] It is understood that in this embodiment of the application, the substrate management controller is used to detect whether the optical interconnect link meets the establishment conditions and control it to enter the training process. This not only ensures that the optoelectronic conversion device can operate in the best condition, but also allows the linear direct-drive optoelectronic converter to adapt to different communication rates, thereby better meeting diverse needs.

[0054] Specifically, such as Figure 2 As shown, since the internal TIA and Driver of the optoelectronic converter need to undergo temperature compensation and bias current adjustment after power-on to ensure stable operation of the optoelectronic converter, the onboard management device needs to connect the switching component and the optoelectronic converter to control their timing relationship, ensuring that the switching component waits for the optoelectronic converter to operate stably before establishing the PCIe optical interconnect link. The onboard management device judges the stable operation of the optoelectronic converter based on: (1) monitoring whether the internal state machine of the optoelectronic converter has reached the valid state of the data link; (2) monitoring whether there is an alarm in the internal state register of the optoelectronic converter to avoid internal abnormal state affecting the transmission and reception of optical signals; (3) monitoring the optical power of the controlled end to ensure that the optoelectronic converters at both ends of the optical fiber are in a stable operating state.

[0055] According to the optical interconnect link component proposed in the embodiments of this application, when the switching components successfully receive the detection character sequence sent by the other party, the optical interconnect link is allowed to enter the training process to ensure normal communication of the optical interconnect link, thereby improving the success rate and reliability of link establishment. Only when the bidirectional communication link of the optical interconnect link is successfully established can the actual data transmission function be performed through the optical interconnect link. This not only solves the limitations of traditional copper cable interconnection, but also provides an efficient and stable data transmission link based on PCIe optical interconnection, realizing low-loss and long-distance transmission.

[0056] Embodiments of this application also provide a method for training an optical interconnect link.

[0057] Figure 3This is a flowchart of the training method for optical interconnect links provided in the embodiments of this application.

[0058] like Figure 3 As shown, the training method for this optical interconnect link is applied to the first switching component 130 in the optical interconnect link component 10, wherein the method includes:

[0059] In step S101, the first photoelectric converter is controlled to continuously send the first detection character sequence, identify the first electrical signal output by the first photoelectric converter device, and determine whether the second response character sequence is received within a preset time period. The second switching component outputs the second response character sequence after receiving the first detection character sequence.

[0060] The preset duration can be set according to actual needs without specific limitations.

[0061] It is understood that the embodiments of this application determine whether a second response character sequence is received within a preset time period by controlling the first photoelectric converter to continuously send the first detection character sequence and identifying the first electrical signal output by the first photoelectric converter device. This not only ensures the existence and normal operation of the devices at both ends of the link, but also significantly improves the success rate of link establishment and enhances the stability and reliability of the link.

[0062] In step S102, after determining that the second response character sequence is received within a preset time period, the internal state machine of the first switching component is controlled to be set to an effective state, wherein the effective state indicates that the optical interconnect link is allowed to enter the training process.

[0063] It is understood that, in this embodiment of the application, after receiving the second response character sequence within a preset time period, the internal state machine of the first exchange component is set to an effective state, ensuring that the training process is only allowed when the link conditions are met. This improves the reliability and efficiency of link establishment, enhances the stability and robustness of the system, and supports efficient link training.

[0064] It should be noted that the response character sequence is a series of specific characters generated by the controlled end after receiving the detection character sequence sent by the master end, and then sent back to the master end. These characters are used to verify the communication capability and status between the devices at both ends of the link. The main purpose is to ensure that the devices at both ends of the link can recognize and communicate with each other, and are ready to enter the subsequent link training phase. This is achieved in traditional electrical interconnects by detecting physical layer characteristics (such as DC common-mode impedance), but it cannot be directly applied in optical interconnects due to the presence of optoelectronic conversion devices. Therefore, a character sequence is used as an alternative.

[0065] In this embodiment of the application, before controlling the internal state machine of the first switching component to be set to an effective state, the method further includes: identifying the first continuous reception quantity of the second response character sequence within a preset time period; if the first continuous reception quantity is greater than or equal to the preset quantity, controlling the internal state machine of the first switching component to be set to an effective state; if the first continuous reception quantity is less than the preset quantity, controlling the internal state machine of the first switching component to be set to an invalid state, wherein the invalid state indicates that the optical interconnect link is not allowed to enter the training process.

[0066] The preset quantity can be set according to actual needs, without specific limitations.

[0067] It is understood that, before setting the internal state machine of the first switching component to an effective state, this application embodiment adds the identification of the first continuous reception quantity of the second response character sequence within a preset time period, and determines the state of the state machine based on the comparison result with the preset quantity. This not only enhances the robustness of the system in the face of abnormal situations, but also optimizes the efficiency of resource utilization and simplifies the troubleshooting process.

[0068] In this embodiment, the method further includes: identifying a first electrical signal output by a first photoelectric conversion device, determining a second detection character sequence output by a second switching component, and controlling the first photoelectric conversion device to send a first response character sequence.

[0069] It is understood that the embodiments of this application can identify the first electrical signal output by the first optoelectronic conversion device, determine the second detection character sequence output by the second switching component, and control the first optoelectronic conversion device to send the first response character sequence. This can determine that the switching component at the other end of the optical interconnect link exists and can receive data sent by the switching component at the other end through the optical interconnect link. This solves the problem of detection device failure in traditional electrical interconnects and provides a reliable foundation for subsequent link training, data transmission, and system management.

[0070] exist Figure 2 In the process shown, the switching component itself can be used as a USP device or a DSP device. That is, when both ends of the optical interconnect link are switching components, the switching component can skip the process of monitoring the USP device or DSP device and directly send device signals to the optoelectronic converter.

[0071] In this embodiment of the application, before controlling the first photoelectric converter to continuously send the first detection character sequence, the method further includes: determining whether to detect the first device signal of the master control terminal based on the type of the first switching component; if the type of the first switching component is a switching type, then controlling the first photoelectric converter to continuously send the first detection character sequence; if the type of the first switching component is a relay type, then after detecting the first device signal of the master control terminal, controlling the first photoelectric converter to continuously send the first detection character sequence.

[0072] It is understood that the embodiments of this application can adopt different link establishment strategies according to the switching type of the first switching component, thereby performing differentiated link initialization control based on different switching component types, realizing compatible design for multiple interconnection topologies, and improving the flexibility and adaptability of the system architecture.

[0073] According to the training method for optical interconnect links proposed in this application, the first switching component sends a detection character sequence, and the second switching component returns a response character sequence after receiving it. Only after the first switching component confirms the response is the training process allowed to begin. This ensures that both ends of the link are ready, avoids training failure due to one side not being ready, significantly improves the success rate and stability of link training, and constructs an efficient, reliable, and intelligent optical interconnect link establishment mechanism. It solves the core problem of detection device failure in PCIe optical interconnects and significantly improves the success rate of link training and the stability of the system.

[0074] Embodiments of this application also provide a method for training an optical interconnect link.

[0075] Figure 4 This is a flowchart of the training method for optical interconnect links provided in the embodiments of this application.

[0076] like Figure 4 As shown, the training method for the optical interconnect link is applied to the second switching component 140 of the optical interconnect link component 10, wherein the method includes:

[0077] In step S201, the second photoelectric converter is controlled to continuously send the second detection character sequence, identify the second electrical signal output by the second photoelectric converter device, and determine whether the first response character sequence is received within a preset time period. The first switching component outputs the first response character sequence after receiving the second detection character sequence.

[0078] The preset duration can be set according to actual needs without specific limitations.

[0079] It is understood that the embodiments of this application can determine whether the first response character sequence is received within a preset time by controlling the second photoelectric converter to continuously send the second detection character sequence and identifying the second electrical signal output by it. This not only ensures the existence and readiness of the link peer device, but also significantly improves the success rate of link establishment and the stability of the system.

[0080] In step S202, after determining that the first response character sequence is received within a preset time period, the internal state machine of the second switching component is controlled to be set to an effective state, wherein the effective state indicates that the optical interconnect link is allowed to enter the training process.

[0081] It is understood that, according to the embodiments of this application, after determining that the first response character sequence is received within a preset time period, the internal state machine of the second exchange component is set to an effective state. This design ensures that the training process is only allowed when the link conditions are met. This not only improves the reliability and efficiency of link establishment, but also enhances the stability and robustness of the system, supports efficient link training, and simplifies management and maintenance.

[0082] In this embodiment of the application, before controlling the internal state machine of the second switching component to be set to an effective state, the method further includes: identifying a second continuous reception quantity of the first response character sequence within a preset time period; if the second continuous reception quantity is greater than or equal to a preset quantity, controlling the internal state machine of the second switching component to be set to an effective state; if the second continuous reception quantity is less than a preset quantity, controlling the internal state machine of the second switching component to be set to an invalid state, wherein the invalid state indicates that the optical interconnect link is not allowed to enter the training process.

[0083] The preset quantity can be set according to actual needs, without specific limitations.

[0084] It is understood that, in the embodiments of this application, a mechanism for statistical analysis and judgment of the number of first response character sequences received can be added before the internal state machine of the second switching component is set to an effective state. This not only enhances the system's ability to resist interference and occasional failures, but also improves resource utilization efficiency.

[0085] It's important to note that a single successful reception of a response character sequence is insufficient to determine link stability and reliability. By setting a threshold, the system ensures that both ends of the link are considered to be in normal communication only after a continuous and stable reception of response signals. This approach avoids misjudgments caused by momentary interference, noise, or occasional errors, improving the reliability of link establishment. In complex electromagnetic environments or long-distance transmission, data loss or bit errors may occur. By identifying the number of consecutively received response character sequences, the system can distinguish between occasional packet loss and link anomalies. If the number of received sequences is insufficient, the link quality is considered substandard, preventing it from entering the training process and thus enhancing the system's anti-interference capability.

[0086] In this embodiment, the method further includes: identifying a second electrical signal output by the second photoelectric conversion device, determining a first detection character sequence output by the first switching component, and controlling the second photoelectric conversion device to send a second response character sequence.

[0087] It is understood that the embodiments of this application can identify the second electrical signal output by the second photoelectric conversion device, determine the first detection character sequence output by the first switching component, and control the second photoelectric conversion device to send the second response character sequence. Only when the first photoelectric converter successfully identifies the first detection character sequence from the other end and can correctly return the first response character sequence can it be said that the link has the basic conditions to enter the training phase. This avoids blindly entering the training process when the link is abnormal or the equipment is not ready, and significantly improves the success rate and stability of link training.

[0088] In this embodiment of the application, before controlling the second photoelectric converter to continuously send the second detection character sequence, the method further includes: determining whether to detect the second device signal of the controlled end based on the type of the second switching component; if the type of the second switching component is a switching type, then controlling the second photoelectric converter to continuously send the second detection character sequence; if the type of the second switching component is a relay type, then after detecting the second device signal of the master control end, controlling the second photoelectric converter to continuously send the second detection character sequence.

[0089] It is understood that the embodiments of this application can adopt different link establishment strategies according to the switching type of the second switching component, thereby performing differentiated link initialization control based on different switching component types, realizing compatible design for multiple interconnection topologies, and improving the flexibility and adaptability of the system architecture.

[0090] The training method for optical interconnect links proposed in this application constructs an optical interconnect link establishment mechanism by detecting controlled end device signals, sending detection character sequences, recognizing response character sequences, and controlling the state machine of the optoelectronic converter. This not only solves the problem of failure of traditional physical layer detection mechanisms in optical interconnects, but also significantly improves the reliability, stability, and intelligence level of link establishment. It is an important technical support for realizing high-performance, low-latency, large-scale GPU-accelerated cluster systems.

[0091] The following will combine Figure 2 The training method for the optical interconnect link in this application is described in detail below:

[0092] When the switching components are working normally and the PCIe optical interconnect link is not established, the system actively monitors whether a USP (User Services Platform) device or a DSP device is connected to the switching components via a PCIe electrical interconnect link. When both ends detect USP and DSP devices, the switching components at both ends continuously and actively send specific detection character sequences to the opto-converter to monitor whether the PCIe optical interconnect link can communicate normally. If the controlled end of the switching component at one end of the PCIe optical interconnect link detects N consecutive specific detection character sequences sent by the opto-converter within a specific time, it is considered that data sent by the other end can be received through this PCIe optical interconnect link. If the specified number of detection character sequences are not received after the timeout, the PCIe optical interconnect link is considered abnormal, character sequence sending stops, the PCIe optical interconnect link training RX detection (detection device, receiver detection) process exits, and the RX is restarted only after the onboard management device sends a command. The detection process involves the following steps: First, upon receiving a continuous sequence of specific detection characters, the main control unit of the switching component actively sends a specific response character sequence to the connected opto-converter, informing the peer switching component that it can receive data transmitted via the optical interconnect link. Second, while the switching component sends the specific response character sequence at its main control unit, the controlled end continuously monitors whether it receives the specific response character sequence from the opto-converter. If M consecutive specific response character sequences are received within a specified time, it is considered that the switching components at both ends of the link exist and are functioning normally, and can transmit data normally via the PCIe optical interconnect link. If the specified number and format of characters are not received within the specified time, the process is considered complete. If the sequence is not found, it is considered that the photoelectric converter / peer-end switching component is in an abnormal state and cannot establish a PCIe optical interconnect link normally. In this case, it will return to the stage of sending a specific detection character sequence. In order to avoid the power timing of the switching components at both ends of the link being out of sync, resulting in the chips at both ends of the link receiving the detection character sequence out of sync, if one end of the switching component has not yet received the required N specific detection character sequences, that is, when the specific response sequence sent by the peer end is detected, it will directly enter the stage of judging whether the specific response sequence meets the requirements, and then judge whether the switching components at both ends of the link exist and work normally, and whether data can be sent normally through the PCIe optical interconnect link.

[0093] If it has been confirmed that the switching components at both ends of the PCIe optical interconnect link exist and are working properly, and data can be sent normally through the PCIe optical interconnect link, when the USP device or DSP device actively sends a device signal, the switching components directly inform the USP device or DSP device that the receiving device exists, and the next link training polling process can be carried out directly.

[0094] The embodiments of this application provide an accelerated cluster system, and the system is described in detail in conjunction with the operation process of the accelerated cluster system.

[0095] Specifically, Figure 5 This is a block diagram of the accelerated cluster system provided in the embodiments of this application.

[0096] like Figure 5 As shown, the accelerated cluster system 20 is based on the optical interconnect link component 10, wherein the system includes: an optical interconnect computing unit 210, an optical interconnect acceleration unit 220 and an optical interconnect switching unit 230.

[0097] The optical interconnect computing unit 210, the optical interconnect acceleration unit 220, and the optical interconnect switching unit 230 include at least one, wherein the optical interconnect computing unit 210 provides computing resources, the optical interconnect acceleration unit 220 provides acceleration resources, at least one of the optical interconnect switching unit 230 and the optical interconnect computing unit 210 and the optical interconnect acceleration unit 220 is connected via an optical interconnect link, and the optical interconnect acceleration unit 220 is connected to at least one optical interconnect computing unit 210 via an optical interconnect link.

[0098] It is understood that in this embodiment, the optical interconnect computing unit focuses on CPU (Central Processing Unit) resources, the optical interconnect acceleration unit provides GPU resources, and the optical interconnect switching unit independently undertakes data routing. Based on the characteristics of PCIe optical signals in optical fibers, such as low attenuation, anti-interference, and long-distance transmission, the three can be deployed across racks, breaking through the physical limitations of a single rack and supporting the on-demand distributed layout of computing and acceleration resources. Data is transmitted through optical interconnect links within the rack, across racks, and between switching boards and resources. The number of resource pools can be flexibly adjusted according to the application scenario to realize the on-demand allocation of computing, switching, and network resources, which can meet diverse deployment needs. The optical interconnect acceleration units are directly connected by optical links to achieve microsecond-level latency communication between GPUs, avoiding network congestion and realizing long-distance, low-loss, low-latency, and high-speed PCIe data transmission.

[0099] like Figure 6As shown, a single optical interconnect switching unit connects to one GPU acceleration rack, and four GPU acceleration racks connect to a Layer 4 optical interconnect switching unit within a PCIe optical interconnect switching unit in a PCIe switch rack. The GPU acceleration racks and PCIe switch racks are decoupled and can be independently distributed in different locations within different data centers. The racks are connected via photoelectric converters and optical fibers to achieve data transmission of PCIe optical signals between racks. Each single-layer optical interconnect switching unit supports 10 PCIe x16 optical interconnect links externally, with 8 PCIe optical interconnect links connecting to GPU acceleration racks and 2 PCIe x16 optical interconnect links connecting to the top-level PCIe optical interconnect switching unit to achieve larger-scale GPU acceleration resource expansion. Internally, the 4 PCIe x16 interconnect ports of the optical interconnect switching unit can be used for interconnecting switching components within the switching unit. Based on different cable connections, 1x2, 2x2, 2x3, and 2x4 optical interconnect switching unit interconnection topologies can be implemented, allowing GPU accelerator cards in different GPU acceleration racks to access GPU accelerator cards in other GPU acceleration racks through these interconnect ports. When the two sets of PCIe x16 connected to the top-level PCIe optical interconnect switching unit are changed to connect to the optical interconnect computing unit, a single optical interconnect computing unit can support simultaneous communication with 128 GPU accelerator cards.

[0100] In this embodiment of the application, if there are multiple optical interconnect switching units 230, the multiple optical interconnect switching units 230 include at least one main switching unit and at least one secondary switching unit. The main switching unit and the at least one secondary switching unit are connected through an optical interconnect link, and the main switching units are connected to each other through an optical interconnect link.

[0101] It is understood that the embodiments of this application can hierarchically distribute the communication traffic of a large-scale GPU cluster through the hierarchical interconnection of the main switching unit and the secondary switching unit, so as to avoid a single switching unit becoming a performance bottleneck. The main switching units are directly connected to each other, and the secondary switching units are mounted under the main units, so as to ensure that GPUs in any acceleration rack can access the resources of other racks through optical links, and realize the full connectivity of cross-unit resource pools. When the secondary switching unit fails, the main switching unit can dynamically switch the routing path. The multi-link interconnection between the main units provides redundancy, avoids single point failure leading to cluster paralysis, and enhances the fault tolerance of the system.

[0102] Specifically, such as Figure 7As shown, in a 1000-GPU acceleration cluster, a single PCIe optical interconnect switching unit connects 8*PCIe x16 optical interconnect links to the optical interconnect switching unit within the top-level PCIe optical interconnect switching unit. A single top-level PCIe optical interconnect switching unit can support connections to 4 PCIe optical interconnect switching units, which are respectively connected to the 4 layers of optical interconnect switching units within the top-level PCIe optical interconnect switching unit. Each layer of optical interconnect switching unit within the top-level PCIe optical interconnect switching unit, in addition to aggregating the 8*PCIe x16 optical interconnect links from the PCIe optical interconnect switching unit, has two additional sets of PCIe x16 links used for interconnection with the optical interconnect switching units within the top-level PCIe switching unit of another PCIe switch rack. A single top-level PCIe optical interconnect switching unit is responsible for cross-rack PCIe data exchange for 512 GPU accelerator cards within the system, while two top-level PCIe optical interconnect switching units jointly handle the data exchange tasks for 1024 GPU accelerator cards within the system.

[0103] In the embodiments of this application, at least one of the main switching unit and the secondary switching unit is connected to at least one optical interconnect computing unit via an optical interconnect link.

[0104] It is understood that in this embodiment of the application, the main switching unit and the secondary switching unit are connected to the optical interconnect computing unit through optical interconnect links, which can significantly improve data transmission efficiency and reduce latency. By optimizing the connection method between the optical interconnect computing unit and the switching unit, the computing resources in the cluster can be utilized more effectively, thereby improving the overall system operating efficiency. The use of optical interconnect technology reduces data communication latency compared to network interconnection schemes, and high-speed, low-latency data transmission can be achieved even in cross-rack situations. This not only enhances the system's data processing capabilities and response speed, but also provides greater flexibility and reliability.

[0105] Specifically, when the interconnect port between the two top-level PCIe optical interconnect switching units is changed to connect to the optical interconnect computing unit, each optical interconnect computing unit can access 512 GPU accelerator cards distributed in different locations through the top-level PCIe optical interconnect switching unit. The 512 GPU accelerator cards can also perform PCIe data communication through the PCIe optical interconnect switching unit and the top-level PCIe optical interconnect switching unit.

[0106] In this embodiment of the application, the optical interconnect switching unit 230 includes at least one optical interconnect switching board, wherein when the optical interconnect switching board connects to computing resources based on optical interconnect links, the optical interconnect switching board connects to at least one optical interconnect computing unit 210.

[0107] It is understood that the embodiments of this application can uniformly aggregate and manage PCIe data streams from multiple optical interconnect computing units through an optical interconnect switching board, avoiding topology chaos caused by scattered connections of computing resources. When the optical interconnect switching board is connected to an optical interconnect computing unit, direct data exchange is achieved through the PCIe optical interconnect link, which significantly improves communication efficiency and data processing capabilities.

[0108] In this embodiment of the application, when the optical interconnect switching board is based on optical interconnect link connection acceleration resources, the optical interconnect switching board is connected to at least one optical interconnect acceleration unit 220.

[0109] It is understood that in the embodiments of this application, the optical interconnect switching board connects at least one optical interconnect acceleration unit through an optical interconnect link, which not only solves the distance limitation problem of traditional electrical interconnect, but also significantly improves the scalability, flexibility and overall performance of the system, enabling large-scale GPU acceleration clusters to operate more efficiently in data center environments.

[0110] Specifically, such as Figure 8 As shown, a single GPU acceleration rack contains four optical interconnect acceleration units and at least one optical interconnect computing unit, housing 32 GPU accelerator cards. Each optical interconnect computing unit can communicate with all 32 GPU accelerator cards. Inside the rack, each optical interconnect computing unit outputs eight PCIe x16 ports, connecting to eight switching components within the four optical interconnect acceleration units, forming eight PCIe optical interconnect links between the optical interconnect computing units and the optical interconnect acceleration units. Each optical interconnect acceleration unit configures its two externally connected PCIe x16 ports as four PCIe x8 interconnect ports for interconnecting switching components between acceleration resource pools within the rack, enabling high-speed, low-latency communication between the 32 GPU accelerator cards in the same rack. In addition to the ports used for internal interconnection, each switching component outputs one PCIe x16 optical signal for cross-rack connections to PCIe optical interconnect switching units within a PCIe switch rack. A single GPU acceleration rack outputs eight PCIe x16 optical signals to connect to the same layer of PCIe optical interconnect switching boards within a PCIe optical interconnect switching unit. GPU accelerator cards distributed across different optical interconnect acceleration units can communicate not only through the interconnect ports between resource pools within the rack, but also via data relay components on the PCIe optical interconnect switching board. Similar to the 8-card system mentioned above, this rack can also be used as a standalone 32-GPU accelerator card system. In this case, one PCIe x16 connector connected to the PCIe optical interconnect switching board can be used for internal rack interconnection, optimizing the GPU interconnect topology.

[0111] In this embodiment of the application, the optical interconnect computing unit 210 includes at least one central processing unit and at least one switching component.

[0112] The central processing unit is connected to at least one switching component, and the switching component is connected to at least one photoelectric converter. The photoelectric converter inside the computing unit and the photoelectric converter outside the computing unit form an optical interconnect link, and the central processing unit provides computing resources through the optical interconnect link.

[0113] It is understood that the central processing unit in this application embodiment connects to external devices through optical interconnect links, enabling local computing resources to communicate efficiently with other nodes across physical distances, supporting low-latency, high-bandwidth data interaction between multiple nodes, improving the overall system's computing power and scalability, thereby breaking through the physical limitations of traditional electrical interconnects, and also realizing cross-node sharing and flexible expansion of computing resources, with high stability, low power consumption, easy maintenance, and strong scalability.

[0114] It should be noted that the photoelectric converter is responsible for converting between electrical signals and optical signals. Specifically, at the main control end, the photoelectric converter converts electrical signals into optical signals for transmission through optical fibers; at the controlled end, the photoelectric converter then converts the received optical signals back into electrical signals for subsequent equipment processing.

[0115] In this embodiment, the optical interconnect computing unit 210 further includes at least one substrate management controller, which is connected to at least one switching component and at least one optoelectronic converter, and is used to manage the timing relationship between the switching component and the optoelectronic converter and the establishment of the optical interconnect link.

[0116] It is understood that, in the embodiments of this application, the baseboard management controller can precisely control the startup sequence of the switching components and the optoelectronic converters, as well as the working timing between the switching components and the optoelectronic converters, ensuring that the transimpedance amplifiers and drivers of the optoelectronic conversion devices can perform PCIe link training after temperature compensation and bias current adjustment are completed, thereby ensuring the stable establishment of the link; reducing link failures caused by timing mismatches and improving system stability.

[0117] In this embodiment, the substrate management controller detects the operating parameters of the photoelectric converter and establishes an optical interconnect link when the photoelectric converter is in a stable state based on the operating parameters.

[0118] It is understood that in this embodiment, the substrate management controller detects the operating parameters of the optoelectronic converter. Based on the operating parameters, it determines that the optoelectronic converter is in a stable state and establishes an optical interconnect link. Only when all key parameters meet the expected standards will the substrate management controller trigger the link establishment process, thereby effectively avoiding link instability or failure caused by the device not being fully ready. If some parameters are detected to deviate from the normal range, corresponding correction measures can be automatically executed, such as adjusting the bias current or notifying the administrator for maintenance, to ensure that the optoelectronic converter is always in the best working state.

[0119] It's important to note that this scenario assumes a large-scale GPU-accelerated cluster where multiple optical interconnect computing units need to be interconnected to form a high-efficiency data processing network. Each computing unit contains at least one photoelectric converter for photoelectric signal conversion. Without a substrate management controller to monitor the photoelectric converter's operational status, frequent link interruptions or data loss may occur during the initial link establishment phase because the photoelectric converters have not yet reached a stable operating state. However, with the intervention of the substrate management controller, it will only allow link establishment after all key parameters of the photoelectric converter (such as temperature, bias current, and optical power) meet the requirements, thus ensuring the reliability and stability of data transmission.

[0120] In this embodiment, the operating parameters include at least one of the state machine state, register state, and controlled-end optical power within the optoelectronic conversion device. When the substrate management controller determines that the operating parameters meet the target conditions, it determines that the optoelectronic conversion device is in a stable state. The target conditions include: the state machine state is valid, wherein the substrate management controller identifies the valid flag bit of the state machine and determines that the state machine state is valid; the register state is normal, wherein the substrate management controller identifies the normal flag bit of the register and determines that the register state is normal; and the controlled-end optical power is greater than a preset optical power.

[0121] The preset optical power can be set according to actual needs without specific limitations.

[0122] It is understood that the embodiments of this application can identify the valid flag bit of the state machine to confirm whether the state machine is in a valid state; identify the normal flag bit in the register to confirm whether the register is in a normal state without alarms or abnormalities; ensure that the optical power of the controlled end is greater than a preset value to ensure the quality of signal transmission. Only when all target conditions are met will the substrate management controller determine that the optoelectronic conversion device is in a stable state and allow the establishment of an optical interconnect link. By strictly checking the key operating parameters of the optoelectronic conversion device, the substrate management controller can prevent the link from being established if the device is not fully ready, thereby avoiding potential link instability or failure problems. This not only improves the reliability and stability of the system, but also enhances the flexibility, adaptability and management convenience of the system.

[0123] It is important to note that during the photoelectric conversion process, any unstable factors can lead to signal distortion, affecting the final data transmission quality. By ensuring that the photoelectric conversion device is in a stable state before establishing the link, signal distortion can be minimized, guaranteeing high-speed, low-latency data transmission. If the substrate management controller detects that certain parameters deviate from the normal range (e.g., excessively high temperature, insufficient optical power, or register alarms), it can trigger corresponding corrective measures, such as adjusting the bias current or notifying the administrator for maintenance, to ensure that the equipment is always in optimal operating condition. The substrate management controller can dynamically adjust the operating parameters of the photoelectric conversion device according to changes in the actual operating environment (e.g., temperature fluctuations, load changes), enabling it to maintain optimal performance under different conditions.

[0124] Specifically, such as Figure 7 As shown, the optical interconnect computing resource pool contains two CPU processors. Each CPU processor outputs multiple sets of PCIe x16 electrical signals. Each set of PCIe x16 connects to a switching component, and each switching component outputs two sets of PCIe x8 signals, which are connected to two opto-converter connectors. Simultaneously, an onboard management device connects to all switching components and opto-converter devices within the computing unit, controlling the initialization and link establishment process of the same PCIe optical interconnect link. The switching component optimizes the PCIe signals emitted by the opto-converters and assists in the PCIe optical interconnect link training process. Two opto-converters connected to the same switching component can be used as 1*PCIe x16 or 2*PCIe x8, supporting more connection topologies in the system. Based on the above, the optical interconnect computing unit can expand externally with multiple sets of PCIe optical interconnect links, which can be connected to optical interconnect switching units or optical interconnect acceleration resources, realizing the expansion of computing resources in the form of PCIe optical signals.

[0125] In this embodiment of the application, the optical interconnect switching unit 230 includes at least one optical interconnect switching board.

[0126] The optical interconnect switching board includes a substrate connector and at least one switching component, wherein the substrate connector is connected to at least one switching component; the switching component is connected to at least one high-speed connector and at least one optoelectronic converter, wherein the optical interconnect switching board communicates with other optical interconnect switching boards through the high-speed connector, and the optoelectronic converters of the optical interconnect switching board form an optical interconnect link with other optoelectronic converters.

[0127] It is understood that in this embodiment, the optical interconnect switching boards communicate with each other through high-speed connectors, supporting high-speed data exchange within and between boards, improving the data transmission efficiency between modules within the switching unit, reducing communication latency, and supporting the rapid scheduling requirements of computing resources in large-scale GPU clusters. Each switching component connects to at least one optoelectronic converter to establish an optical interconnect link between local and remote devices. The optoelectronic converters inside the optical interconnect switching board and the optoelectronic converters of other external units form an optical interconnect link, enabling high-speed data transmission between different racks via fiber optics. This breaks the distance limitations of traditional copper cable PCIe connections, supports flexible deployment across racks and even floors, provides lower signal attenuation and higher anti-interference capabilities, ensures stable communication over long distances, and achieves efficient resource collaboration in large-scale GPU acceleration clusters.

[0128] Specifically, the optical interconnect switching unit mainly consists of four optical interconnect switching units, and the connection topology of a single optical interconnect switching unit is as follows: Figure 10 As shown. Each optical interconnect switching unit contains two switching components that support optical interconnect links. Each switching component supports connecting nine sets of PCIe Gen5 x16 ports. Four of these PCIe x16 ports are internally connected to MCIO x16 connectors for interconnection within the switching unit, allowing them to exchange data transmitted from the external ports of the switching components. The other five PCIe x16 ports are externally connected to ten optoelectronic converter connectors.

[0129] In this embodiment of the application, the switching component includes multiple internal ports and multiple external ports.

[0130] The internal port is connected to a high-speed connector, and the external port is connected to a photoelectric converter.

[0131] It is understood that in this embodiment, the internal port is connected to the high-speed connector and the external port is connected to the optoelectronic converter, supporting multi-level network topology, meeting the high-efficiency communication requirements between large-scale GPU cards, and enabling seamless interconnection between local and remote resources. This not only improves the system's bandwidth utilization and communication efficiency, but also enhances its scalability, stability and manageability.

[0132] It should be noted that the internal ports connect to high-speed connectors for high-speed data forwarding within the optical interconnect switching unit or between different switching boards; the external ports connect to optoelectronic converters for establishing optical interconnect links with external devices such as computing units and acceleration units. The ratio of internal to external ports can be flexibly configured according to application needs to adapt to cluster deployments of different sizes.

[0133] In this embodiment of the application, the configuration type of the external port includes any one of an uplink port, a downlink port, and an interconnect port, wherein the uplink port connects to computing resources, the downlink port connects to acceleration resources, and the interconnect port connects to other switching components.

[0134] It is understood that each external port in this application embodiment can be configured as an uplink, downlink or interconnect port according to actual deployment needs, adapting to different application scenarios. Through flexible configuration, the system can build a variety of network topologies to meet the different needs of communication paths in heterogeneous computing environments, which not only improves the communication efficiency and resource utilization of the system, but also enhances its scalability, reliability and manageability.

[0135] It should be noted that the uplink port directly connects to computing resources: ensuring that the CPU can access remote GPU acceleration resources with the shortest path and reducing intermediate forwarding layers; the downlink port directly connects to acceleration resources: avoiding bandwidth bottlenecks caused by shared buses in traditional switching component architectures and improving communication efficiency between GPUs; the interconnect port builds a high-speed channel: establishing dedicated links between multiple switching components for internal load balancing or cross-processor data forwarding.

[0136] Specifically, such as Figure 11 As shown, each external optoelectronic converter connector supports PCIe x8. Depending on the application scenario, the switching component can change the external PCIe x16 port to two independent PCIe x8 ports. A single PCIe switching chip can support up to 10 external PCIe Gen5 x8 ports. Each x8 or x16 port can be flexibly configured as an uplink port (connecting computing resources), a downlink port (connecting acceleration resources), or an interconnect port (interconnecting switching components). If all external ports of a switching component are configured as PCIe x16, with one as an uplink port and the rest as downlink ports, one set of PCIe x16 optical interconnect links can be expanded to four sets of PCIe x16 optical interconnect links. When all external ports are configured as PCIe x8, the expansion performance can be improved proportionally. Based on differences in application scenarios, the internal port connection topology of the optical interconnect switching unit can be flexibly changed, the external port configuration type can be modified, the allocation ratio of uplink ports, downlink ports, and interconnect ports in the external ports can be adjusted, and the data routing method between switching components can be optimized. It can support multiple interconnect topologies of PCIe optical interconnect systems and meet the PCIe interconnect resource requirements of large-scale clusters of GPU accelerator cards.

[0137] In this embodiment, the optical interconnect switching unit 230 further includes at least one clock buffer.

[0138] The clock buffer is connected to the clock port of the switching component and is used to amplify and distribute the clock signal of the switching component.

[0139] It is understood that, in the embodiments of this application, the clock buffer is used to enhance and distribute the clock signal of the switching component, which can enhance the amplitude of the original clock signal and improve the clock signal and stability integrity.

[0140] It's important to note that this scenario assumes a large-scale GPU-accelerated cluster where multiple switching components need to work collaboratively to achieve efficient distributed computing. Each processor relies on a precise clock signal for data synchronization and processing scheduling. Without a clock buffer, as the system scales and signal transmission distances increase, the clock signal may attenuate or become distorted, leading to synchronization failures and data transmission errors. Signal attenuation is a particularly problematic issue in high-frequency applications. Clock buffers compensate for attenuation during transmission by amplifying the signal, thereby ensuring signal quality.

[0141] In this embodiment, the optical interconnect acceleration unit 220 includes at least one graphics processor and at least one switching component.

[0142] The switching component is connected to at least one graphics processor, and the switching component is connected to at least one optoelectronic converter. The optoelectronic converter of the optical interconnect acceleration unit forms an optical interconnect link with other optoelectronic converters. At least one graphics processor provides acceleration resources through the optical interconnect link.

[0143] It is understood that, in this embodiment of the application, by introducing photoelectric converters and switching components, the GPU can establish a connection with remote computing resources through optical fiber, breaking through the distance limitations of traditional PCIe electrical connections; the photoelectric converters of multiple optical interconnect acceleration units form optical interconnect links with other photoelectric converters, forming a large-scale GPU cluster, which can flexibly allocate the number of GPUs according to the business load requirements, meet the elastic resource scheduling requirements in different scenarios, and improve the overall system utilization; the switching component is responsible for converting the PCIe protocol signals emitted by the GPU into a format suitable for transmission on the optical link, and restoring them to the original PCIe signals at the controlled end, reducing communication latency and improving heterogeneous computing efficiency.

[0144] It should be noted that the optical interconnect acceleration units in this application are all optical interconnect acceleration resource pools. Acceleration cards such as GPUs / NPUs (Graphics Processing Units / Neural Processing Units), which were originally bound to a single server, are centrally deployed through optical interconnect links. This decouples physical resources and constructs a virtualized acceleration resource pool. By combining high-speed optical interconnect technology with intelligent protocol processors, centralized management, on-demand scheduling, and remote transparent access to heterogeneous computing resources are achieved. It not only breaks the physical limitations of the traditional PCIe architecture but also improves resource utilization, system flexibility, and operational efficiency, serving as the core infrastructure for building next-generation high-performance, low-latency distributed heterogeneous computing systems.

[0145] In this embodiment, the switching component includes multiple ports, which are configured as internal ports and external ports. The internal ports are connected to the graphics processor, and the external ports are connected to the optoelectronic converter. The configuration type of the external ports includes interconnect ports, which are connected to other switching components inside the optical interconnect acceleration unit.

[0146] It is understood that the embodiments of this application can meet the communication path requirements of different application scenarios by reasonably configuring the ports of the switching components, thereby achieving efficient scheduling and communication management of GPU acceleration resources and improving the communication efficiency and resource utilization of the system.

[0147] It should be noted that the internal ports connect to the GPU: ensuring that the GPU can directly access the switching components through the PCIe interface for local data processing; the external ports connect to the optoelectronic converter: enabling high-speed optical link communication with external computing resources and switching devices; the interconnect ports construct on-chip / on-board interconnect channels: allowing multiple switching components within the same acceleration unit to establish efficient data paths, enabling seamless data flow from a single GPU to multiple GPUs, and from local to remote.

[0148] Specifically, the PCIe optical interconnect acceleration unit connection topology is as follows: Figure 12As shown, the resource pool consists of two switching components that support optical interconnect links and eight GPU accelerator cards. Each switching component connects to eight optical-to-electrical converter connectors with four PCIe x16 ports. Internally, the four PCIe x16 ports are connected to four GPU accelerator cards. One set of PCIe x16 ports is configured as an interconnect port for interconnecting the two internal switching components. The external opto-converter of the PCIe optical interconnect acceleration unit can be directly connected to the optical interconnect computing unit to achieve a one-to-one correspondence between the CPU processor and the GPU accelerator card, ensuring high-speed, low-latency PCIe data interconnection. Alternatively, it can be connected to the PCIe optical interconnect switching unit, and then connected to the optical interconnect computing unit through the PCIe optical interconnect switching unit. Based on the PCIe resource expansion and internal PCIe data exchange functions of the PCIe optical interconnect switching unit, it can support the expansion of the number of GPU accelerator cards that a single CPU processor can access. GPU accelerator cards within the cluster can also achieve high-speed communication through the PCIe optical interconnect switching unit. Alternatively, depending on application requirements, some ports can be configured as uplink ports to connect to the optical interconnect computing unit, while the remaining ports are configured as interconnect ports for interconnection between optical interconnect acceleration units. In this case, a single CPU processor can communicate with more GPU accelerator cards, and GPU accelerator cards can also transmit PCIe data through the interconnect ports between the switching components, which can improve the communication bandwidth of GPU accelerator cards across resource pools.

[0149] The accelerated cluster system proposed in this application replaces traditional copper cables with optical interconnect links. Utilizing the low attenuation and electromagnetic interference resistance characteristics of optical fibers, it supports resource pool interconnection across racks, floors, and even regions. Optical interconnect computing units, optical interconnect acceleration units, and optical interconnect switching units can be dynamically combined as needed, and the number of resource pools can be freely increased or decreased according to the scenario. Low-power, long-distance, and low-latency PCIe data transmission is supported between resource pools within / between racks, enabling flexible expansion of the GPU accelerated cluster scale based on PCIe communication.

[0150] The following will combine Figure 12 The accelerated cluster system of this application is described in detail below:

[0151] This system leverages the long-distance transmission capability of PCIe optical interconnects to distribute a 1000-card GPU acceleration cluster across 34 racks. 32 racks are dedicated to GPU accelerator cards, and 2 are PCIe switch racks. Each GPU accelerator card rack contains 4 optical interconnect acceleration units, resulting in a total of 128 optical interconnect acceleration units across the entire GPU acceleration cluster, supporting 128 * 8 = 1024 GPU accelerator cards. The racks are connected by fiber optic cables to transmit PCIe optical signals, enabling long-distance, distributed deployment of a 1000-card GPU acceleration cluster based on PCIe interconnects.

[0152] Each resource pool in this application contains a different number of switching components and connector interfaces for connecting optoelectronic converters. The switching components and connector interfaces have pre-installed onboard PCIe links. When resource pools are connected via optoelectronic converters and optical fibers, multiple PCIe optical interconnect links as mentioned above can be formed. Based on the characteristics of PCIe optical signals—low attenuation, interference resistance, and long-distance transmission in optical fibers—it can support distributed deployment of optical interconnect computing resource pools, optical interconnect switching units, and optical interconnect acceleration units across racks, floors, and regions, enabling large-scale expansion of GPU accelerator card clusters based on the PCIe protocol.

[0153] The embodiments of this application also provide an optical interconnect computing unit, and the system is described in detail in conjunction with the operation process of the optical interconnect computing unit.

[0154] Specifically, Figure 13 This is a block diagram of the optical interconnect computing unit provided in an embodiment of this application.

[0155] like Figure 13 As shown, the optical interconnect computing unit 210 is based on the optical interconnect link component 10 and includes: at least one central processing unit and at least one switching component.

[0156] The central processing unit is connected to at least one switching component, and the switching component is connected to at least one photoelectric converter. The photoelectric converter inside the computing unit and the photoelectric converter outside the computing unit form an optical interconnect link.

[0157] It should be noted that the foregoing explanation of the accelerated cluster system embodiment also applies to the optical interconnect computing unit of this embodiment, and will not be repeated here.

[0158] According to the optical interconnect computing unit proposed in the embodiments of this application, the photoelectric converter and the photoelectric converters of other external units (such as optical interconnect acceleration units, optical interconnect switching units, etc.) form an optical interconnect link through optical fiber, thereby realizing cross-unit data communication and meeting the high bandwidth requirements of large-scale GPU clusters.

[0159] The embodiments of this application also provide an optical interconnect switching unit, and the system is described in detail in conjunction with the operation process of the optical interconnect switching unit.

[0160] Specifically, Figure 14 This is a block diagram of an optical interconnect switching unit provided in an embodiment of this application.

[0161] like Figure 14 As shown, the optical interconnect switching unit 230 is based on the optical interconnect link component 10 and includes at least one optical interconnect switching board.

[0162] The optical interconnect switching board includes: a substrate connector and at least one switching component, wherein the substrate connector is connected to at least one switching component; the switching component is connected to at least one high-speed connector and at least one optoelectronic converter, wherein the optical interconnect switching board communicates with other optical interconnect switching boards through the high-speed connector, and the optoelectronic converter inside the optical interconnect switching board and the optoelectronic converter outside the optical interconnect switching board form an optical interconnect link.

[0163] It should be noted that the foregoing explanation of the accelerated cluster system embodiment also applies to the optical interconnect switching unit of this embodiment, and will not be repeated here.

[0164] According to the optical interconnect switching unit proposed in the embodiments of this application, data transmission between optical interconnect switching boards is carried out through high-speed connectors to ensure high-speed signal integrity. The switching component is connected to a processor responsible for routing and forwarding, and performs intelligent routing of data packets between multiple ports to reduce cross-device communication latency. Each switching component is connected to at least one optoelectronic converter to establish an optical interconnect link between local and remote devices. The optoelectronic converters inside the optical interconnect switching board and the optoelectronic converters of other external units form an optical interconnect link, enabling data to be transmitted at high speed in fiber optic mode between different cabinets. This breaks the distance limitation of traditional copper cable PCIe connections, supports flexible deployment across cabinets and even across floors, provides lower signal attenuation and higher anti-interference capability, ensures stable communication over long distances, and realizes efficient resource collaboration in large-scale GPU acceleration clusters.

[0165] The embodiments of this application also provide an optical interconnect acceleration unit, and the system is described in detail in conjunction with the operation process of the optical interconnect acceleration unit.

[0166] Specifically, Figure 15 This is a block diagram of the optical interconnect acceleration unit provided in an embodiment of this application.

[0167] like Figure 15 As shown, the optical interconnect acceleration unit 220 is based on the optical interconnect link component 10 and includes: at least one graphics processor and at least one switching component.

[0168] The switching component is connected to at least one graphics processor and at least one optoelectronic converter. The optoelectronic converter inside the optical interconnect acceleration unit and the optoelectronic converter outside the optical interconnect acceleration unit form an optical interconnect link. At least one graphics processor provides acceleration resources through the optical interconnect link.

[0169] Specifically, such as Figure 16As shown, each optical interconnect acceleration resource pool contains 8 GPUs. Each switching component internally connects 4 PCIe x16 ports to the 4 GPUs, 1 PCIe x16 port is used for interconnection within the resource pool's switching components, and externally, 1 PCIe x16 port is used to connect compute nodes, 1 PCIe x16 port is used to connect PCIe optical interconnect switching units, and 4 PCIe x8 ports are configured as interconnect ports for interconnection within the GPU Chassis's internal optical interconnect acceleration resource pool. Based on the design method described above, the external connection of the optical interconnect acceleration resource pool uses photoelectric converters to convert PCIe electrical signals into optical signals for transmission in optical fibers. Each photoelectric converter can transmit PCIe x8 signals. Depending on the application scenario, the port allocation ratio for connecting PCIe optical interconnect switching units and interconnect ports can be adjusted to flexibly adapt to application requirements. Each GPU acceleration resource pool can be independently connected to a single optical interconnect compute unit, allowing one compute unit to communicate with 8 GPU accelerator cards.

[0170] It should be noted that the foregoing explanation of the accelerated cluster system embodiment also applies to the optical interconnect acceleration unit of this embodiment, and will not be repeated here.

[0171] According to the optical interconnect acceleration unit proposed in the embodiments of this application, by introducing a photoelectric converter and a switching component, the GPU can establish a connection with remote computing resources through optical fiber, breaking through the distance limitation of traditional PCIe electrical connection; multiple optical interconnect acceleration units can be interconnected through optical interconnect links to form a large-scale GPU cluster, which can flexibly allocate the number of GPUs according to the business load requirements, meet the elastic resource scheduling requirements in different scenarios, and improve the overall system utilization; the switching component is responsible for converting the PCIe protocol signal emitted by the GPU into a format suitable for transmission on the optical link, and restoring it to the original PCIe signal at the controlled end, reducing communication latency and improving heterogeneous computing efficiency.

[0172] Embodiments of this application also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the training method for optical interconnect links when running.

[0173] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0174] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the above-described embodiment of the training method for optical interconnect links.

[0175] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described method embodiment for a shared memory resource pool.

[0176] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] The above provides a detailed description of a shared memory resource pool system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. An optical interconnect link component, characterized in that, include: The system comprises a first photoelectric converter, a second photoelectric converter, a first switching component, and a second switching component. The first and second photoelectric converters form an optical interconnect link, and the first and second switching components interact through this optical interconnect link. The first and second photoelectric converters have identical structures, each including a linear transimpedance amplifier and a linear driver. The linear transimpedance amplifier converts the optical signal output from the photodetector into an electrical signal with a target linearity. The linear driver converts the electrical signal received by the photodetector into a continuously modulated current, driving the laser to operate in the linear region to obtain an optical signal. The first switching component is configured to continuously send a first detection character sequence, send a first response character sequence upon receiving a second detection character sequence, determine whether a second response character sequence is received within a preset time period, and if the second response character sequence is received within the preset time period, control the internal state machine of the first switching component to set to an active state. The active state indicates that the optical interconnect link is allowed to enter the training process, while the inactive state indicates that the optical interconnect link is not allowed to enter the training process. Before setting the internal state machine of the first switching component to an active state, a first continuous reception count of the second response character sequence within the preset time period is identified. If the first continuous reception count is greater than or equal to a preset count, the internal state machine of the first switching component is set to an active state; if the first continuous reception count is less than the preset count, the internal state machine of the first switching component is set to an inactive state. The second switching component is configured to continuously send a second detection character sequence, send a second response character sequence upon receiving a first detection character sequence, determine whether a first response character sequence has been received within a preset time period, and if the first response character sequence has been received within the preset time period, control the internal state machine of the second switching component to set it to an active state, wherein the active state indicates that the optical interconnect link is allowed to enter the training process. Before controlling the internal state machine of the second switching component to set it to an active state, the component further includes: identifying a second continuous reception quantity of the first response character sequence within the preset time period; if the second continuous reception quantity is greater than or equal to a preset quantity, controlling the internal state machine of the second switching component to set it to an active state; if the second continuous reception quantity is less than the preset quantity, controlling the internal state machine of the second switching component to set it to an inactive state. After the first switching component and the second switching component have completed their interaction, the optical interconnect link enters the training process.

2. The optical interconnect link component according to claim 1, characterized in that, The first switching component processes the first electrical signal output by the first photoelectric converter based on the electrical specifications of the optical interconnect protocol. The second switching component processes the second electrical signal output by the second photoelectric converter based on the electrical specifications of the optical interconnect protocol.

3. The optical interconnect link component according to claim 1, characterized in that, Both the first photoelectric converter and the second photoelectric converter are connected to the substrate management controller. When the substrate management controller detects that the optical interconnect link meets the link establishment conditions, it establishes the optical interconnect link and controls the optical interconnect link to enter the training process.

4. A training method for an optical interconnect link, characterized in that, The method is applied to a first switching component of the optical interconnect link assembly according to any one of claims 1-3, wherein the method comprises: The first photoelectric converter is controlled to continuously send a first detection character sequence, the first electrical signal output by the first photoelectric converter is identified, and it is determined whether a second response character sequence is received within a preset time period. The second switching component outputs a second response character sequence after receiving the first detection character sequence. After receiving the second response character sequence within a preset time period, the internal state machine of the first switching component is controlled to be set to an effective state, wherein the effective state indicates that the optical interconnect link is allowed to enter the training process.

5. The training method for optical interconnect links according to claim 4, characterized in that, Before controlling the internal state machine of the first switching component to be set to the valid state, the following steps are also included: Identify the first continuous number of times the second response character sequence is received within the preset time period; If the first continuous reception quantity is greater than or equal to a preset quantity, then control the internal state machine of the first switching component to be set to an active state; If the first continuous reception quantity is less than a preset quantity, the internal state machine of the first switching component is controlled to be set to an invalid state, wherein the invalid state indicates that the optical interconnect link is not allowed to enter the training process.

6. The training method for optical interconnect links according to claim 5, characterized in that, Also includes: The system identifies the first electrical signal output by the first photoelectric conversion device, determines the second detection character sequence output by the second switching component, and controls the first photoelectric conversion device to send the first response character sequence.

7. The training method for optical interconnect links according to claim 4, characterized in that, Before controlling the first photoelectric converter to continuously send the first detection character sequence, the method further includes: Whether to detect the first device signal at the master control end is determined based on the type of the first switching component. If the first switching component is of the switching type, then the first photoelectric converter is controlled to continuously send the first detection character sequence; If the first switching component is a relay type, then after detecting the first device signal at the master control end, the first photoelectric converter is controlled to continuously send the first detection character sequence.

8. A training method for an optical interconnect link, characterized in that, The method is applied to a second switching component of the optical interconnect link assembly according to any one of claims 1-3, wherein the method comprises: The second photoelectric converter is controlled to continuously send the second detection character sequence, the second electrical signal output by the second photoelectric converter device is identified, and it is determined whether the first response character sequence is received within a preset time period. The first switching component outputs the first response character sequence after receiving the second detection character sequence. After receiving the first response character sequence within a preset time period, the internal state machine of the second switching component is controlled to be set to an active state, wherein the active state indicates that the optical interconnect link is allowed to enter the training process.

9. The training method for an optical interconnect link according to claim 8, characterized in that, Before controlling the internal state machine of the second switching component to be set to the valid state, the following is also included: Identify the second continuous reception quantity of the first response character sequence within the preset time period; If the second continuous reception quantity is greater than or equal to the preset quantity, then control the internal state machine of the second switching component to be set to the valid state; If the second continuous reception quantity is less than the preset quantity, the internal state machine of the second switching component is controlled to be set to an invalid state, wherein the invalid state indicates that the optical interconnect link is not allowed to enter the training process.

10. The training method for an optical interconnect link according to claim 9, characterized in that, Also includes: The second electrical signal output by the second photoelectric conversion device is identified, the first detection character sequence output by the first switching component is determined, and the second photoelectric conversion device is controlled to send the second response character sequence.

11. The training method for an optical interconnect link according to claim 8, characterized in that, Before controlling the second photoelectric converter to continuously send the second detection character sequence, the following is also included: Whether to detect the second device signal at the controlled end is determined based on the type of the second switching component; If the type of the second switching component is switching, then control the second photoelectric converter to continuously send the second detection character sequence; If the second switching component is a relay type, then after detecting the second device signal at the master control end, the second photoelectric converter is controlled to continuously send the second detection character sequence.

12. An accelerated cluster system, characterized in that, The accelerated cluster system is based on the optical interconnect link component according to any one of claims 1-3, wherein the system includes: At least one optical interconnect computing unit and at least one optical interconnect acceleration unit, wherein the optical interconnect computing unit provides computing resources and the optical interconnect acceleration unit provides acceleration resources; At least one optical interconnect switching unit, wherein the optical interconnect switching unit is connected to at least one of the optical interconnect computing units and the optical interconnect acceleration units via an optical interconnect link, and / or the optical interconnect acceleration unit is connected to at least one of the optical interconnect computing units via an optical interconnect link.

13. The accelerated cluster system according to claim 12, characterized in that, If there are multiple optical interconnect switching units, the multiple optical interconnect switching units include at least one main switching unit and at least one secondary switching unit. The main switching unit and at least one secondary switching unit are connected by an optical interconnect link, and the main switching units are connected by an optical interconnect link.

14. The accelerated cluster system according to claim 13, characterized in that, At least one of the main switching unit and the secondary switching unit is connected to at least one optical interconnect computing unit via an optical interconnect link.

15. The accelerated cluster system according to claim 14, characterized in that, The optical interconnect computing unit includes at least one central processing unit and at least one switching component, wherein the central processing unit is connected to at least one switching component, the switching component is connected to at least one optoelectronic converter, the optoelectronic converter inside the computing unit and the optoelectronic converter outside the computing unit form an optical interconnect link, and the central processing unit provides computing resources through the optical interconnect link.

16. The accelerated cluster system according to claim 15, characterized in that, The optical interconnect computing unit further includes at least one substrate management controller, which is connected to the at least one switching component and the at least one optoelectronic converter, and is used to manage the timing relationship between the switching component and the optoelectronic converter and the establishment of the optical interconnect link.

17. The accelerated cluster system according to claim 12, characterized in that, The optical interconnect switching unit includes: at least one optical interconnect switching board, wherein the optical interconnect switching board includes: A substrate connector and at least one switching component, wherein the substrate connector is connected to the at least one switching component; The switching component is connected to at least one high-speed connector and at least one optoelectronic converter, wherein the optical interconnect switching board communicates with other optical interconnect switching boards through the high-speed connector, and the optoelectronic converters of the optical interconnect switching board form an optical interconnect link with other optoelectronic converters.

18. The accelerated cluster system according to claim 17, characterized in that, The switching component includes multiple internal ports and multiple external ports, wherein the internal ports are connected to the high-speed connector and the external ports are connected to the photoelectric converter.

19. The accelerated cluster system according to claim 18, characterized in that, The configuration type of the external port includes any one of uplink port, downlink port, and interconnect port, wherein the uplink port connects to computing resources, the downlink port connects to acceleration resources, and the interconnect port connects to other switching components.

20. The accelerated cluster system according to claim 17, characterized in that, The optical interconnect switching unit further includes at least one clock buffer, wherein the clock buffer is connected to the clock port of the switching component and is used to amplify and distribute the clock signal of the switching component.

21. The accelerated cluster system according to claim 12, characterized in that, The optical interconnect acceleration unit includes at least one graphics processor and at least one switching component, wherein the switching component is connected to at least one graphics processor and at least one optoelectronic converter, the optoelectronic converter of the optical interconnect acceleration unit forms an optical interconnect link with other optoelectronic converters, and the at least one graphics processor provides acceleration resources through the optical interconnect link.

22. The accelerated cluster system according to claim 21, characterized in that, The switching component includes multiple ports, which are configured as internal ports and external ports. The internal ports are connected to the graphics processor, and the external ports are connected to the opto-converter. The configuration type of the external ports includes interconnect ports, which are connected to other switching components.

23. An optical interconnect computing unit, characterized in that, The optical interconnect computing unit is composed of the optical interconnect link component according to any one of claims 1-3, comprising: At least one central processing unit and at least one switching component, wherein the central processing unit is connected to at least one switching component, the switching component is connected to at least one optoelectronic converter, and the optoelectronic converter inside the computing unit and the optoelectronic converter outside the computing unit form an optical interconnect link.

24. An optical interconnect switching unit, characterized in that, The optical interconnect switching unit is composed of the optical interconnect link component according to any one of claims 1-3, comprising: At least one optical interconnect switching board, wherein the optical interconnect switching board comprises: A substrate connector and at least one switching component, wherein the substrate connector is connected to at least one of the switching components; The switching component is connected to at least one high-speed connector and at least one optoelectronic converter. The optical interconnect switching board communicates with other optical interconnect switching boards through the high-speed connector. The optoelectronic converter inside the optical interconnect switching board and the optoelectronic converter outside the optical interconnect switching board form an optical interconnect link.

25. An optical interconnect acceleration unit, characterized in that, The optical interconnect acceleration unit is based on the optical interconnect link component according to any one of claims 1-3, comprising: The system includes at least one graphics processor and at least one switching component, wherein the switching component is connected to at least one graphics processor and at least one optoelectronic converter, the optoelectronic converter inside the optical interconnect acceleration unit and the optoelectronic converter outside the optical interconnect acceleration unit form an optical interconnect link, and the at least one graphics processor provides acceleration resources through the optical interconnect link.

26. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the training method for the optical interconnect link according to any one of claims 4-11.

27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the training method for the optical interconnect link according to any one of claims 4-11.

Citation Information

Patent Citations

  • Link detection method, apparatus and system

    CN101340320A

  • PCIe extension line, PCIe device with extension line and signal transmission method

    CN119473975A