Communication method between graphics processors, product, equipment and medium
By employing a central optoelectronic hybrid switching chip in a standalone system, and utilizing electrical and optical switching matrices to transmit control and data flows respectively, the problems of bandwidth attenuation and topology rigidity in inter-GPU communication are solved, achieving low-latency, high-bandwidth communication and improving the efficiency of multi-GPU collaborative computing.
Patent Information
- Application Number
- CN202511518861.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
In current stand-alone systems, communication between graphics processors suffers from severe bandwidth attenuation with distance and topology rigidity, making it difficult to meet the low-latency, high-bandwidth communication requirements for gradient synchronization and scientific computing in artificial intelligence training.
A central optoelectronic hybrid switching chip based on the interconnection of electrical and optical switching matrices is adopted. Control flow and data flow are transmitted through optical and electrical links respectively, realizing direct interconnection between graphics processors. The low latency of the electrical link and the high bandwidth of the optical link are utilized to achieve physical separation transmission.
It reduces communication latency, improves the efficiency and performance of multi-GPU collaborative computing, and meets the real-time requirements of control flow and the large data volume transmission requirements of data flow.
Smart Images

Figure CN120997027A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of high-performance computing and artificial intelligence acceleration technology, and in particular to communication methods, products, devices and media between graphics processors. Background Technology
[0002] Current single-machine systems suffer from severe bandwidth attenuation with distance and topology rigidity in inter-GPU communication, making it difficult to meet the low-latency, high-bandwidth communication requirements for gradient synchronization and scientific computing in AI (Artificial Intelligence) training. In current NVLink interconnect solutions, the NVLink mesh topology uses NVIDIA DGX A100 and NVSwitch to connect multiple GPUs, forming a fully connected network. However, due to PCB (Printed Circuit Board) wiring density limitations, single-hop communication distances must be <10cm; otherwise, bandwidth will be severely attenuated. In PCIe (Peripheral Component Interconnect Express) interconnect solutions, the PCIe tree topology requires non-blocking communication via the Central Processing Unit (CPU), severely impacting communication efficiency and sharing bus bandwidth.
[0003] It is evident that optimizing the interconnection between graphics processors to improve communication efficiency is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this invention is to provide a method, apparatus, device, and medium for communication between graphics processors, optimizing the interconnection between graphics processors to improve communication efficiency. The specific solution is as follows: In a first aspect, the present invention discloses a method for communication between graphics processors, wherein the standalone system includes multiple graphics processors and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each graphics processor is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link, respectively, to realize the interconnection between the graphics processors; the method includes: The source graphics processor is determined from each of the graphics processors; The data to be transmitted from the source graphics processor is classified to determine the data type of the data to be transmitted. If the data type of the data to be transmitted is control flow type, then the source graphics processor is controlled to send the data to be transmitted to the electrical switching matrix through the electrical link, and the electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors; If the data type of the data to be transmitted is a data stream type, then the source graphics processor is controlled to send the data to be transmitted to the optical switching matrix through the optical link, and the optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors.
[0005] Optionally, classifying the data to be transmitted from the source graphics processor to determine the data type includes: Extract the data features of the data to be transmitted from the source graphics processor, and use a hardware-level traffic classifier to determine whether the data features meet a preset small data condition; wherein, the preset small data condition is that the data volume of the data to be transmitted is less than a preset data volume threshold and the latency requirement of the data to be transmitted is less than a preset requirement threshold. If the data characteristics satisfy the preset small data condition, then the data type of the data to be transmitted is determined to be control flow type; If the data characteristics do not meet the preset small data conditions, then the data type of the data to be transmitted is determined to be a data stream type.
[0006] Optionally, the step of extracting the data features of the source graphics processor's data to be transmitted includes: The amount of data to be transmitted from the source graphics processor is extracted using a data length counter, and the delay requirement of the data to be transmitted is determined according to the delay requirement mark in the transmission instruction of the data to be transmitted.
[0007] Optionally, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: The data to be transmitted is parsed to identify the destination address, and the destination graphics processor corresponding to the destination address is determined from each of the graphics processors. The electrical switching matrix is controlled to route the data to be transmitted to the electrical link port corresponding to the destination address.
[0008] Optionally, controlling the optical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: Wavelength identification is performed on the data to be transmitted to determine the target wavelength corresponding to the data to be transmitted. Determine the target graphics processor corresponding to the target wavelength from among the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor.
[0009] Optionally, controlling the source graphics processor to send the data to be transmitted to the electrical switching matrix via the electrical link includes: The source graphics processor is controlled to process the data to be transmitted through the electrical link to obtain the data to be transmitted in the form of an electrical signal, and then the data to be transmitted in the form of an electrical signal is sent to the electrical switching matrix.
[0010] Optionally, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: The electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors through the electrical link, so that the destination graphics processor generates control instructions and confirmation instructions based on the data to be transmitted, uses the control instructions to regulate the processing mechanism of the current task to be processed, and feeds back the confirmation instructions to the source graphics processor.
[0011] Optionally, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors via the electrical link includes: The target graphics processor is determined from each of the graphics processors; The electrical switching matrix is controlled to route the data to be transmitted to the target electrical link corresponding to the destination graphics processor. The target electrical link is controlled to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted, and the first target-processed data to be transmitted is sent to the target graphics processor.
[0012] Optionally, the electrical link includes a cross-group amplifier, a clock data recovery circuit, and a PAM4 modulator; controlling the target electrical link to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted includes: The cross-group amplifier is controlled to amplify the data to be transmitted to obtain the first signal-processed data to be transmitted. The clock data recovery circuit is used to eliminate the signal jitter noise in the first signal-processed data to be transmitted to obtain the second signal-processed data to be transmitted. The PAM4 modulator is controlled to restore the data to be transmitted after the second signal processing to binary data, so as to obtain the data to be transmitted after the first target processing.
[0013] Optionally, controlling the source graphics processor to send the data to be transmitted to the optical switching matrix via the optical link includes: The silicon photonics emission module in the source graphics processor is controlled to convert the data to be transmitted into optical signal form, and then send the optical signal form of the data to be transmitted to the optical switching matrix through the optical link.
[0014] Optionally, controlling the optical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: The target graphics processor is determined from each of the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the target optical link corresponding to the destination graphics processor. The data to be transmitted is routed to the target graphics processor via the target optical link.
[0015] Optionally, the optical link includes a microring resonator array; routing the data to be transmitted to the destination graphics processor via the target optical link includes: The micro-ring resonator array is controlled to route the data to be transmitted to the wavelength channel of the target graphics processor, so that the germanium-silicon detector of the target graphics processor demodulates the data to be transmitted into an electrical signal form to obtain the data to be transmitted after third signal processing, and restores the data to be transmitted after third signal processing into parallel data to obtain the data to be transmitted after second target processing.
[0016] Optionally, after routing the data to be transmitted to the destination graphics processor via the target optical link, the method further includes: The destination graphics processor is controlled to perform integrity verification on the data to be transmitted, so as to obtain the verification result, and the verification result is fed back to the source graphics processor; wherein, the verification result is a cyclic redundancy check result or a hash comparison result.
[0017] Optionally, the electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before controlling the micro-ring resonator array to route the data to be transmitted to the wavelength channel of the destination graphics processor, the method further includes: Different wavelength channels are assigned to the optical ports of each graphics processor.
[0018] Optionally, the electrical switching matrix is constructed based on a non-blocking cross switch structure, and the internal components of the electrical switching matrix are connected to each first input port and each first output port via metal interconnects. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal components of the optical switching matrix are connected to each second input port and each second output port via optical waveguides. The electrical switching matrix and the optical switching matrix are vertically connected via through-silicon vias.
[0019] Optionally, the communication hop count between any two graphics processors is no greater than 2, the electrical link between any graphics processor and the central optoelectronic hybrid switching chip is a dual-link, and the optical link between any graphics processor and the central optoelectronic hybrid switching chip is a single-link.
[0020] Secondly, the present invention discloses an inter-graphics processor communication device applied to a standalone system. The standalone system includes multiple graphics processors and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each graphics processor is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link, respectively, to achieve interconnection between the graphics processors. The device includes: A determining module is used to determine the source graphics processor from each of the graphics processors; A classification module is used to classify the data to be transmitted from the source graphics processor in order to determine the data type of the data to be transmitted. The first communication module is configured to, if the data type of the data to be transmitted is control flow type, control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors. The second communication module is used to control the source graphics processor to send the data to be transmitted to the optical switching matrix through the optical link if the data type of the data to be transmitted is a data stream type, and to control the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors.
[0021] Thirdly, the present invention discloses an electronic device, comprising: Memory, used to store computer programs; A processor for executing computer programs to implement the steps of the aforementioned disclosed inter-graphics processor communication method.
[0022] Fourthly, the present invention discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed inter-graphics processor communication method.
[0023] Fifthly, the present invention discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned disclosed inter-graphics processor communication method.
[0024] Therefore, this invention is applied to a standalone system, which includes multiple graphics processors (GPUs) and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each GPU is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link to achieve interconnection between the GPUs. The method includes: determining a source GPU from each GPU; classifying the data to be transmitted from the source GPU to determine the data type; if the data type of the data to be transmitted is control flow type, controlling the source GPU to send the data to be transmitted to the electrical switching matrix via the electrical link, and controlling the electrical switching matrix to route the data to be transmitted to the destination GPU among the GPUs; if the data type of the data to be transmitted is data stream type, controlling the source GPU to send the data to be transmitted to the optical switching matrix via the optical link, and controlling the optical switching matrix to route the data to be transmitted to the destination GPU among the GPUs.
[0025] The beneficial effects are as follows: This invention interconnects multiple graphics processors (GPUs) in a single-machine system with a central optoelectronic hybrid switching chip built on an electrical switching matrix and an optical switching matrix, respectively, via optical links and electrical links. In other words, the GPUs are directly interconnected through the central chip, reducing the number of communication hops. During data transmission, the data to be transmitted from the source GPU is first classified. Control flow types are routed to the destination GPU via electrical links and electrical switching matrices, while data flow types are routed to the destination GPU via optical links and optical switching matrices. This achieves physical separation of control flow and data flow transmission. The control flow meets real-time requirements thanks to the low latency of the electrical links, while the data flow meets the requirements for large data volume transmission thanks to the high bandwidth of the optical links, reducing transmission latency and improving the efficiency and performance of multi-GPU collaborative computing within a single-machine system. Attached Figure Description
[0026] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A flowchart of an inter-graphics processor communication method provided in an embodiment of the present invention; Figure 2 A specific single-machine system schematic diagram provided for an embodiment of the present invention; Figure 3 This is a schematic diagram of a specific optoelectronic hybrid transmission path provided in an embodiment of the present invention; Figure 4 This is a specific connection diagram of a graphics processor and a chip provided in an embodiment of the present invention; Figure 5 A schematic diagram of a specific electrical connection module provided in an embodiment of the present invention; Figure 6 A schematic diagram of a specific optical connection module provided in an embodiment of the present invention; Figure 7 A specific data communication diagram is provided for an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an inter-graphics processor communication device provided in an embodiment of the present invention; Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0029] Current single-machine systems suffer from severe bandwidth attenuation with distance and topology rigidity in inter-GPU communication, making it difficult to meet the low-latency, high-bandwidth communication requirements for gradient synchronization and scientific computing in AI training. In current NVLink interconnect solutions, the NVLink mesh topology uses NVIDIA DGX A100 and NVSwitch to connect multiple GPUs, forming a fully connected network. However, due to PCB wiring density limitations, the single-hop communication distance must be less than 10cm; otherwise, bandwidth will be severely attenuated. In PCIe interconnect solutions, the PCIe tree topology requires non-blocking communication via the central processing unit, severely impacting communication efficiency and sharing bus bandwidth.
[0030] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.
[0031] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] Next, we will describe in detail a communication scheme between graphics processors provided by an embodiment of the present invention. Figure 1 This invention provides a method for communication between graphics processors (GPUs), applied to a standalone system. The standalone system includes multiple GPUs and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each GPU is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link, respectively, to achieve interconnection between the GPUs. The method includes: Step S11: Identify the source graphics processor from each of the graphics processors.
[0033] The standalone system uses a hybrid optoelectronic switching architecture, such as... Figure 2 As shown, taking a single-machine system with 8 graphics processors as an example, the system achieves full interconnection of 8 GPUs. The system consists of three core modules: a central optoelectronic hybrid switching chip, 8 GPU computing nodes, and an intelligent control platform. The central optoelectronic hybrid switching chip uses 3D heterogeneous integration technology to vertically stack a 16×16 electrical switching matrix and an 8×8 optical switching matrix to form a unified switching plane. Each GPU node is equipped with a dual-mode network interface, connected to the central switching chip via a high-speed electrical link (NVLink compatible) and an optical link (fixed wavelength). The intelligent control platform monitors the system status in real time and dynamically optimizes the topology configuration and resource allocation. The system adopts a star topology to ensure that the maximum hop count between any two GPUs does not exceed 2. Simultaneously, through an optoelectronic collaborative transmission strategy, it intelligently allocates control flow (electrical transmission) and data flow (optical transmission).
[0034] The source graphics processor is determined from among the graphics processors. The source graphics processor is the one that currently contains data that needs to be transmitted. In other words, the source graphics processor is the source of communication. The number of source graphics processors can be single or multiple, depending on the specific communication situation.
[0035] In this embodiment, the electrical switching matrix is constructed based on a non-blocking cross switch structure, and the internal components of the electrical switching matrix are connected to each first input port and each first output port through metal interconnects. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal components of the optical switching matrix are connected to each second input port and each second output port through optical waveguides. The electrical switching matrix and the optical switching matrix are vertically connected through through-silicon vias.
[0036] The electrical switching matrix is constructed based on a non-blocking cross switch structure. Internally, the first input port and the first output port are connected by metal interconnects, enabling non-blocking routing of control flow between ports and ensuring low-latency transmission of control flow. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism. Internally, the second input port and the second output port are connected by optical waveguides. Relying on a micro-ring resonator array, direct routing of optical signals in the optical domain is achieved, meeting the high-bandwidth transmission requirements of data flow. The electrical and optical switching matrices are vertically connected through through-silicon vias (TSVs), achieving efficient coordinated scheduling of electrical and optical signals. This structural design allows the electrical and optical switching matrices to leverage their respective advantages in control flow and data flow transmission while achieving tight interconnection through TSVs, reducing signal transmission loss, improving the integration and communication efficiency of the entire central optoelectronic hybrid switching chip, and thus enhancing the overall performance of multi-GPU interconnection in a single system.
[0037] In this embodiment, the communication hop count between any two graphics processors is no greater than 2, the electrical link between any graphics processor and the central optoelectronic hybrid switching chip is a dual link, and the optical link between any graphics processor and the central optoelectronic hybrid switching chip is a single link.
[0038] The communication hop count between any two graphics processors is no greater than 2. This means that the data to be transmitted from the source graphics processor can reach the destination graphics processor after being routed through the central optoelectronic hybrid switching chip, without the need for multi-level forwarding. The electrical link between any graphics processor and the central optoelectronic hybrid switching chip is a dual-link (supporting bidirectional communication, with a unidirectional bandwidth of 112Gbps), which can ensure the bidirectional and efficient transmission of the control flow. The optical link between any graphics processor and the central optoelectronic hybrid switching chip is a single-link (corresponding to a dedicated wavelength, with a bandwidth of 112Gbps), which is adapted to the unidirectional high-bandwidth transmission requirements of the data flow. This design reduces transmission latency by reducing the number of communication hops. The configuration of dual electrical links and a single optical link not only meets the low-latency requirements of bidirectional interaction of the control flow, but also ensures the high-bandwidth characteristics of unidirectional transmission of the data flow. At the same time, it improves the link redundancy capability and resource utilization efficiency, and enhances the stability and overall performance of multi-graphics processor collaborative computing.
[0039] Step S12: Classify the data to be transmitted from the source graphics processor to determine the data type of the data to be transmitted.
[0040] In this embodiment, classifying the data to be transmitted from the source graphics processor to determine the data type includes: extracting data features of the data to be transmitted from the source graphics processor, and using a hardware-level traffic classifier to determine whether the data features meet a preset small data condition; wherein, the preset small data condition is that the data volume of the data to be transmitted is less than a preset data volume threshold and the latency requirement of the data to be transmitted is less than a preset requirement threshold; if the data features meet the preset small data condition, the data type of the data to be transmitted is determined to be a control flow type; if the data features do not meet the preset small data condition, the data type of the data to be transmitted is determined to be a data stream type.
[0041] The data characteristics of the data to be transmitted from the source graphics processor are extracted. These characteristics include data size and latency requirements. A hardware-level traffic classifier is used to determine whether the data characteristics meet a preset small data condition. The preset small data condition is that the data size to be transmitted is less than a preset data size threshold and the latency requirement is less than a preset requirement threshold. Specifically, the preset data size threshold can be 4KB, and the preset requirement threshold is 35ns. That is, if the data size is less than 4KB and the latency requirement is less than 35ns, then the preset small data condition is met. If the data characteristics meet the preset small data condition, the data type of the data to be transmitted is determined to be control flow; if the data characteristics do not meet the preset small data condition, the data type of the data to be transmitted is determined to be data stream. This classification method achieves accurate determination of data type through hardware-level real-time processing, ensuring that control flow can be transmitted through electrical links adapted to its low latency requirements, and data stream can be transmitted through optical links adapted to its high bandwidth requirements. This avoids conflicts caused by different types of data being transmitted on the same link. Furthermore, hardware-level classification is fast, adds almost no extra latency, and effectively improves the overall efficiency of data transmission.
[0042] In this embodiment, extracting the data features of the data to be transmitted from the source graphics processor includes: extracting the data volume of the data to be transmitted from the source graphics processor using a data length counter, and determining the delay requirement of the data to be transmitted based on the delay requirement marker in the transmission instruction of the data to be transmitted.
[0043] By utilizing the data length counter integrated into the source graphics processor, the data to be transmitted output by the source graphics processor is counted in real time to directly obtain the specific data volume, such as the number of bytes. Simultaneously, the transmission instructions corresponding to the data to be transmitted are read. By identifying the latency requirement markers carried in the instructions, such as the "low latency marker" in gradient synchronization instructions in AI training and the "high bandwidth marker" in memory block migration instructions, it is determined whether the data to be transmitted needs to meet the preset requirement of end-to-end latency <35ns, thus clarifying its latency requirement. This data feature extraction method relies on the hardware module (i.e., the data length counter) and the markers carried in the instructions, without the need for additional software calculations or data parsing steps. It can quickly and accurately obtain the two core features of data volume and latency requirement, laying the foundation for the subsequent hardware-level traffic classifier to accurately determine the data type. At the same time, it avoids the additional latency caused by software extraction and ensures the overall efficiency of data transmission.
[0044] Furthermore, the preset small data conditions can also include communication between the source and destination GPUs within the same PCB board, meaning the distance between them is less than 10cm. In other words, the preset small data conditions are: the amount of data to be transmitted is less than a preset data amount threshold, the latency requirement is less than a preset requirement threshold, and the source and destination GPUs communicate within the same PCB board. For example, if the data to be transmitted is high-frequency small data (<4KB), such as gradient synchronization signals, barrier synchronization signals, or atomic operation signals, and has low latency requirements, such as real-time control instructions requiring a latency of less than 35ns, and communication is within the same PCB board (GPU spacing <10cm), then the data type of the data to be transmitted is control flow type. Conversely, if the data to be transmitted is large data (>4KB), such as model weight update data, memory copy data, video frame transmission data, or bandwidth-intensive operation data, such as requiring a continuous throughput of >100GB / s, with insensitive latency, or cross-board communication (GPU spacing >10cm, connected via onboard fiber optics), then the data type of the transmitted data is data stream type. Taking Verilog as an example: always @(packet_header) begin; if (packet.distance <= 10 || packet.latency<35 ||packet.size <4096); / / Connection distance less than 10cm or latency requirement less than 35ns or data packet size less than 4k; route_to_electrical(); / / Data travels via the electrical link; else; route_to_optical(); / / Data routing via optical path; end.
[0045] Step S13: If the data type of the data to be transmitted is control flow type, then control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors.
[0046] In this embodiment, controlling the source graphics processor to send the data to be transmitted to the electrical switching matrix via the electrical link includes: controlling the source graphics processor to perform signal processing on the data to be transmitted via the electrical link to obtain the data to be transmitted in the form of an electrical signal, and sending the data to be transmitted in the form of an electrical signal to the electrical switching matrix.
[0047] The control source GPU activates the PAM4 modulator in the electrical link, encodes the data to be transmitted into a 4-level electrical signal, completes signal processing to obtain the data to be transmitted in electrical signal form, and then controls the transmission of this electrical signal through the metal interconnect of the electrical link to the first input port of the electrical switching matrix, sending the data to be transmitted in electrical signal form to the electrical switching matrix. This signal processing and transmission method relies on PAM4 modulation technology to improve the bandwidth efficiency of the electrical link. Combined with the low loss characteristics of the metal interconnect, it ensures high signal integrity and low latency (end-to-end <35ns) of the control flow during transmission, effectively meeting the needs of real-time interaction of the control flow and ensuring the rapid transmission of control commands between multiple GPUs.
[0048] In this embodiment, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: parsing the data to be transmitted to identify the destination address, and determining the destination graphics processor corresponding to the destination address from each of the graphics processors; controlling the electrical switching matrix to route the data to be transmitted to the electrical link port corresponding to the destination address.
[0049] After receiving the data to be transmitted (control flow type) in the form of electrical signals, the electrical switching matrix parses it through its built-in address resolution module, extracts the destination address information (such as the destination GPU identifier), and determines the corresponding destination GPU from multiple GPUs based on the destination address. Then, it controls the non-blocking crossbar switch array of the electrical switching matrix to operate. Based on the connection relationship of the metal interconnects, the data to be transmitted is routed non-blockingly from the current input port to the electrical link output port corresponding to the destination address, so that the data can be transmitted to the destination GPU through the electrical link. This achieves fast directional transmission of control flow with a routing delay of <10ns, ensuring that control commands can be accurately delivered to the target in a short time, and improving the real-time performance and reliability of control interaction between multiple GPUs.
[0050] In this embodiment, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors through the electrical link, so that the destination graphics processor generates control instructions and confirmation instructions based on the data to be transmitted, uses the control instructions to regulate the processing mechanism of the current task to be processed, and feeds back the confirmation instructions to the source graphics processor.
[0051] After receiving data (control flow type) in the form of electrical signals, the electrical switching matrix parses it using its built-in address resolution module, extracting the destination address information, such as the destination GPU identifier. Based on this destination address, it determines the corresponding target GPU from among multiple GPUs. Subsequently, it controls the non-blocking crossbar switch array of the electrical switching matrix to operate. Based on the connection relationship of the metal interconnects, the data to be transmitted is routed non-blockingly from the current input port to the electrical link output port corresponding to the destination address, so that the data can be transmitted to the target GPU through the electrical link. This routing method, relying on address resolution and non-blocking crossbar switch design, achieves fast directional transmission of control flow with a routing latency of <10ns, ensuring that control commands can be accurately delivered to the target in a short time, and improving the real-time performance and reliability of control interaction between multiple GPUs.
[0052] In this embodiment, controlling the electrical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors via the electrical link includes: determining the destination graphics processor from among the graphics processors; controlling the electrical switching matrix to route the data to be transmitted to the target electrical link corresponding to the destination graphics processor; controlling the target electrical link to perform signal processing on the data to be transmitted to obtain first target-processed data to be transmitted, and sending the first target-processed data to be transmitted to the destination graphics processor.
[0053] Based on the destination address parsed from the data to be transmitted, the corresponding target graphics processor is determined from each graphics processor. The non-blocking cross switch structure of the electrical switching matrix is controlled to operate. Based on the metal interconnect connection relationship, the data to be transmitted (control flow in the form of electrical signals) is routed from the input port to the target electrical link corresponding to the target graphics processor. Then, the target electrical link is controlled to amplify and clock-restore the electrical signal to obtain the first target-processed data to be transmitted with optimized signal integrity, and then transmitted to the target graphics processor through the electrical link.
[0054] In this embodiment, the electrical link includes a cross-group amplifier, a clock data recovery circuit, and a PAM4 modulator. Controlling the target electrical link to perform signal processing on the data to be transmitted to obtain first target-processed data to be transmitted includes: controlling the cross-group amplifier to amplify the data to be transmitted to obtain first signal-processed data to be transmitted; eliminating signal jitter noise in the first signal-processed data to be transmitted through the clock data recovery circuit to obtain second signal-processed data to be transmitted; and controlling the PAM4 modulator to restore the second signal-processed data to be transmitted into binary data to obtain the first target-processed data to be transmitted.
[0055] The electrical link includes a transimpedance amplifier (TIA), a clock and data recovery circuit (CDR), and a PAM4 modulator (4-level pulse amplitude modulation modulator). The TIA amplifies the data to be transmitted (control flow in electrical signal form) routed to the target electrical link via the electrical switching matrix, compensates for signal attenuation during transmission, and obtains the first signal-processed data to be transmitted. The clock and data recovery circuit performs clock synchronization and jitter noise elimination on the first signal-processed data to be transmitted, corrects timing deviations in signal transmission, and obtains the second signal-processed data to be transmitted. The PAM4 modulator demodulates the processed 4-level signal to restore it to binary data, recovers the original control flow information, and obtains the first target-processed data to be transmitted, effectively ensuring the signal integrity and accuracy of the control flow in the electrical link transmission.
[0056] Step S14: If the data type of the data to be transmitted is a data stream type, then control the source graphics processor to send the data to be transmitted to the optical switching matrix through the optical link, and control the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors.
[0057] In this embodiment, controlling the source graphics processor to send the data to be transmitted to the optical switching matrix via the optical link includes: controlling the silicon photonics emission module in the source graphics processor to convert the data to be transmitted into optical signal form, and sending the optical signal form of the data to be transmitted to the optical switching matrix via the optical link.
[0058] The silicon photonics emission module in the control source graphics processor, which is bound to the GPU's dedicated wavelength (such as the corresponding wavelengths in λ1-λ8), converts the electrical signal of the data to be transmitted (data stream type) into an optical signal of the corresponding wavelength, forming the data to be transmitted in the form of an optical signal. Then, the optical signal is controlled to be transmitted through the optical waveguide of the optical link to the second input port corresponding to the optical switching matrix of the central optoelectronic hybrid switching chip, and the data to be transmitted in the form of an optical signal is sent to the optical switching matrix. The efficient conversion of electrical to optical signals is achieved by relying on the silicon photonics emission module. With the optical waveguide transmission of the dedicated wavelength, the data stream is ensured to be transmitted with a high throughput of 896Gbps total bandwidth in the optical link, and the optical signal transmission loss is low, which meets the bandwidth requirements of large data volume transmission and improves the efficiency of large-scale data interaction between multiple GPUs.
[0059] In this embodiment, controlling the optical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors includes: performing wavelength identification on the data to be transmitted to determine the target wavelength corresponding to the data to be transmitted; determining the destination graphics processor corresponding to the target wavelength from each of the graphics processors; and controlling the optical switching matrix to route the data to be transmitted to the destination graphics processor.
[0060] After receiving the data to be transmitted in the form of an optical signal, the optical switching matrix identifies the target wavelength (such as a specific wavelength among λ1-λ8) through a wavelength detection unit. According to the preset wavelength-graphics processor mapping relationship (each wavelength uniquely corresponds to one graphics processor), the target graphics processor corresponding to the target wavelength is determined from among the graphics processors. Then, the non-blocking optical cross-switching action of the micro-ring resonator array in the optical switching matrix is controlled. By adjusting the resonant state of the corresponding micro-ring, the optical signal of the wavelength is routed from the current input optical waveguide to the output optical waveguide corresponding to the target graphics processor, thereby routing the data to be transmitted to the target graphics processor. This wavelength identification-based routing method relies on the fast response characteristics of the micro-ring resonator array to achieve direct routing of the data stream in the optical domain without photoelectric conversion. Moreover, the total bandwidth of the parallel transmission of 8 dedicated wavelengths reaches 896Gbps, meeting the high bandwidth requirements of large data volume transmission and improving the efficiency and real-time performance of large-scale data interaction between multiple GPUs.
[0061] In this embodiment, controlling the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors includes: determining the destination graphics processor from among the graphics processors; controlling the optical switching matrix to route the data to be transmitted to the target optical link corresponding to the destination graphics processor; and routing the data to be transmitted to the destination graphics processor through the target optical link.
[0062] Based on the destination identifier of the data to be transmitted, the corresponding target graphics processor is determined from each graphics processor. The micro-ring resonator array in the optical switching matrix is controlled to adjust the resonant frequency of a specific micro-ring, so that the data to be transmitted (i.e., the data stream in the form of optical signal) is switched from the current input optical waveguide to the target optical link bound to the target graphics processor. Then, the data to be transmitted in the form of optical signal is directly transmitted to the optical receiving port of the target graphics processor through the optical waveguide of the target optical link, realizing data routing. Due to the direct optical domain switching characteristic of the micro-ring resonator, there is no need for intermediate photoelectric conversion links, and the dedicated wavelength design of the target optical link ensures high bandwidth transmission of the data stream, meeting the needs of efficient transmission of large amounts of data.
[0063] In this embodiment, the optical link includes a microring resonator array; the step of routing the data to be transmitted to the target graphics processor through the target optical link includes: controlling the microring resonator array to route the data to be transmitted to the wavelength channel of the target graphics processor, so that the germanium-silicon detector of the target graphics processor demodulates the data to be transmitted into an electrical signal form to obtain the data to be transmitted after third signal processing, and restores the data to be transmitted after third signal processing into parallel data to obtain the data to be transmitted after second target processing.
[0064] The micro-ring resonator array in the target optical link is controlled to adjust the bias voltage of the corresponding micro-rings so that they resonate with the wavelength of the optical signal of the data to be transmitted. The optical signal is then precisely routed from the output of the optical switching matrix to the wavelength channel dedicated to the target graphics processor. The germanium-silicon detector of the target graphics processor receives the optical signal and demodulates it into an electrical signal, obtaining the data to be transmitted after the third signal processing. The electrical signal is then restored to parallel data through a serial-to-parallel conversion circuit, obtaining the data to be transmitted after the second target processing. This achieves low-loss transmission and reduces the additional overhead of signal conversion and transmission.
[0065] In this embodiment, after routing the data to be transmitted to the destination graphics processor via the target optical link, the method further includes: controlling the destination graphics processor to perform integrity verification on the data to be transmitted to obtain a verification result, and feeding the verification result back to the source graphics processor; wherein the verification result is a cyclic redundancy check result or a hash comparison result.
[0066] The target GPU initiates its built-in verification module to perform integrity verification on the received, processed data to be transmitted from the second target. Specifically, this is done by calculating the cyclic redundancy check (CRC) value or hash value of the data and comparing it with the verification information pre-attached to the data by the source GPU. The verification result is either a pass or a fail. Subsequently, the target GPU feeds back the verification result to the source GPU via an electrical link. The verification result is either a CRC result or a hash comparison result. This verification mechanism can promptly detect data errors that may occur during optical link transmission due to signal attenuation, crosstalk, etc., ensuring the accuracy of large-scale data stream transmission. Furthermore, the verification result is fed back via a low-latency electrical link, enabling the source GPU to respond quickly. This ensures the reliability of data transmission while preventing the accumulation of errors from affecting the correctness of multi-GPU collaborative computing, thus improving the overall stability of the single-machine system.
[0067] In this embodiment, the electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before controlling the micro-ring resonator array to route the data to be transmitted to the wavelength channel of the target graphics processor, the embodiment further includes: allocating different wavelength channels to the optical ports of each graphics processor.
[0068] Electrical links are high-speed differential signal lines or electrical connections used to carry low-latency transmission of control flow, while optical links are onboard optical fibers or silicon-based optical waveguides adapted to the high-bandwidth transmission requirements of data flow. Before the control micro-ring resonator array routes the data to be transmitted to the wavelength channel of the destination graphics processor, different wavelength channels are allocated to the optical ports of each graphics processor based on wavelength division multiplexing technology. For example, 8 GPUs correspond to λ1-λ8 respectively, with a wavelength interval of 100GHz. This link type selection and wavelength channel allocation design allows electrical links and optical links to adapt to the transmission characteristics of control flow and data flow respectively. At the same time, crosstalk between different GPU data transmissions in the optical link is avoided through dedicated wavelength channels. Combined with the low-loss characteristics of onboard optical fibers or silicon-based optical waveguides, collision-free transmission with a total bandwidth of 896Gbps is achieved, improving the efficiency and reliability of data interaction between multiple GPUs.
[0069] Therefore, this invention is applied to a standalone system, which includes multiple graphics processors (GPUs) and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each GPU is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link to achieve interconnection between the GPUs. The method includes: determining a source GPU from each GPU; classifying the data to be transmitted from the source GPU to determine the data type; if the data type of the data to be transmitted is control flow type, controlling the source GPU to send the data to be transmitted to the electrical switching matrix via the electrical link, and controlling the electrical switching matrix to route the data to be transmitted to the destination GPU among the GPUs; if the data type of the data to be transmitted is data stream type, controlling the source GPU to send the data to be transmitted to the optical switching matrix via the optical link, and controlling the optical switching matrix to route the data to be transmitted to the destination GPU among the GPUs.
[0070] The beneficial effects are as follows: This invention interconnects multiple graphics processors (GPUs) in a single-machine system with a central optoelectronic hybrid switching chip built on an electrical switching matrix and an optical switching matrix, respectively, via optical links and electrical links. In other words, the GPUs are directly interconnected through the central chip, reducing the number of communication hops. During data transmission, the data to be transmitted from the source GPU is first classified. Control flow types are routed to the destination GPU via electrical links and electrical switching matrices, while data flow types are routed to the destination GPU via optical links and optical switching matrices. This achieves physical separation of control flow and data flow transmission. The control flow meets real-time requirements thanks to the low latency of the electrical links, while the data flow meets the requirements for large data volume transmission thanks to the high bandwidth of the optical links, reducing transmission latency and improving the efficiency and performance of multi-GPU collaborative computing within a single-machine system.
[0071] The following describes the inter-graphics processor communication of the present invention. Figure 3 The diagram shown is a specific example of a hybrid optoelectronic transmission path. Figure 3 This includes GPU0, GPU1, GPU4, and GPU7, all of which are connected to the central optoelectronic hybrid switching chip via electrical and optical links. For example... Figure 4As shown, each GPU node features a highly integrated dual-mode interconnect interface chip, integrating electrical control and optical data channels through 3D packaging. The electrical channel supports the 112Gbps NVLink protocol and employs adaptive equalization technology to ensure signal integrity. The optical channel integrates a tunable laser and a germanium-silicon detector, supporting parallel transmission of eight fixed wavelengths (λ1-λ8), ultimately connecting to a central optoelectronic hybrid switching chip. The intelligent offloading controller performs hardware-level packet analysis. Taking an 8-GPU system as an example, a performance comparison was conducted, and the results are shown in Table 1 below. Table 1 Performance Comparison Results In this embodiment, the electrical connection module adopts an improved NVLink mechanism, using PAM4 modulation, with a single-channel rate of 112Gbps. Figure 5 As shown, the electrical connection module adopts a layered architecture, integrating a protocol engine, hardware offloading unit, redundant controller, SerDes array, adaptive equalization module, PCB interface, and spare traces, and connecting to the central optoelectronic hybrid switching chip via main and spare traces. Table 2 shows a comparison of the transmission distance between the improved NVLink mechanism and the original NVLink mechanism. Table 2 Comparison of transmission distances In this embodiment, the optical connectivity module adopts a silicon photonics integration mechanism, overcoming the limitations of optical module size and energy efficiency, and providing a high-bandwidth, low-latency data channel for full 8-GPU interconnection, such as... Figure 6 As shown, the optical connection module includes a laser array, a micro-ring modulator, a germanium-silicon detector, a wavelength division multiplexer, an optical fiber coupler, a clock recovery circuit, and a TIA array. Figure 7 As shown, the photoelectric conversion system module can communicate the data to be transmitted to the corresponding graphics processor according to the electrical signal flow and optical signal flow. To verify the communication performance advantages of the present invention based on the photoelectric hybrid interconnection mechanism, a comprehensive test of key indicators was conducted based on an 8-GPU single-machine full interconnection scenario, and a comparative analysis was performed with a pure electrical interconnection scheme (NVLink multi-hop, PCIe tree topology). The test environment included: 1) Hardware platform: 8×NVIDIA H20 GPUs, central photoelectric hybrid switching chip, silicon photonics module (λ1-λ8@112Gbps); 2) Software configuration: CUDA 12.4, custom communication protocol stack, AllReduce benchmark tool; 3) Comparison scheme: NVIDIA DGX H20 (NVSwitch full connection), PCIe 5.0 x16 tree topology; The multi-dimensional performance comparison results are shown in Table 3. Table 3. Multi-dimensional performance comparison results In the test scenario of AI model training, this embodiment reduces gradient synchronization time by 42% in ResNet-152 distributed training; and reduces communication overhead by 58% in scientific computing, such as CFD simulation, and shortens the overall task completion time by 35%.
[0072] This embodiment achieves both ultra-high bandwidth and low latency. Through the collaborative operation of the electrical plane (NVLink compatible) and the optical plane (fixed wavelength routing), the control flow (<4KB data) uses electrical transmission to achieve an end-to-end latency of 35ns, while the data flow (>4KB) achieves an aggregated bandwidth of 896Gbps through 8 wavelength optical channels. This solves the problem of severe bandwidth attenuation with distance in traditional pure electrical interconnect solutions. The AllReduce operation bandwidth is increased to 896GB / s (a 198% improvement over the previous NVLink multi-hop solution), and the worst-case end-to-end latency is reduced to 55ns (a 54% improvement over the previous NVLink solution), significantly accelerating AI training and scientific computing tasks.
[0073] This embodiment has the advantage of breakthrough optimization in energy efficiency ratio. Silicon-based photonic integration technology significantly reduces photoelectric conversion energy consumption, and the optical channel energy efficiency ratio is reduced to 2.8pJ / bit (33% energy saving compared to traditional electrical interconnects). At the same time, 3D heterogeneous packaging shortens the interconnect distance and reduces signal transmission loss.
[0074] This embodiment enhances topology flexibility and scalability, breaks through PCB wiring density limitations, and supports GPU spacing of up to 20cm within a single machine (100% larger than NVLink); the optoelectronic dual-plane architecture supports dynamic switching of various topology modes such as full connectivity, multicast, and ring, and intelligent routing strategies to adapt to different load scenarios.
[0075] Figure 8 This is a schematic diagram of a communication device between graphics processors provided in an embodiment of the present invention, applied to a standalone system. The standalone system includes multiple graphics processors and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each graphics processor is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link, respectively, to realize the interconnection between the graphics processors. The device includes: Determining module 11 is used to determine the source graphics processor from each of the graphics processors; The classification module 12 is used to classify the data to be transmitted from the source graphics processor to determine the data type of the data to be transmitted. The first communication module 13 is configured to, if the data type of the data to be transmitted is control flow type, control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors. The second communication module 14 is used to control the source graphics processor to send the data to be transmitted to the optical switching matrix through the optical link if the data type of the data to be transmitted is a data stream type, and to control the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors.
[0076] Furthermore, embodiments of this application also disclose an electronic device, Figure 9 This is a structural diagram of an electronic device according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. Specifically, the electronic device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the inter-graphics processor communication method disclosed in any of the foregoing embodiments. Furthermore, the electronic device in this embodiment may specifically be an electronic computer.
[0077] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0078] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0079] The operating system 221 is used to manage and control the various hardware devices on the electronic device and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the inter-graphics processor communication method executed by the electronic device as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.
[0080] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned inter-graphics processor communication method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0081] Furthermore, embodiments of this application also disclose a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the graphics processor inter-communication method disclosed in any of the foregoing embodiments.
[0082] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0083] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0084] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0085] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0086] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for communication between graphics processors, characterized in that, The method is applied to a standalone system, which includes multiple graphics processors (GPUs) and a central optoelectronic hybrid switching chip built based on the interconnection of electrical and optical switching matrices. Each GPU is connected to the central optoelectronic hybrid switching chip via optical and electrical links to achieve interconnection between the GPUs. The source graphics processor is determined from each of the graphics processors; The data to be transmitted from the source graphics processor is classified to determine the data type of the data to be transmitted. If the data type of the data to be transmitted is control flow type, then the source graphics processor is controlled to send the data to be transmitted to the electrical switching matrix through the electrical link, and the electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors; If the data type of the data to be transmitted is a data stream type, then the source graphics processor is controlled to send the data to be transmitted to the optical switching matrix through the optical link, and the optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors.
2. The inter-graphics processor communication method according to claim 1, characterized in that, The step of classifying the data to be transmitted from the source graphics processor to determine the data type includes: Extract the data features of the data to be transmitted from the source graphics processor, and use a hardware-level traffic classifier to determine whether the data features meet a preset small data condition; wherein, the preset small data condition is that the data volume of the data to be transmitted is less than a preset data volume threshold and the latency requirement of the data to be transmitted is less than a preset requirement threshold. If the data characteristics satisfy the preset small data condition, then the data type of the data to be transmitted is determined to be control flow type; If the data characteristics do not meet the preset small data conditions, then the data type of the data to be transmitted is determined to be a data stream type.
3. The inter-graphics processor communication method according to claim 2, characterized in that, The extraction of data features from the source graphics processor's data to be transmitted includes: The amount of data to be transmitted from the source graphics processor is extracted using a data length counter, and the delay requirement of the data to be transmitted is determined according to the delay requirement mark in the transmission instruction of the data to be transmitted.
4. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The data to be transmitted is parsed to identify the destination address, and the destination graphics processor corresponding to the destination address is determined from each of the graphics processors. The electrical switching matrix is controlled to route the data to be transmitted to the electrical link port corresponding to the destination address.
5. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the optical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: Wavelength identification is performed on the data to be transmitted to determine the target wavelength corresponding to the data to be transmitted. Determine the target graphics processor corresponding to the target wavelength from among the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor.
6. The inter-graphics processor communication method according to claim 1, characterized in that, The step of controlling the source graphics processor to send the data to be transmitted to the electrical switching matrix via the electrical link includes: The source graphics processor is controlled to process the data to be transmitted through the electrical link to obtain the data to be transmitted in the form of an electrical signal, and then the data to be transmitted in the form of an electrical signal is sent to the electrical switching matrix.
7. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors through the electrical link, so that the destination graphics processor generates control instructions and confirmation instructions based on the data to be transmitted, uses the control instructions to regulate the processing mechanism of the current task to be processed, and feeds back the confirmation instructions to the source graphics processor.
8. The inter-graphics processor communication method according to claim 7, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors via the electrical link includes: The target graphics processor is determined from each of the graphics processors; The electrical switching matrix is controlled to route the data to be transmitted to the target electrical link corresponding to the destination graphics processor. The target electrical link is controlled to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted, and the first target-processed data to be transmitted is sent to the target graphics processor.
9. The inter-graphics processor communication method according to claim 8, characterized in that, The electrical link includes a cross-group amplifier, a clock data recovery circuit, and a PAM4 modulator; controlling the target electrical link to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted includes: The cross-group amplifier is controlled to amplify the data to be transmitted to obtain the first signal-processed data to be transmitted. The clock data recovery circuit is used to eliminate the signal jitter noise in the first signal-processed data to be transmitted to obtain the second signal-processed data to be transmitted. The PAM4 modulator is controlled to restore the data to be transmitted after the second signal processing to binary data, so as to obtain the data to be transmitted after the first target processing.
10. The inter-graphics processor communication method according to claim 1, characterized in that, The step of controlling the source graphics processor to send the data to be transmitted to the optical switching matrix via the optical link includes: The silicon photonics emission module in the source graphics processor is controlled to convert the data to be transmitted into optical signal form, and then send the optical signal form of the data to be transmitted to the optical switching matrix through the optical link.
11. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the optical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The target graphics processor is determined from each of the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the target optical link corresponding to the destination graphics processor. The data to be transmitted is routed to the target graphics processor via the target optical link.
12. The inter-graphics processor communication method according to claim 11, characterized in that, The optical link includes a microring resonator array; routing the data to be transmitted to the destination graphics processor via the target optical link includes: The micro-ring resonator array is controlled to route the data to be transmitted to the wavelength channel of the target graphics processor, so that the germanium-silicon detector of the target graphics processor demodulates the data to be transmitted into an electrical signal form to obtain the data to be transmitted after third signal processing, and restores the data to be transmitted after third signal processing into parallel data to obtain the data to be transmitted after second target processing.
13. The inter-graphics processor communication method according to claim 11, characterized in that, After routing the data to be transmitted to the destination graphics processor via the target optical link, the method further includes: The destination graphics processor is controlled to perform integrity verification on the data to be transmitted, so as to obtain the verification result, and the verification result is fed back to the source graphics processor; wherein, the verification result is a cyclic redundancy check result or a hash comparison result.
14. The inter-graphics processor communication method according to claim 12, characterized in that, The electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before controlling the micro-ring resonator array to route the data to be transmitted to the wavelength channel of the target graphics processor, the method further includes: Different wavelength channels are assigned to the optical ports of each graphics processor.
15. The inter-graphics processor communication method according to claim 1, characterized in that, The electrical switching matrix is constructed based on a non-blocking cross switch structure, and the internal components of the electrical switching matrix are connected to each first input port and each first output port via metal interconnects. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal components of the optical switching matrix are connected to each second input port and each second output port via optical waveguides. The electrical switching matrix and the optical switching matrix are vertically connected via through-silicon vias.
16. The inter-graphics processor communication method according to claim 1, characterized in that, The communication hop count between any two graphics processors is no greater than 2, the electrical link between any graphics processor and the central optoelectronic hybrid switching chip is a dual link, and the optical link between any graphics processor and the central optoelectronic hybrid switching chip is a single link.
17. A communication device between graphics processors, characterized in that, The device is applied to a standalone system, which includes multiple graphics processors and a central optoelectronic hybrid switching chip built based on the interconnection of electrical and optical switching matrices. Each graphics processor is connected to the central optoelectronic hybrid switching chip via optical and electrical links to achieve interconnection between the graphics processors. The device includes: A determining module is used to determine the source graphics processor from each of the graphics processors; A classification module is used to classify the data to be transmitted from the source graphics processor in order to determine the data type of the data to be transmitted. The first communication module is configured to, if the data type of the data to be transmitted is control flow type, control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors. The second communication module is used to control the source graphics processor to send the data to be transmitted to the optical switching matrix through the optical link if the data type of the data to be transmitted is a data stream type, and to control the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors.
18. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the inter-graphics processor communication method as described in any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the inter-graphics processor communication method as described in any one of claims 1 to 16.
20. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the inter-graphics processor communication method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Photoelectric hybrid switching method and device based on QoS flow classification in data center
CN113472685A
Distributed comprehensive reconfigurable electronic system platform architecture
CN114697321A
Flow group scheduling method in photoelectric hybrid data center network
CN114827782A
Photoelectric hybrid switching system, transmission method and device, GPU server and medium
CN118826874A
Low-delay transmission method and system for ultra-high bandwidth OLED (Organic Light Emitting Diode) display drive
CN120412475A
Cited By
Photoelectric double-plane cross-layer interconnection scheduling method based on flow behavior perception
CN121357078A
A pop-behavior-based optoelectronic biplane cross-layer interconnect scheduling method
CN121357078B
Parallel processor cluster, topology switching method, electronic equipment and storage medium
CN121455890A