Inter-graphics processor communication method, product, device and medium
By employing a central optoelectronic hybrid switching chip in a standalone system, and utilizing electrical and optical switching matrices to transmit control and data streams respectively, the problems of bandwidth attenuation and topology rigidity in inter-GPU communication are solved, achieving low-latency, high-bandwidth communication and improving the efficiency of GPU collaborative computing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-03-20
AI Technical Summary
In current stand-alone systems, communication between graphics processors suffers from severe bandwidth attenuation with distance and topology rigidity, making it difficult to meet the low-latency, high-bandwidth communication requirements for gradient synchronization and scientific computing in artificial intelligence training.
A central optoelectronic hybrid switching chip based on the interconnection of electrical and optical switching matrices is adopted. Control flow and data flow are transmitted through optical and electrical links respectively, realizing direct interconnection between graphics processors. The low latency of the electrical link and the high bandwidth of the optical link are utilized to achieve physical separation transmission.
It reduces communication latency, improves the efficiency and performance of multi-GPU collaborative computing, and meets the real-time requirements of control flow and the large data volume transmission requirements of data flow.
Smart Images

Figure CN120997027B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of high-performance computing and artificial intelligence acceleration, in particular to a graphics processor intercommunication method, product, device and medium. BACKGROUND
[0002] The current single-machine system multi-graphics processor (Graphics Processing Unit, GPU) intercommunication has problems such as serious bandwidth attenuation with distance and topology rigidity, which is difficult to meet the low-latency high-bandwidth communication needs of gradient synchronization in AI (Artificial Intelligence) training and scientific computing. In the NVLink interconnection scheme, the NVLink mesh topology is NVIDIA DGX A100 and connects multiple graphics processors using NVSwitch to form a fully connected network. Limited by the wiring density of the PCB (Printed Circuit Board), the single-hop communication distance needs to be <10 cm, otherwise the bandwidth will be severely attenuated. In the PCIe (Peripheral Component Interconnect Express) interconnection scheme, the PCIe tree topology needs to pass through the central processor (Central Processing Unit, CPU) for non-blocking communication, which seriously affects the communication efficiency and shared bus bandwidth.
[0003] It can be seen that how to optimize the interconnection between graphics processors to improve the communication efficiency between graphics processors is a problem to be solved by those skilled in the art. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a graphics processor intercommunication method, device, equipment and medium, which optimizes the interconnection between graphics processors to improve the communication efficiency between graphics processors. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a graphics processor intercommunication method, the single-machine system includes a plurality of graphics processors and a central optoelectronic hybrid switching chip constructed based on an electrical switching matrix and an optical switching matrix, each graphics processor is connected with the central optoelectronic hybrid switching chip through an optical link and an electrical link to realize the interconnection between each graphics processor; the method comprises:
[0006] determining a source graphics processor from each graphics processor;
[0007] classifying the to-be-transmitted data of the source graphics processor to determine the data type of the to-be-transmitted data;
[0008] if the data type of the data to be transmitted is a control flow type, then controlling the source GPU to send the data to be transmitted to the electrical switch matrix through the electrical link, and controlling the electrical switch matrix to route the data to be transmitted to a destination GPU in each of the GPUs;
[0009] if the data type of the data to be transmitted is a data flow type, then controlling the source GPU to send the data to be transmitted to the optical switch matrix through the optical link, and controlling the optical switch matrix to route the data to be transmitted to a destination GPU in each of the GPUs.
[0010] Optionally, the data classification of the data to be transmitted by the source GPU comprises:
[0011] extracting a data feature of the data to be transmitted by the source GPU, and determining whether the data feature satisfies a preset small data condition by using a hardware-level traffic classifier; wherein the preset small data condition is that a data volume of the data to be transmitted is less than a preset data volume threshold and a delay requirement of the data to be transmitted is less than a preset requirement threshold;
[0012] if the data feature satisfies the preset small data condition, then determining that the data type of the data to be transmitted is a control flow type;
[0013] if the data feature does not satisfy the preset small data condition, then determining that the data type of the data to be transmitted is a data flow type.
[0014] Optionally, the extraction of the data feature of the data to be transmitted by the source GPU comprises:
[0015] extracting a data volume of the data to be transmitted by the source GPU by using a data length counter, and determining a delay requirement of the data to be transmitted according to a delay requirement flag in a transmission instruction of the data to be transmitted.
[0016] Optionally, the control of the electrical switch matrix to route the data to be transmitted to a destination GPU in each of the GPUs comprises:
[0017] parsing the data to be transmitted to identify a destination address, and determining the destination GPU corresponding to the destination address from each of the GPUs;
[0018] controlling the electrical switch matrix to route the data to be transmitted to an electrical link port corresponding to the destination address.
[0019] Optionally, the control of the optical switch matrix to route the data to be transmitted to a destination GPU in each of the GPUs comprises:
[0020] performing wavelength identification on the to-be-transmitted data to determine a target wavelength corresponding to the to-be-transmitted data;
[0021] determining a destination graphic processor corresponding to the target wavelength from the graphic processors;
[0022] controlling the optical switch matrix to route the to-be-transmitted data to the destination graphic processor.
[0023] Optionally, the controlling the source graphic processor to send the to-be-transmitted data to the electrical switch matrix through the electrical link comprises:
[0024] controlling the source graphic processor to perform signal processing on the to-be-transmitted data through the electrical link to obtain to-be-transmitted data in the form of electrical signals, and sending the to-be-transmitted data in the form of electrical signals to the electrical switch matrix.
[0025] Optionally, the controlling the electrical switch matrix to route the to-be-transmitted data to the destination graphic processor in the graphic processors comprises:
[0026] controlling the electrical switch matrix to route the to-be-transmitted data to the destination graphic processor in the graphic processors through the electrical link, so that the destination graphic processor generates a control instruction and an acknowledgement instruction based on the to-be-transmitted data, uses the control instruction to regulate a processing mechanism of a current to-be-processed task, and feeds back the acknowledgement instruction to the source graphic processor.
[0027] Optionally, the controlling the electrical switch matrix to route the to-be-transmitted data to the destination graphic processor in the graphic processors comprises:
[0028] determining the destination graphic processor from the graphic processors;
[0029] controlling the electrical switch matrix to route the to-be-transmitted data to a target electrical link corresponding to the destination graphic processor;
[0030] controlling the target electrical link to perform signal processing on the to-be-transmitted data to obtain first target processed to-be-transmitted data, and sending the first target processed to-be-transmitted data to the destination graphic processor.
[0031] Optionally, the electrical link comprises a cross-group amplifier, a clock data recovery circuit, and a PAM4 modulator; and the controlling the target electrical link to perform signal processing on the to-be-transmitted data to obtain first target processed to-be-transmitted data comprises:
[0032] The cross-group amplifier is controlled to perform signal amplification processing on the to-be-transmitted data, to obtain first signal-processed to-be-transmitted data, and the clock data recovery circuit is used to eliminate signal jitter noise in the first signal-processed to-be-transmitted data, to obtain second signal-processed to-be-transmitted data.
[0033] The PAM4 modulator is controlled to restore the second signal-processed to-be-transmitted data to binary data, to obtain first target-processed to-be-transmitted data.
[0034] Optionally, the control of the source graphics processor to send the to-be-transmitted data to the optical switching matrix through the optical link includes:
[0035] The silicon light emitting module in the source graphics processor is controlled to convert the to-be-transmitted data into optical signal form to-be-transmitted data, and the optical signal form to-be-transmitted data is sent to the optical switching matrix through the optical link.
[0036] Optionally, the control of the optical switching matrix to route the to-be-transmitted data to a destination graphics processor in each graphics processor includes:
[0037] A destination graphics processor is determined from each graphics processor.
[0038] The optical switching matrix is controlled to route the to-be-transmitted data to a target optical link corresponding to the destination graphics processor.
[0039] The to-be-transmitted data is routed to the destination graphics processor through the target optical link.
[0040] Optionally, the optical link includes a microring resonator array; and the routing of the to-be-transmitted data to the destination graphics processor through the target optical link includes:
[0041] The microring resonator array is controlled to route the to-be-transmitted data to a wavelength channel of the destination graphics processor, so that a germanium-silicon detector of the destination graphics processor demodulates the to-be-transmitted data into an electrical signal form to obtain third signal-processed to-be-transmitted data, and restores the third signal-processed to-be-transmitted data to parallel data to obtain second target-processed to-be-transmitted data.
[0042] Optionally, after the routing of the to-be-transmitted data to the destination graphics processor through the target optical link, the method further includes:
[0043] The destination graphics processor is controlled to perform integrity verification on the to-be-transmitted data to obtain a verification result, and the verification result is fed back to the source graphics processor; wherein the verification result is a cyclic redundancy check result or a hash comparison result.
[0044] Optionally, the electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before the control of the micro-ring resonator array to route the data to be transmitted to a wavelength channel of the destination graphic processor, the method further comprises:
[0045] allocating different wavelength channels to optical ports of each graphic processor.
[0046] Optionally, the electrical switching matrix is constructed based on a non-blocking crossbar structure, and the internal of the electrical switching matrix is connected with each first input port and each first output port through a metal interconnection line; the optical switching matrix is constructed based on a silicon-based photonic integrated mechanism, and the internal of the optical switching matrix is connected with each second input port and each second output port through an optical waveguide; the electrical switching matrix and the optical switching matrix are connected vertically through a through-silicon via.
[0047] Optionally, the communication hop number between any two graphic processors is not greater than 2, and the electrical link between any graphic processor and the central optoelectronic hybrid switching chip is a double link, and the optical link between any graphic processor and the central optoelectronic hybrid switching chip is a single link.
[0048] In a second aspect, the present application discloses a graphic processor intercommunication device applied to a single computer system, wherein the single computer system comprises a plurality of graphic processors and a central optoelectronic hybrid switching chip constructed based on an electrical switching matrix and an optical switching matrix interconnection; each graphic processor is connected with the central optoelectronic hybrid switching chip through an optical link and an electrical link to realize the interconnection between the graphic processors; the device comprises:
[0049] a determination module configured to determine a source graphic processor from the graphic processors;
[0050] a classification module configured to classify the data to be transmitted of the source graphic processor to determine a data type of the data to be transmitted;
[0051] a first communication module configured to, if the data type of the data to be transmitted is a control flow type, control the source graphic processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to a destination graphic processor among the graphic processors;
[0052] a second communication module configured to, if the data type of the data to be transmitted is a data flow type, control the source graphic processor to send the data to be transmitted to the optical switching matrix through the optical link, and control the optical switching matrix to route the data to be transmitted to a destination graphic processor among the graphic processors.
[0053] In a third aspect, the present application discloses an electronic device, comprising:
[0054] a memory for storing a computer program;
[0055] a processor for executing the computer program to implement the steps of the above disclosed inter-graphic processor communication method.
[0056] In a fourth aspect, the present application discloses a computer readable storage medium for storing a computer program; wherein the computer program is executed by a processor to implement the steps of the above disclosed inter-graphic processor communication method.
[0057] In a fifth aspect, the present application discloses a computer program product comprising computer programs / instructions, which are executed by a processor to implement the steps of the above disclosed inter-graphic processor communication method.
[0058] Therefore, the present application is applied to a single system, the single system comprises a plurality of graphic processors and a central opto-electric hybrid switching chip constructed based on an electrical switching matrix and an optical switching matrix, each graphic processor is connected with the central opto-electric hybrid switching chip through an optical link and an electrical link to realize the interconnection between the graphic processors; the method comprises: determining a source graphic processor from the graphic processors; classifying the to-be-transmitted data of the source graphic processor to determine the data type of the to-be-transmitted data; if the data type of the to-be-transmitted data is a control flow type, controlling the source graphic processor to send the to-be-transmitted data to the electrical switching matrix through the electrical link, and controlling the electrical switching matrix to route the to-be-transmitted data to a destination graphic processor in the graphic processors; if the data type of the to-be-transmitted data is a data flow type, controlling the source graphic processor to send the to-be-transmitted data to the optical switching matrix through the optical link, and controlling the optical switching matrix to route the to-be-transmitted data to the destination graphic processor in the graphic processors.
[0059] The beneficial effect is that: the application realizes the interconnection of each graphic processor by connecting the multiple graphic processors in the single machine system with the central optoelectronic hybrid switching chip based on the electrical switching matrix and the optical switching matrix through optical links and electrical links, that is, each graphic processor is directly interconnected through the central chip, so that the communication hop count is reduced, and when data is transmitted, the source graphic processor is classified, the control flow type is routed to the destination graphic processor through the electrical switching matrix through the electrical link, and the data flow type is routed to the destination graphic processor through the optical switching matrix through the optical link, so that the physical separation transmission of the control flow and the data flow can be realized, the control flow meets the real-time requirement by means of the low-delay characteristic of the electrical link, the data flow meets the large data transmission requirement by relying on the high-bandwidth characteristic of the optical link, the transmission delay is reduced, and the efficiency and performance of the collaborative calculation of the multiple graphic processors in the single machine system are improved. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0061] Figure 1 A graphic processor intercommunication method flow chart is provided for the embodiments of the present application.
[0062] Figure 2 A specific single machine system schematic diagram is provided for the embodiments of the present application.
[0063] Figure 3 A specific optoelectronic hybrid transmission path schematic diagram is provided for the embodiments of the present application.
[0064] Figure 4 A specific graphic processor and chip connection schematic diagram is provided for the embodiments of the present application.
[0065] Figure 5 A specific electrical connection module schematic diagram is provided for the embodiments of the present application.
[0066] Figure 6 A specific optical connection module schematic diagram is provided for the embodiments of the present application.
[0067] Figure 7 A specific data communication schematic diagram is provided for the embodiments of the present application.
[0068] Figure 8 A graphic processor intercommunication device structure schematic diagram is provided for the embodiments of the present application.
[0069] Figure 9This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0070] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0071] Current single-machine systems suffer from severe bandwidth attenuation with distance and topology rigidity in inter-GPU communication, making it difficult to meet the low-latency, high-bandwidth communication requirements for gradient synchronization and scientific computing in AI training. In current NVLink interconnect solutions, the NVLink mesh topology uses NVIDIA DGX A100 and NVSwitch to connect multiple GPUs, forming a fully connected network. However, due to PCB wiring density limitations, the single-hop communication distance must be less than 10cm; otherwise, bandwidth will be severely attenuated. In PCIe interconnect solutions, the PCIe tree topology requires non-blocking communication via the central processing unit, severely impacting communication efficiency and sharing bus bandwidth.
[0072] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.
[0073] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0074] Next, we will describe in detail a communication scheme between graphics processors provided by an embodiment of the present invention. Figure 1 This invention provides a method for communication between graphics processors (GPUs), applied to a standalone system. The standalone system includes multiple GPUs and a central optoelectronic hybrid switching chip constructed based on the interconnection of an electrical switching matrix and an optical switching matrix. Each GPU is connected to the central optoelectronic hybrid switching chip via an optical link and an electrical link, respectively, to achieve interconnection between the GPUs. The method includes:
[0075] Step S11: Identify the source graphics processor from each of the graphics processors.
[0076] The standalone system uses a hybrid optoelectronic switching architecture, such as... Figure 2As shown, taking 8 GPUs included in a single system as an example, single-machine 8-GPU full interconnection is implemented, and the system is composed of three core modules: a central optoelectronic hybrid switching chip, 8 GPU computing nodes, and an intelligent control platform. Among them, the central optoelectronic hybrid switching chip vertically stacks a 16x16 electrical switching matrix and an 8x8 optical switching matrix through 3D heterogeneous integration technology to form a unified switching plane. Each GPU node is equipped with a dual-mode network interface, which connects the central switching chip through a high-speed electrical link (NVLink compatible) and an optical link (fixed wavelength). The intelligent control platform monitors the system state in real time, dynamically optimizes the topology configuration and resource allocation. The system adopts a star topology structure to ensure that the maximum hop count between any two GPUs does not exceed 2, and through an optoelectronic collaborative transmission strategy, the control flow (electrical transmission) and data flow (optical transmission) are intelligently allocated.
[0077] The source graphics processor is determined from the graphics processors, and the source graphics processor is the current data that needs to be transmitted, that is, the source graphics processor is the source of communication, and the number of source graphics processors can be single or multiple, depending on the specific communication situation.
[0078] In this embodiment, the electrical switching matrix is constructed based on a non-blocking crossbar structure, and the internal electrical switching matrix connects each first input port and each first output port through a metal interconnection line. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal optical switching matrix connects each second input port and each second output port through an optical waveguide. The electrical switching matrix and the optical switching matrix are connected vertically through a through-silicon via.
[0079] The electrical switching matrix is constructed based on a non-blocking crossbar structure, and the internal electrical switching matrix connects each first input port and each first output port through a metal interconnection line. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal optical switching matrix connects each second input port and each second output port through an optical waveguide. The electrical switching matrix and the optical switching matrix are connected vertically through a through-silicon via.
[0080] In the embodiment, the communication hop number between any two of the graphic processors is not greater than 2, the electrical link between any of the graphic processors and the central optoelectronic hybrid switching chip is a double link, and the optical link between any of the graphic processors and the central optoelectronic hybrid switching chip is a single link.
[0081] The communication hop number between any two of the graphic processors is not greater than 2, that is, the data to be transmitted of the source graphic processor can reach the destination graphic processor after being routed by the central optoelectronic hybrid switching chip without multi-stage forwarding, the electrical link between any of the graphic processors and the central optoelectronic hybrid switching chip is a double link (supporting bidirectional communication, unidirectional bandwidth 112 Gbps), which can guarantee bidirectional efficient transmission of the control flow, the optical link between any of the graphic processors and the central optoelectronic hybrid switching chip is a single link (corresponding to a dedicated wavelength, bandwidth 112 Gbps), which adapts to the unidirectional high-bandwidth transmission requirement of the data flow. This design reduces transmission delay by reducing communication hop number, and the configuration of double electrical links and single optical link not only meets the low-delay requirement of bidirectional interaction of the control flow, but also guarantees the high-bandwidth characteristic of unidirectional transmission of the data flow, while improving the link redundancy capability and resource utilization efficiency, and enhancing the stability and overall performance of multi-graphic processor collaborative computing.
[0082] Step S12: classifying the data to be transmitted of the source graphic processor to determine the data type of the data to be transmitted.
[0083] In the embodiment, the classifying the data to be transmitted of the source graphic processor to determine the data type of the data to be transmitted includes: extracting data features of the data to be transmitted of the source graphic processor, and judging whether the data features satisfy a preset small data condition by using a hardware-level traffic classifier; wherein the preset small data condition is that the data amount of the data to be transmitted is less than a preset data amount threshold and the delay requirement of the data to be transmitted is less than a preset requirement threshold; if the data features satisfy the preset small data condition, the data type of the data to be transmitted is determined as a control flow type; and if the data features do not satisfy the preset small data condition, the data type of the data to be transmitted is determined as a data flow type.
[0084] The data characteristics of the data to be transmitted from the source graphics processor are extracted. These characteristics include data size and latency requirements. A hardware-level traffic classifier is used to determine whether the data characteristics meet a preset small data condition. The preset small data condition is that the data size to be transmitted is less than a preset data size threshold and the latency requirement is less than a preset requirement threshold. Specifically, the preset data size threshold can be 4KB, and the preset requirement threshold is 35ns. That is, if the data size is less than 4KB and the latency requirement is less than 35ns, then the preset small data condition is met. If the data characteristics meet the preset small data condition, the data type of the data to be transmitted is determined to be control flow; if the data characteristics do not meet the preset small data condition, the data type of the data to be transmitted is determined to be data stream. This classification method achieves accurate determination of data type through hardware-level real-time processing, ensuring that control flow can be transmitted through electrical links adapted to its low latency requirements, and data stream can be transmitted through optical links adapted to its high bandwidth requirements. This avoids conflicts caused by different types of data being transmitted on the same link. Furthermore, hardware-level classification is fast, adds almost no extra latency, and effectively improves the overall efficiency of data transmission.
[0085] In this embodiment, extracting the data features of the data to be transmitted from the source graphics processor includes: extracting the data volume of the data to be transmitted from the source graphics processor using a data length counter, and determining the delay requirement of the data to be transmitted based on the delay requirement marker in the transmission instruction of the data to be transmitted.
[0086] By utilizing the data length counter integrated into the source graphics processor, the data to be transmitted output by the source graphics processor is counted in real time to directly obtain the specific data volume, such as the number of bytes. Simultaneously, the transmission instructions corresponding to the data to be transmitted are read. By identifying the latency requirement markers carried in the instructions, such as the "low latency marker" in gradient synchronization instructions in AI training and the "high bandwidth marker" in memory block migration instructions, it is determined whether the data to be transmitted needs to meet the preset requirement of end-to-end latency <35ns, thus clarifying its latency requirement. This data feature extraction method relies on the hardware module (i.e., the data length counter) and the markers carried in the instructions, without the need for additional software calculations or data parsing steps. It can quickly and accurately obtain the two core features of data volume and latency requirement, laying the foundation for the subsequent hardware-level traffic classifier to accurately determine the data type. At the same time, it avoids the additional latency caused by software extraction and ensures the overall efficiency of data transmission.
[0087] Furthermore, the preset small data conditions can also include communication between the source and destination GPUs within the same PCB board, meaning the distance between them is less than 10cm. In other words, the preset small data conditions are: the amount of data to be transmitted is less than a preset data amount threshold, the latency requirement is less than a preset requirement threshold, and the source and destination GPUs communicate within the same PCB board. For example, if the data to be transmitted is high-frequency small data (<4KB), such as gradient synchronization signals, barrier synchronization signals, or atomic operation signals, and has low latency requirements, such as real-time control instructions requiring a latency of less than 35ns, and communication is within the same PCB board (GPU spacing <10cm), then the data type of the data to be transmitted is control flow type. Conversely, if the data to be transmitted is large data (>4KB), such as model weight update data, memory copy data, video frame transmission data, or bandwidth-intensive operation data, such as requiring a continuous throughput of >100GB / s, with insensitive latency, or cross-board communication (GPU spacing >10cm, connected via onboard fiber optics), then the data type of the transmitted data is data stream type. Taking Verilog as an example:
[0088] always @(packet_header) begin;
[0089] if (packet.distance <= 10 || packet.latency<35 ||packet.size <4096);
[0090] / / Connection distance less than 10cm or latency requirement less than 35ns or data packet size less than 4k;
[0091] route_to_electrical(); / / Data travels via the electrical link;
[0092] else;
[0093] route_to_optical(); / / Data routing via optical path;
[0094] end.
[0095] Step S13: If the data type of the data to be transmitted is control flow type, then control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor in each of the graphics processors.
[0096] In this embodiment, the control of the source graphics processor to send the to-be-transmitted data to the electrical switching matrix through the electrical link includes: controlling the source graphics processor to perform signal processing on the to-be-transmitted data through the electrical link to obtain to-be-transmitted data in the form of electrical signals, and sending the to-be-transmitted data in the form of electrical signals to the electrical switching matrix.
[0097] The source graphics processor is controlled to activate the PAM4 modulator in the electrical link, encode the to-be-transmitted data into electrical signals in 4-level states, complete signal processing to obtain to-be-transmitted data in the form of electrical signals, and then control the electrical signals to be transmitted through the metal interconnection line of the electrical link to the first input port of the electrical switching matrix, so as to send the to-be-transmitted data in the form of electrical signals to the electrical switching matrix. This signal processing and transmission mode relies on PAM4 modulation technology to improve the bandwidth efficiency of the electrical link, and cooperates with the low-loss characteristics of the metal interconnection line to ensure high signal integrity and low delay (end-to-end < 35 ns) of the control flow in the transmission process, effectively meeting the real-time interaction demand of the control flow and ensuring the rapid transmission of control instructions between multiple GPUs.
[0098] In this embodiment, the control of the electrical switching matrix to route the to-be-transmitted data to the destination graphics processor in each graphics processor includes: analyzing the to-be-transmitted data to identify the destination address, and determining the destination graphics processor corresponding to the destination address from each graphics processor; and controlling the electrical switching matrix to route the to-be-transmitted data to the electrical link port corresponding to the destination address.
[0099] After the electrical switching matrix receives the to-be-transmitted data in the form of electrical signals (control flow type), the electrical switching matrix analyzes the to-be-transmitted data through the built-in address analysis module, extracts the destination address information (such as the destination GPU identifier) contained in the to-be-transmitted data, and determines the corresponding destination graphics processor from multiple graphics processors according to the destination address. Then, the non-blocking crossbar array of the electrical switching matrix is controlled to act, and based on the connection relationship of the metal interconnection line, the to-be-transmitted data is routed from the current input port to the electrical link output port corresponding to the destination address without blocking, so that the data is transmitted to the destination graphics processor through the electrical link, realizing the rapid directional transmission of the control flow, the routing delay is < 10 ns, ensuring that the control instructions can be accurately delivered to the target in a short time, and improving the real-time performance and reliability of the control interaction between multiple GPUs.
[0100] In the embodiment, the control of the electrical switching matrix to route the to-be-transmitted data to the destination graphics processor in each graphics processor comprises: control of the electrical switching matrix to route the to-be-transmitted data to the destination graphics processor in each graphics processor through the electrical link, so that the destination graphics processor generates control instructions and confirmation instructions based on the to-be-transmitted data, uses the control instructions to regulate the processing mechanism of the current to-be-processed task, and feeds back the confirmation instructions to the source graphics processor.
[0101] After the electrical switching matrix receives the to-be-transmitted data in the form of electrical signals (control flow type), the electrical switching matrix analyzes the to-be-transmitted data through the built-in address resolution module, extracts the destination address information such as the destination GPU identifier contained in the to-be-transmitted data, determines the corresponding destination graphics processor from the plurality of graphics processors according to the destination address, and then controls the non-blocking crossbar array of the electrical switching matrix to route the to-be-transmitted data from the current input port to the electrical link output port corresponding to the destination address based on the connection relationship of the metal interconnection line, so that the data is transmitted to the destination graphics processor through the electrical link. This routing method relies on address resolution and non-blocking crossbar design to realize fast directional transmission of control flow, with a routing delay of <10 ns, ensuring that the control instructions can be accurately delivered to the target in a short time, and improving the real-time performance and reliability of control interaction between multiple GPUs.
[0102] In the embodiment, the control of the electrical switching matrix to route the to-be-transmitted data to the destination graphics processor in each graphics processor comprises: determination of the destination graphics processor from each graphics processor; control of the electrical switching matrix to route the to-be-transmitted data to the target electrical link corresponding to the destination graphics processor; control of the target electrical link to perform signal processing on the to-be-transmitted data to obtain first target processed to-be-transmitted data, and transmission of the first target processed to-be-transmitted data to the destination graphics processor.
[0103] According to the destination address resolved from the to-be-transmitted data, the corresponding destination graphics processor is determined from each graphics processor, the non-blocking crossbar structure of the electrical switching matrix is controlled to route the to-be-transmitted data (control flow in the form of electrical signals) from the input port to the target electrical link corresponding to the destination graphics processor based on the connection relationship of the metal interconnection line, and then the target electrical link is controlled to perform amplification and clock recovery processing on the electrical signal to obtain first target processed to-be-transmitted data with optimized signal integrity, and transmit the first target processed to-be-transmitted data to the destination graphics processor through the electrical link.
[0104] In this embodiment, the electrical link includes a transimpedance amplifier, a clock data recovery circuit, and a PAM4 modulator; and the control of the target electrical link to perform signal processing on the to-be-transmitted data to obtain first target processed to-be-transmitted data includes: control of the transimpedance amplifier to perform signal amplification processing on the to-be-transmitted data to obtain first signal processed to-be-transmitted data, elimination of signal jitter noise in the first signal processed to-be-transmitted data by the clock data recovery circuit to obtain second signal processed to-be-transmitted data; and control of the PAM4 modulator to restore the second signal processed to-be-transmitted data to binary data to obtain the first target processed to-be-transmitted data.
[0105] The electrical link includes a transimpedance amplifier (TIA), a clock and data recovery circuit (CDR), and a PAM4 modulator; the transimpedance amplifier is controlled to perform signal amplification on to-be-transmitted data (control flow in the form of electrical signals) routed to the target electrical link by the electrical switching matrix, to compensate for signal attenuation in the transmission process, to obtain first signal processed to-be-transmitted data; the clock and data recovery circuit is controlled to perform clock synchronization and jitter noise elimination on the first signal processed to-be-transmitted data, to correct timing deviation in signal transmission, to obtain second signal processed to-be-transmitted data; and the PAM4 modulator is controlled to demodulate the processed 4-level signal to binary data to recover the original control flow information, to obtain the first target processed to-be-transmitted data, effectively ensuring the signal integrity and accuracy of the control flow in the electrical link transmission.
[0106] Step S14: If the data type of the to-be-transmitted data is a data flow type, the source graphic processor is controlled to send the to-be-transmitted data to the optical switching matrix through the optical link, and the optical switching matrix is controlled to route the to-be-transmitted data to a destination graphic processor in each graphic processor.
[0107] In this embodiment, the control of the source graphic processor to send the to-be-transmitted data to the optical switching matrix through the optical link includes: control of a silicon light emitting module in the source graphic processor to convert the to-be-transmitted data into optical signal form to-be-transmitted data, and send the optical signal form to-be-transmitted data to the optical switching matrix through the optical link.
[0108] The silicon light emitting module in the control source graphics processor bound with the GPU exclusive wavelength (such as the corresponding wavelength in λ1-λ8) converts the electrical signal of the data to be transmitted (data stream type) into an optical signal of the corresponding wavelength, forms the data to be transmitted in the form of an optical signal, and then controls the optical signal to be transmitted through the optical waveguide of the optical link to the second input port of the optical switching matrix of the central optoelectronic hybrid switching chip, sends the data to be transmitted in the form of an optical signal to the optical switching matrix, realizes efficient conversion of electrical-optical signals relying on the silicon light emitting module, and cooperates with the exclusive wavelength optical waveguide transmission to ensure high-throughput transmission of the data stream in the optical link at a total bandwidth of 896 Gbps, and the optical signal transmission loss is low, meeting the bandwidth demand of large data transmission and improving the efficiency of large-scale data interaction between multiple GPUs.
[0109] In this embodiment, the control of the optical switching matrix to route the data to be transmitted to the destination graphics processor in each graphics processor includes wavelength identification of the data to be transmitted to determine the target wavelength corresponding to the data to be transmitted, determination of the destination graphics processor corresponding to the target wavelength from each graphics processor, and control of the optical switching matrix to route the data to be transmitted to the destination graphics processor.
[0110] After the optical switching matrix receives the data to be transmitted in the form of an optical signal, it identifies the wavelength of the data to be transmitted through a wavelength detection unit to determine the target wavelength (such as a specific wavelength in λ1-λ8) corresponding to the optical signal, determines the destination graphics processor corresponding to the target wavelength from each graphics processor according to the preset wavelength-graphics processor mapping relationship (each wavelength uniquely corresponds to one graphics processor), and then controls the non-blocking optical crossbar composed of a micro-ring resonator array in the optical switching matrix to route the optical signal of the wavelength from the current input optical waveguide to the output optical waveguide corresponding to the destination graphics processor by adjusting the resonance state of the corresponding micro-ring, thereby realizing routing of the data to be transmitted to the destination graphics processor. This wavelength-identified routing method relies on the fast response characteristics of the micro-ring resonator array to realize direct routing of the data stream in the optical domain without the need for optical-electric conversion, and the total bandwidth of the parallel transmission of the eight exclusive wavelengths reaches 896 Gbps, meeting the high-bandwidth demand of large data transmission and improving the efficiency and real-time performance of large-scale data interaction between multiple GPUs.
[0111] In this embodiment, the control of the optical switching matrix to route the data to be transmitted to the destination graphics processor in each graphics processor includes determination of the destination graphics processor from each graphics processor, control of the optical switching matrix to route the data to be transmitted to the target optical link corresponding to the destination graphics processor, and routing of the data to be transmitted to the destination graphics processor through the target optical link.
[0112] According to the destination identifier of the to-be-transmitted data, the corresponding destination graphic processor is determined from the graphic processors, the micro-ring resonator array in the optical switching matrix is controlled to adjust the resonant frequency of a specific micro-ring, the to-be-transmitted data (i.e. data stream in the form of optical signals) is switched from the current input optical waveguide to the target optical link bound with the destination graphic processor, and then the to-be-transmitted data in the form of optical signals is directly transmitted to the optical receiving port of the destination graphic processor through the optical waveguide of the target optical link, so that data routing is realized. Due to the optical domain direct switching characteristic of the micro-ring resonator, no intermediate optoelectronic conversion link is needed, and the exclusive wavelength design of the target optical link guarantees high-bandwidth transmission of the data stream, thereby meeting the efficient transmission demand of large data volume.
[0113] In the embodiment, the optical link includes a micro-ring resonator array; and the routing of the to-be-transmitted data to the destination graphic processor through the target optical link includes: controlling the micro-ring resonator array to route the to-be-transmitted data to a wavelength channel of the destination graphic processor, so that a germanium-silicon detector of the destination graphic processor demodulates the to-be-transmitted data into an electrical signal form to obtain third signal-processed to-be-transmitted data, and restores the third signal-processed to-be-transmitted data into parallel data to obtain second target-processed to-be-transmitted data.
[0114] The micro-ring resonator array in the target optical link is controlled to adjust the bias voltage of the corresponding micro-ring, so that it resonates with the optical signal wavelength of the to-be-transmitted data, the optical signal is accurately routed from the output end of the optical switching matrix to the wavelength channel exclusive to the destination graphic processor, so that the germanium-silicon detector of the destination graphic processor receives the optical signal and demodulates it into an electrical signal form to obtain third signal-processed to-be-transmitted data, and then the electrical signal is restored into parallel data through a serial-parallel conversion circuit to obtain second target-processed to-be-transmitted data, thereby realizing low-loss transmission and reducing the additional overhead of signal conversion and transmission.
[0115] In the embodiment, after the routing of the to-be-transmitted data to the destination graphic processor through the target optical link, the destination graphic processor is further controlled to perform integrity verification on the to-be-transmitted data to obtain a verification result, and the verification result is fed back to the source graphic processor; wherein the verification result is a cyclic redundancy check result or a hash comparison result.
[0116] The control destination graphic processor starts a built-in verification module to perform integrity verification on the received second target processed data to be transmitted, specifically by calculating a cyclic redundancy check value or a hash value of the data, and comparing the check value or the hash value with the check information attached in the data by the source graphic processor in advance to obtain a check result, the check result being pass or fail, and then the control destination graphic processor feeds back the check result to the source graphic processor through the electrical link; wherein the check result is a cyclic redundancy check result or a hash comparison result, this verification mechanism can timely find data errors caused by signal attenuation, crosstalk and the like in the optical link transmission process, ensure the accuracy of large-scale data flow transmission, and feed back the check result through a low-delay electrical link, so that the source graphic processor can quickly respond, ensure the reliability of data transmission, avoid the influence of error accumulation on the correctness of multi-GPU collaborative calculation, and improve the overall stability of the single-machine system.
[0117] In the embodiment, the electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before the control of the micro-ring resonator array routing the data to be transmitted to the wavelength channel of the destination graphic processor, the method further comprises: allocating different wavelength channels to the optical ports of each graphic processor.
[0118] The electrical link is a high-speed differential signal line or an electrical connection line, used for low-delay transmission of control flow, and the optical link is an on-board optical fiber or a silicon-based optical waveguide, suitable for high-bandwidth transmission requirements of data flow; before the control of the micro-ring resonator array routing the data to be transmitted to the wavelength channel of the destination graphic processor, different wavelength channels are allocated to the optical ports of each graphic processor based on wavelength division multiplexing technology, such as 8 GPUs corresponding to λ1-λ8, with a wavelength interval of 100GHz; this link type selection and wavelength channel allocation design makes the electrical link and the optical link adapt to the transmission characteristics of control flow and data flow respectively, and avoids crosstalk of data transmission of different GPUs in the optical link through exclusive wavelength channels, and realizes conflict-free transmission of 896Gbps total bandwidth with the low-loss characteristics of on-board optical fiber or silicon-based optical waveguide, improving the efficiency and reliability of data interaction between multiple GPUs.
[0119] Therefore, the application is applied to a single machine system, the single machine system includes a plurality of graphic processors and a central optoelectronic hybrid switching chip constructed based on an electrical switching matrix and an optical switching matrix interconnection, each graphic processor is connected with the central optoelectronic hybrid switching chip through an optical link and an electrical link to realize the interconnection between the graphic processors, the method includes: determining a source graphic processor from the graphic processors; classifying the to-be-transmitted data of the source graphic processor to determine the data type of the to-be-transmitted data; if the data type of the to-be-transmitted data is a control flow type, controlling the source graphic processor to send the to-be-transmitted data to the electrical switching matrix through the electrical link, and controlling the electrical switching matrix to route the to-be-transmitted data to a destination graphic processor in the graphic processors; if the data type of the to-be-transmitted data is a data flow type, controlling the source graphic processor to send the to-be-transmitted data to the optical switching matrix through the optical link, and controlling the optical switching matrix to route the to-be-transmitted data to the destination graphic processor in the graphic processors.
[0120] The application has the beneficial effects that: the plurality of graphic processors in the single machine system are connected with the central optoelectronic hybrid switching chip constructed based on the electrical switching matrix and the optical switching matrix through the optical link and the electrical link, the graphic processors are interconnected, that is, the graphic processors are directly interconnected through the central chip, the communication hop count is reduced, the to-be-transmitted data of the source graphic processor is classified when data transmission, the control flow type is routed to the destination graphic processor through the electrical switching matrix through the electrical link, the data flow type is routed to the destination graphic processor through the optical switching matrix through the optical link, the physical separation transmission of the control flow and the data flow is realized, the control flow meets the real-time requirement by virtue of the low-delay characteristic of the electrical link, the data flow meets the large data transmission requirement by virtue of the high-bandwidth characteristic of the optical link, the transmission delay is reduced, and the efficiency and performance of the collaborative calculation of the plurality of graphic processors in the single machine system are improved.
[0121] The communication between the graphic processors of the application will be described below. Figure 3 As shown in a specific optoelectronic hybrid transmission path diagram, Figure 3 The GPU0, the GPU1, the GPU4 and the GPU7 are connected with the central optoelectronic hybrid switching chip through the electrical link and the optical link. Figure 4As shown, each GPU node features a highly integrated dual-mode interconnect interface chip, integrating electrical control and optical data channels through 3D packaging. The electrical channel supports the 112Gbps NVLink protocol and employs adaptive equalization technology to ensure signal integrity. The optical channel integrates a tunable laser and a germanium-silicon detector, supporting parallel transmission of eight fixed wavelengths (λ1-λ8), ultimately connecting to a central optoelectronic hybrid switching chip. The intelligent offloading controller performs hardware-level packet analysis. Taking an 8-GPU system as an example, a performance comparison was conducted, and the results are shown in Table 1 below.
[0122] Table 1 Performance Comparison Results
[0123]
[0124] In this embodiment, the electrical connection module adopts an improved NVLink mechanism, using PAM4 modulation, with a single-channel rate of 112Gbps. Figure 5 As shown, the electrical connection module adopts a layered architecture, integrating a protocol engine, hardware offloading unit, redundant controller, SerDes array, adaptive equalization module, PCB interface, and spare traces, and connecting to the central optoelectronic hybrid switching chip via main and spare traces. Table 2 shows a comparison of the transmission distance between the improved NVLink mechanism and the original NVLink mechanism.
[0125] Table 2 Comparison of transmission distances
[0126]
[0127] In this embodiment, the optical connectivity module adopts a silicon photonics integration mechanism, overcoming the limitations of optical module size and energy efficiency, and providing a high-bandwidth, low-latency data channel for full 8-GPU interconnection, such as... Figure 6 As shown, the optical connection module includes a laser array, a micro-ring modulator, a germanium-silicon detector, a wavelength division multiplexer, an optical fiber coupler, a clock recovery circuit, and a TIA array. Figure 7As shown, the photoelectric conversion system module can communicate the data to be transmitted to the corresponding graphics processor according to the electrical signal flow and the optical signal flow; in order to verify the communication performance advantage of the optical-electric hybrid interconnection mechanism based on the application, based on the 8 GPU single machine full interconnection scene, the key indicators are comprehensively tested, and the pure electric interconnection scheme (NVLink multi-hop, PCIe tree topology) is compared and analyzed. The test environment includes: 1) the hardware platform is 8xNVIDIA H20 GPU, central optical-electric hybrid switching chip, silicon optical module (λ1-λ8@112Gbps); 2) software configuration is CUDA 12.4, custom communication protocol stack, AllReduce benchmark test tool; 3) the comparison scheme is NVIDIA DGX H20 (NVSwitch full connection), PCIe5.0 x16 tree topology; the multi-dimensional performance comparison results are shown in Table 3:
[0128] Table 3 Multi-dimensional performance comparison results
[0129]
[0130] In the test scene of the AI model training, the gradient synchronization time is reduced by 42% in the ResNet-152 distributed training; when performing scientific computing, the CFD simulation communication overhead can be reduced by 58%, and the overall task completion time is shortened by 35%.
[0131] The embodiment can realize the coexistence of ultra-high bandwidth and low delay, through the cooperative work of the electrical plane (NVLink compatible) and the optical plane (fixed wavelength routing), the control flow (<4KB data) adopts electrical transmission to realize 35ns end-to-end delay, and the data flow (>4KB) realizes 896Gbps aggregated bandwidth through 8 wavelength optical channels, solving the problem of serious bandwidth attenuation with distance in the traditional pure electric interconnection scheme. The AllReduce operation bandwidth is improved to 896GB / s (198% higher than the improved NVLink multi-hop scheme), and the end-to-end worst delay is reduced to 55ns (54% lower than the improved NVLink), which significantly accelerates the AI training and scientific computing tasks.
[0132] The embodiment has the advantage of energy efficiency ratio breakthrough optimization, and the silicon-based photonic integrated technology greatly reduces the photoelectric conversion energy consumption, and the optical channel energy efficiency ratio is reduced to 2.8pJ / bit (33% energy saving compared with the traditional electric interconnection), and at the same time, the interconnection distance is shortened through 3D heterogeneous packaging, and the signal transmission loss is reduced.
[0133] The embodiment can enhance the flexibility and scalability of topology, break through the limitation of PCB wiring density, and support GPU spacing in a single machine to 20 cm (100% higher than NVLink extension); the optoelectronic double-plane architecture supports dynamic switching of various topology modes such as full connection, multicast, and ring, and intelligent routing strategies to adapt to different load scenarios.
[0134] Figure 8 A structural schematic diagram of an inter-graphic processor communication device provided by the embodiment of the application is applied to a single machine system, the single machine system includes a plurality of graphic processors and a central optoelectronic hybrid switching chip constructed based on an electrical switching matrix and an optical switching matrix, each graphic processor is connected with the central optoelectronic hybrid switching chip through an optical link and an electrical link to realize the interconnection between the graphic processors; the device includes:
[0135] A determination module 11 is configured to determine a source graphic processor from the graphic processors;
[0136] A classification module 12 is configured to classify the to-be-transmitted data of the source graphic processor to determine the data type of the to-be-transmitted data;
[0137] A first communication module 13 is configured to, if the data type of the to-be-transmitted data is a control flow type, control the source graphic processor to send the to-be-transmitted data to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the to-be-transmitted data to a destination graphic processor in the graphic processors;
[0138] A second communication module 14 is configured to, if the data type of the to-be-transmitted data is a data flow type, control the source graphic processor to send the to-be-transmitted data to the optical switching matrix through the optical link, and control the optical switching matrix to route the to-be-transmitted data to a destination graphic processor in the graphic processors.
[0139] Further, the embodiment of the application further discloses an electronic device, Figure 9 is a structural diagram of an electronic device according to an exemplary embodiment, the content in the figure cannot be considered as any limitation on the use range of the application. The electronic device can specifically include at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, the computer program is loaded and executed by the processor 21 to realize the related steps in the inter-graphic processor communication method disclosed in any of the preceding embodiments. In addition, the electronic device in the embodiment can be an electronic computer.
[0140] In this embodiment, the power supply 23 is configured to provide operating voltage for each hardware device on the electronic device; the communication interface 24 is configured to create a data transmission channel between the electronic device and external devices, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which will not be specifically limited herein; the input / output interface 25 is configured to obtain external input data or output data to the outside, and the specific interface type can be selected according to the specific application needs, which will not be specifically limited herein.
[0141] In addition, the memory 22 as a carrier of resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage mode can be temporary storage or permanent storage.
[0142] The operating system 221 is configured to manage and control each hardware device on the electronic device and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the inter-graphics processor communication method executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 can further include a computer program capable of completing other specific work.
[0143] Further, the present application also discloses a computer readable storage medium for storing a computer program; wherein the computer program is executed by a processor to implement the inter-graphics processor communication method disclosed above. For the specific steps of the method, please refer to the corresponding content disclosed in the foregoing embodiments, which will not be repeated here.
[0144] Further, the present application embodiment also discloses a computer program product, including computer programs / instructions, which are executed by a processor to implement the steps of the inter-graphics processor communication method disclosed in any of the foregoing embodiments.
[0145] In the present specification, each embodiment is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. For the same or similar parts between each embodiment, please refer to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant part is described in the method part.
[0146] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functionality, which has been described generally and symbolically in flow charts. Having thus described the functionality of the examples, a person of ordinary skill in the art will be able to implement such functionality in hardware and / or software, and will recognize that the bounds of the examples are not limited by one approach or the other. The various examples can be realized in a centralized fashion in one computer system or network, or in a distributed fashion where different elements are spread across several computer systems or sub-networks. Any kind of computer system or other apparatus adapted for carrying out the methods described herein is suited. A typical combination of hardware and software could be a general purpose computer system with a computer program that, when being loaded and executed, carries out the methods described herein.
[0147] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, hard disk can be used as a storage medium.
[0148] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and do not imply singular or plural. Moreover, the terms "include", "have", or any other variant thereof are intended to encompass non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise a set of elements not expressly listed are not excluded from the scope of the processes, methods, articles, or apparatuses. In addition, the term "comprise" or "comprises" does not exclude the presence of additional elements or steps.
[0149] The above detailed description of the technical solutions provided by the present application has been described in detail, and the principles and implementation modes of the present application have been described in the text. The above description of the examples is only intended to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in view of the above, the content of the description should not be understood as limiting the present application.
Claims
1. A method for communication between graphics processors, characterized in that, The method is applied to a standalone system, which includes multiple graphics processors (GPUs) and a central optoelectronic hybrid switching chip built based on the interconnection of electrical and optical switching matrices. Each GPU is connected to the central optoelectronic hybrid switching chip via optical and electrical links to achieve interconnection between the GPUs. The source graphics processor is determined from each of the graphics processors; The data to be transmitted from the source graphics processor is classified to determine the data type of the data to be transmitted; If the data type of the data to be transmitted is control flow type, then the source graphics processor is controlled to send the data to be transmitted to the electrical switching matrix through the electrical link, and the electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors; If the data type of the data to be transmitted is a data stream, then the source graphics processor is controlled to send the data to be transmitted to the optical switching matrix through the optical link, and the optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors; The electrical switching matrix is constructed based on a non-blocking cross switch structure, and the internal components of the electrical switching matrix are connected to each first input port and each first output port via metal interconnects. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal components of the optical switching matrix are connected to each second input port and each second output port via optical waveguides. The electrical switching matrix and the optical switching matrix are vertically connected via through-silicon vias.
2. The inter-graphics processor communication method according to claim 1, characterized in that, The step of classifying the data to be transmitted from the source graphics processor to determine the data type includes: Extract the data features of the data to be transmitted from the source graphics processor, and use a hardware-level traffic classifier to determine whether the data features meet a preset small data condition; wherein, the preset small data condition is that the data volume of the data to be transmitted is less than a preset data volume threshold and the latency requirement of the data to be transmitted is less than a preset requirement threshold. If the data characteristics satisfy the preset small data condition, then the data type of the data to be transmitted is determined to be control flow type; If the data characteristics do not meet the preset small data conditions, then the data type of the data to be transmitted is determined to be a data stream type.
3. The inter-graphics processor communication method according to claim 2, characterized in that, The extraction of data features from the source graphics processor's data to be transmitted includes: The amount of data to be transmitted from the source graphics processor is extracted using a data length counter, and the delay requirement of the data to be transmitted is determined according to the delay requirement mark in the transmission instruction of the data to be transmitted.
4. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The data to be transmitted is parsed to identify the destination address, and the destination graphics processor corresponding to the destination address is determined from each of the graphics processors. The electrical switching matrix is controlled to route the data to be transmitted to the electrical link port corresponding to the destination address.
5. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the optical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: Wavelength identification is performed on the data to be transmitted to determine the target wavelength corresponding to the data to be transmitted. Determine the target graphics processor corresponding to the target wavelength from among the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the destination graphics processor.
6. The inter-graphics processor communication method according to claim 1, characterized in that, The step of controlling the source graphics processor to send the data to be transmitted to the electrical switching matrix via the electrical link includes: The source graphics processor is controlled to process the data to be transmitted through the electrical link to obtain the data to be transmitted in the form of an electrical signal, and then the data to be transmitted in the form of an electrical signal is sent to the electrical switching matrix.
7. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The electrical switching matrix is controlled to route the data to be transmitted to the destination graphics processor in each of the graphics processors through the electrical link, so that the destination graphics processor generates control instructions and confirmation instructions based on the data to be transmitted, uses the control instructions to regulate the processing mechanism of the current task to be processed, and feeds back the confirmation instructions to the source graphics processor.
8. The inter-graphics processor communication method according to claim 7, characterized in that, The control of the electrical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors via the electrical link includes: The target graphics processor is determined from each of the graphics processors; The electrical switching matrix is controlled to route the data to be transmitted to the target electrical link corresponding to the destination graphics processor. The target electrical link is controlled to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted, and the first target-processed data to be transmitted is sent to the target graphics processor.
9. The inter-graphics processor communication method according to claim 8, characterized in that, The electrical link includes a cross-group amplifier, a clock data recovery circuit, and a PAM4 modulator; controlling the target electrical link to perform signal processing on the data to be transmitted to obtain the first target-processed data to be transmitted includes: The cross-group amplifier is controlled to amplify the data to be transmitted to obtain the first signal-processed data to be transmitted. The clock data recovery circuit is used to eliminate the signal jitter noise in the first signal-processed data to be transmitted to obtain the second signal-processed data to be transmitted. The PAM4 modulator is controlled to restore the data to be transmitted after the second signal processing to binary data, so as to obtain the data to be transmitted after the first target processing.
10. The inter-graphics processor communication method according to claim 1, characterized in that, The step of controlling the source graphics processor to send the data to be transmitted to the optical switching matrix via the optical link includes: The silicon photonics emission module in the source graphics processor is controlled to convert the data to be transmitted into optical signal form, and then send the optical signal form of the data to be transmitted to the optical switching matrix through the optical link.
11. The inter-graphics processor communication method according to claim 1, characterized in that, The control of the optical switching matrix to route the data to be transmitted to the destination graphics processors in each of the graphics processors includes: The target graphics processor is determined from each of the graphics processors; The optical switching matrix is controlled to route the data to be transmitted to the target optical link corresponding to the destination graphics processor. The data to be transmitted is routed to the target graphics processor via the target optical link.
12. The inter-graphics processor communication method according to claim 11, characterized in that, The optical link includes a microring resonator array; routing the data to be transmitted to the destination graphics processor via the target optical link includes: The micro-ring resonator array is controlled to route the data to be transmitted to the wavelength channel of the target graphics processor, so that the germanium-silicon detector of the target graphics processor demodulates the data to be transmitted into an electrical signal form to obtain the data to be transmitted after third signal processing, and restores the data to be transmitted after third signal processing into parallel data to obtain the data to be transmitted after second target processing.
13. The inter-graphics processor communication method according to claim 11, characterized in that, After routing the data to be transmitted to the destination graphics processor via the target optical link, the method further includes: The destination graphics processor is controlled to perform integrity verification on the data to be transmitted, so as to obtain the verification result, and the verification result is fed back to the source graphics processor; wherein, the verification result is a cyclic redundancy check result or a hash comparison result.
14. The inter-graphics processor communication method according to claim 12, characterized in that, The electrical link is a high-speed differential signal line or an electrical connection line, and the optical link is an on-board optical fiber or a silicon-based optical waveguide; before controlling the micro-ring resonator array to route the data to be transmitted to the wavelength channel of the target graphics processor, the method further includes: Different wavelength channels are assigned to the optical ports of each graphics processor.
15. The inter-graphics processor communication method according to claim 1, characterized in that, The communication hop count between any two graphics processors is no greater than 2, the electrical link between any graphics processor and the central optoelectronic hybrid switching chip is a dual link, and the optical link between any graphics processor and the central optoelectronic hybrid switching chip is a single link.
16. A communication device between graphics processors, characterized in that, The device is applied to a standalone system, which includes multiple graphics processors and a central optoelectronic hybrid switching chip built based on the interconnection of electrical and optical switching matrices. Each graphics processor is connected to the central optoelectronic hybrid switching chip via optical and electrical links to achieve interconnection between the graphics processors. The device includes: A determining module is used to determine the source graphics processor from each of the graphics processors; A classification module is used to classify the data to be transmitted from the source graphics processor in order to determine the data type of the data to be transmitted. The first communication module is configured to, if the data type of the data to be transmitted is control flow type, control the source graphics processor to send the data to be transmitted to the electrical switching matrix through the electrical link, and control the electrical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors. The second communication module is used to control the source graphics processor to send the data to be transmitted to the optical switching matrix through the optical link if the data type of the data to be transmitted is a data stream type, and to control the optical switching matrix to route the data to be transmitted to the destination graphics processor among the graphics processors. The electrical switching matrix is constructed based on a non-blocking cross switch structure, and the internal components of the electrical switching matrix are connected to each first input port and each first output port via metal interconnects. The optical switching matrix is constructed based on a silicon-based photonic integration mechanism, and the internal components of the optical switching matrix are connected to each second input port and each second output port via optical waveguides. The electrical switching matrix and the optical switching matrix are vertically connected via through-silicon vias.
17. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the graphics processor communication method as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when executed by a processor, the computer programs implement the steps of the inter-graphics processor communication method as described in any one of claims 1 to 15.
19. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the inter-graphics processor communication method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Photoelectric hybrid switching method and device based on QoS flow classification in data center
CN113472685A
Photoelectric hybrid switching system, transmission method and device, GPU server and medium
CN118826874A