A chip-to-chip interconnect system for deep learning

By introducing an electro-optical communication port and an arrayed waveguide grating router into the chip system, combined with silicon photonic transceivers and 3D stacking technology, the problems of high communication latency and bandwidth limitations in the chip system are solved, enabling more efficient execution of deep learning tasks.

CN117151183BActive Publication Date: 2026-01-06北京秩联科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311122958.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2026-01-06
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

Existing chip systems cannot meet the needs of large-scale deep learning training due to high communication latency and bandwidth limitations caused by electrical interconnect technology in deep learning tasks.

Method used

By combining an electro-optical communication port and an arrayed waveguide grating router with a silicon photonic transceiver, and encapsulating it using 3D stacking technology, optical signal transmission between the chips is achieved, reducing communication latency.

Benefits of technology

It significantly reduces the communication latency of the chip system, improving the data transmission efficiency and system scale of deep learning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117151183B_ABST
    Figure CN117151183B_ABST
Patent Text Reader

Abstract

This invention provides a chip interconnect system for deep learning, wherein each chip is equipped with an electro-optical communication port, and the chips are divided into CPU chips and GPU chips. The system includes: at least one CPU chip for managing the transmission and reception of data related to deep learning tasks and the task execution process, and the CPU chip is equipped with an electro-optical communication port; multiple GPU chips for executing deep learning tasks based on data related to deep learning tasks, and the GPU chips are equipped with electro-optical communication ports; multiple arrayed waveguide grating routers for processing data related to deep learning tasks carried by light waves, wherein the arrayed waveguide grating routers and the chips are packaged on different chip layers using 3D stacking technology; and multiple silicon photonic transceivers, each of which is used for connection and electro-optical signal conversion between the electro-optical communication port of one chip and the corresponding arrayed waveguide grating router.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to neural network processor architecture and design methods, more specifically to the field of hardware acceleration for neural network model computation, and more specifically to a chip interconnect system for deep learning. Background Technology

[0002] Since 2012, the computational demands of the largest deep learning training tasks have doubled every 3-4 months, placing extremely high demands on hardware computing power. Chip-based systems for deep learning integrate multiple accelerator chips into a single package and interconnect them via extremely high-bandwidth links, representing a promising solution for deep learning training systems.

[0003] As chiplet packaging size increases, interconnectivity becomes more challenging. Currently, chiplets rely on electrical interconnects, requiring trade-offs between bandwidth and transmission distance, making it impossible to simultaneously meet the bandwidth and distance requirements of deep learning training. With 20Gbps bandwidth per I / O pin, the transmission distance is typically on the order of hundreds of micrometers, only allowing interconnection between adjacent chiplets. This results in a very large interconnect network radius and communication latency of tens of nanoseconds per hop, severely limiting the overall chiplet system's scalability. Furthermore, deep training demands extremely high bandwidth to handle MB / GB-level data streams during training, with stringent latency requirements. Traditional multi-hop chiplet networks can lead to significant latency issues. Summary of the Invention

[0004] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a chip interconnect system for deep learning.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] According to a first aspect of the present invention, a chip interconnect system for deep learning is provided, wherein each chip is provided with an electro-optical communication port, the chips are divided into CPU chips and GPU chips, the system comprising: at least one CPU chip for managing the transmission and reception of deep learning task-related data and the task execution process, and the CPU chip is provided with an electro-optical communication port for transmitting deep learning task-related data; multiple GPU chips for executing deep learning tasks based on deep learning task-related data, and the GPU chips are provided with electro-optical communication ports for transmitting deep learning task-related data; and multiple silicon photonic transceivers. Each electro-optical communication port is connected to a corresponding silicon photonics transceiver. Each silicon photonics transceiver transmits data to its connected electro-optical communication port via electrical signals. Each silicon photonics transceiver converts the electrical signals from the electro-optical communication port into optical signals and transmits them. Multiple arrayed waveguide grating routers are used to route the optical signals emitted by the transmitting data chip via the silicon photonics transceivers to the silicon photonics transceivers corresponding to the receiving data chip. The receiving data chip's corresponding silicon photonics transceiver demodulates the optical signals into electrical signals and transmits them to the receiving data chip. The arrayed waveguide grating routers and the data chips are packaged on different chip layers using 3D stacking technology.

[0007] Optionally, the system is configured to: provide a silicon via for transmitting data via electrical signals between each silicon photonic transceiver and its connected electro-optical communication port; and provide an optical waveguide for transmitting data via optical signals between each silicon photonic transceiver and its connected arrayed waveguide grating router.

[0008] Optionally, the system is configured to: control the multiple GPU chips, the multiple silicon photonic transceivers, and the multiple arrayed waveguide grating routers by the CPU chip according to the task execution flow corresponding to the deep learning task, so as to complete the data interaction between GPU chips or between GPU chips and CPU chips according to the set communication topology algorithm required in the current process of the task execution flow.

[0009] Optionally, the system is configured such that: the total number of CPU chips N is less than the total number of GPU chips M, where N≥1, M≥2, and the total number of electro-optical communication ports L on a single CPU chip is less than the number of electro-optical communication ports Q on a single GPU chip, L≥1, Q≥3, the total number of arrayed waveguide grating routers is greater than or equal to the number of electro-optical communication ports Q on a single GPU chip; and the number of ports on a single arrayed waveguide grating router is greater than or equal to the total number of GPU chips and CPU chips.

[0010] Optionally, the system is configured such that: the first to L electro-optical communication ports of each CPU chip are respectively connected to the ports of the first to L arrayed waveguide grating routers; the first to Q electro-optical communication ports of each GPU chip are respectively connected to the ports of the first to Q arrayed waveguide grating routers; wherein, the first arrayed waveguide grating router is used for one-to-many or many-to-many aggregated communication between all CPU chips and all GPU chips, and the optical signals transmitted to the first arrayed waveguide grating router are modulated and demodulated by the corresponding silicon photonic transceiver devices in a wavelength division multiplexing manner.

[0011] Optionally, the system is configured such that: the second arrayed waveguide grating router is used for one-to-many or many-to-many aggregated communication among all GPU chips, and the optical signals transmitted to the second arrayed waveguide grating router are modulated and demodulated by the corresponding silicon photonic transceiver in a wavelength division multiplexing manner.

[0012] Optionally, the system is configured such that: the 3rd-Qth arrayed waveguide grating router is used for one-to-many, many-to-many, or one-to-one aggregated communication among all GPU chips, and the optical signals transmitted to the 3rd-Qth arrayed waveguide grating router are modulated and demodulated by the corresponding silicon photonic transceiver device in a wavelength division multiplexing or single-wavelength manner.

[0013] Optionally, the silicon photonic transceiver is configured to modulate and demodulate at least N+M different wavelengths of optical signals, where N represents the total number of CPU chips and M represents the total number of GPU chips.

[0014] Optionally, the system further includes: dedicated memory for storing data of CPU chips or GPU chips, with each CPU chip connected to at least one dedicated memory and each GPU chip connected to at least one dedicated memory, wherein the dedicated memory is high-bandwidth memory.

[0015] According to a second aspect of the present invention, a deep learning method based on the system described in the first aspect is provided. The method includes: acquiring deep learning task-related data, the deep learning task-related data including an initial deep learning model and training data for training the model; the CPU chip distributing the model parameters of the initial deep learning model and the training data to each GPU chip performing training via the CPU chip's electro-optical communication port, silicon photonics transceiver, and arrayed waveguide grating router; each GPU chip executing the deep learning task of the deep learning model and updating the model parameters, and interacting with the model parameters, training data, and intermediate feature data during deep learning via the GPU chip's electro-optical communication port, silicon photonics transceiver, and arrayed waveguide grating router; and after deep learning is completed, each GPU chip transmitting the trained model parameters to the CPU chip via the electro-optical communication port, silicon photonics transceiver, and arrayed waveguide grating router. Attached Figure Description

[0016] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0017] Figure 1 This is a schematic diagram of a chip-to-particle interconnect system for deep learning according to an embodiment of the present invention;

[0018] Figure 2 This is a simplified schematic diagram illustrating the principle of interconnection between chips via multiple arrayed waveguide grating routers according to an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of an all-to-all interconnect according to an embodiment of the present invention;

[0020] Figure 4 This is a schematic diagram of Allreduce interconnection according to an embodiment of the present invention;

[0021] Figure 5 This is a schematic diagram of the communication connection of Allreduce Interconnect according to an embodiment of the present invention;

[0022] Figure 6 This is a schematic diagram illustrating one-to-many communication using wavelength division multiplexing according to an embodiment of the present invention;

[0023] Figure 7 This is a schematic diagram of one-to-one communication according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0025] As mentioned in the background section, current communication between chips is based on electrical interconnection technology. This results in a large interconnection network radius between chips, leading to significant latency in data transmission related to deep learning tasks. To address this, this invention provides electro-optical communication ports for transmitting data related to deep learning tasks in both the CPU and GPU chips. Unlike traditional technologies where data is transmitted via purely electrical signals from the sending chip to the receiving chip, the electro-optical communication port of this invention transmits data via optical signals after being sent as electrical signals. This process is achieved using multiple Arrayed Waveguide Grating Routers (AWGRs) and multiple silicon photonic transceivers, reducing communication latency for deep learning tasks. Furthermore, the AWGRs and the chips are encapsulated on different chip layers using 3D stacking technology, further shortening the communication path and reducing latency for deep learning tasks.

[0026] According to one embodiment of the present invention, see Figure 1 This invention provides a chip interconnect system for deep learning. Each chip has an electro-optical communication port. The chips are divided into CPU chips and GPU chips. The system includes: a CPU chip, a GPU chip, dedicated memory (HBM), an arrayed waveguide grating router (AWGR), and silicon photonic transceivers (Tx / Rx). Each CPU chip has a control unit (CU) for controlling data processing and / or communication control within that CPU chip. Each GPU chip has a control unit (CU) for controlling data processing and / or communication control within that GPU chip.

[0027] To illustrate, when one chip A (CPU chip or GPU chip) needs to communicate with another chip B (CPU chip or GPU chip), it is done in the following way:

[0028] The electro-optical communication port a1 corresponding to the data transmission chip A transmits the data via an electrical signal through a silicon via a through-silicon via to the silicon photonic transceiver a2 corresponding to the electro-optical communication port a1 of the data transmission chip A.

[0029] The data is modulated into an optical signal by silicon photonic transceiver a2, and the optical signal is transmitted through optical waveguide and arrayed waveguide grating router to the silicon photonic transceiver b2 corresponding to the receiving chip B.

[0030] The silicon photonic transceiver b2 demodulates the data from the optical signal and converts it into an electrical signal that is transmitted through a silicon via a through-silicon via to the electro-optical communication port b1 of the receiving chip B.

[0031] According to one embodiment of the present invention, a chip-to-chip interconnect system for deep learning includes:

[0032] N CPU chips are used to manage the sending and receiving of data related to deep learning tasks and the task execution process, and each CPU chip is equipped with L electro-optical communication ports.

[0033] M GPU chips are used to perform deep learning tasks based on data related to deep learning tasks, and each GPU chip is provided with Q electro-optical communication ports.

[0034] N×L+M×Q silicon photonic transceivers, each of which is used for the connection between the electro-optical communication port of a chip and the corresponding arrayed waveguide grating router, as well as the conversion between electrical and optical signals. Each silicon photonic transceiver transmits data to its connected electro-optical communication port via electrical signals, converts electrical signals from the electro-optical communication port into optical signals and transmits them, and converts optical signals from the arrayed waveguide grating router into electrical signals and sends them to its connected electro-optical communication port.

[0035] Q arrayed waveguide grating routers (numbered A1…A1…A2…A3…A4…A5…A6…A7 ... Q It is used to route the optical signal emitted by the data transmitting chip through the silicon photonic transceiver to the silicon photonic transceiver corresponding to the data receiving chip;

[0036] Where N≥1, M≥2, L≥1, Q≥3. This is illustrative; for example, N can be 1, 2, or 3, L can be 1 or 2, M can be 8, 10, or 20, and Q can be 3, 4, or 8.

[0037] According to one embodiment of the present invention, the system is configured such that the total number of CPU cores N is less than the total number of GPU cores M. M can be set according to the system scale; for example, M can be 1 / 8 or 1 / 16 of the total number of GPU cores N. This embodiment achieves at least the following beneficial technical effects: because deep learning requires the use of a large number of GPU cores for feature map computation, its powerful parallel computing capabilities accelerate deep learning tasks.

[0038] According to one embodiment of the present invention, the system is configured such that the total number L of electro-optical communication ports of a single CPU chip is less than the number Q of electro-optical communication ports of a single GPU chip. This embodiment achieves at least the following beneficial technical effects: the CPU chip only needs to perform a large amount of data interaction with the GPU chip at the beginning and end of deep learning, while the majority of data interaction in deep learning tasks occurs between GPU chips. Therefore, the present invention sets more electro-optical communication ports for the GPU chip to improve the data interaction performance between GPU chips during deep learning.

[0039] According to one embodiment of the present invention, the total number of arrayed waveguide grating routers is greater than or equal to the number Q of electrically powered optical communication ports on a single GPU die. The total number of individual arrayed waveguide grating routers should be at least equal to the number Q of electrically powered optical communication ports on a single GPU die. However, more arrayed waveguide grating routers can be configured to meet the needs of other implementers for data transmission in other aspects using additional arrayed waveguide grating routers.

[0040] According to one embodiment of the present invention, the number of ports of a single arrayed waveguide grating router is greater than or equal to the total number of GPU and CPU chips. The number of ports of a single arrayed waveguide grating router being at least equal to the total number of GPU and CPU chips is sufficient to meet the data interaction requirements of deep learning tasks. However, a greater number of ports can also be configured to meet the needs of other implementers for other aspects of data transmission using the additional ports of the arrayed waveguide grating router.

[0041] According to an embodiment of the present invention, the system is configured to: control the plurality of GPU chips, the plurality of silicon photonic transceivers and the plurality of arrayed waveguide grating routers by the CPU chip according to the task execution flow corresponding to the deep learning task, so as to complete the data interaction between GPU chips or between GPU chips and CPU chips according to the set communication topology algorithm required in the current process of the task execution flow.

[0042] According to one embodiment of the present invention, the electro-optical communication port can be configured in two modes, wherein:

[0043] The first type of mode is the fully connected mode, which can be used to accelerate communication between all-to-all and all-gather operators during training. In this mode, the Tx end (transmitter) of the silicon photonics transceiver modulates the signals from different destination chips onto their corresponding wavelengths and multiplexes them into WDM signals. These WDM signals include O different wavelength signals; where O = M + N. When this WDM signal passes through the AWGR, it is switched to different destination chips according to its wavelength routing rules. The Rx end (receiver) demultiplexes the received WDM signal, demodulating the electrical signals of each wavelength to correspond to different source nodes. Illustratively, the CPU chip achieves full connectivity with all CPU and GPU chips through the AWGR, thus enabling the CPU chip to send data down and the GPU to upload data through this full connectivity.

[0044] The second type of mode is the selective connection mode. This mode allows setting the wavelength used for communication between any two cores according to the desired connection topology. This can be a single-wavelength optical signal or multiple wavelengths of optical signal, or it can be matched according to the allreduce algorithm to accelerate the allreduce operator in deep learning tasks. For the connection relationship between the electro-optical communication port and the port of the arrayed waveguide grating router, and the mode of setting the electro-optical communication port, please refer to the following embodiments.

[0045] According to one embodiment of the present invention, the system is configured such that: the first to L electro-optical communication ports of each CPU chip are respectively connected to the ports of the first to L arrayed waveguide grating routers; that is, the L electro-optical communication ports of the CPU chip are respectively connected to AWGR A1…A through silicon photonic waveguides. L One port; the first to Q electro-optical communication ports of each GPU chip are respectively connected to the ports of the first to Q arrayed waveguide grating routers. That is, the Q electro-optical communication ports of each GPU chip are connected to AWGR A1…A through silicon photonic waveguides. Q One of the ports.

[0046] Illustratively, according to one example of the invention, see [example missing]. Figure 2 It provides a schematic diagram showing the connection between two CPU chips (CPU1 and CPU2) and eight GPU chips (GPU1-GPU8) through four arrayed waveguide grating routers (AWGR1-AWGR4) (where silicon photonic transceivers are not shown, but can be imagined as being placed on the connection between the chips and AWGRs), and HBM stands for high bandwidth memory.

[0047] To enable aggregated communication between CPU chips and GPU chips, according to one embodiment of the present invention, a first arrayed waveguide grating router is used for one-to-many or many-to-many aggregated communication between all CPU chips and all GPU chips. The optical signals transmitted to the first arrayed waveguide grating router are modulated and demodulated by corresponding silicon photonic transceivers using wavelength division multiplexing. For the electro-optical communication port connected to the first arrayed waveguide grating router, a fully connected mode can be configured for one-to-many or many-to-many aggregated communication between all CPU chips and all GPU chips. Correspondingly, the wavelength settings for wavelength division multiplexing communication between relevant chips can be found in Table 1.

[0048] Table 1

[0049] w CPU1 … CPUN GPU1 GPU2 … GPUM CPU1 1 … N N+1 N+2 … N+M … … … … … … … … CPUN N … 2N-1 2N 2N+1 … 1 GPU1 N+1 … 2N 2N+1 2N+2 … 2 GPU2 N+2 … 2N+1 2N+2 2N+3 … 3 … … … … … … … … GPUM N+M … 1 2 3 … N+M-1

[0050] illustrative, see Figure 3 Using a 2-CPU and 8-GPU architecture as an example, the all-to-all interconnect mode is used for aggregate communication, such as all-to-all or all-gather communication. For communication between the CPU and GPU, and other communication between GPUs besides all-reduce, all-to-all optical ports are used. For all-to-all and all-gather communication, GPUs communicate using all-to-all optical ports, ensuring one-hop reachability. See again. Figure 2Each CPU chip has an electro-optical communication port for full-connection mode, connected to AWGR1. Each GPU has two electro-optical communication ports for full-connection mode, connected to AWGR1 and AWGR2. AWGR1 and AWGR2 are in full-connection mode, meaning that single-hop communication between any chip can be achieved via AWGR1 and AWGR2. If CPU1 wants to communicate with GPU1 and GPU2, the data stream from CPU1 to GPU1 is modulated into an optical signal of wavelength w3, and the data stream from CPU1 to GPU2 is modulated into an optical signal of wavelength w4. The two signals of different wavelengths are then multiplexed into a WDM signal. This WDM signal includes the data streams sent to different nodes. When passing through AWGR1, the optical signals of wavelength w3 and wavelength w4 are exchanged to GPU1 and GPU2 respectively, realizing data exchange. Based on Table 1, the wavelength settings for wavelength division multiplexing (WDM) communication between 2 CPU chips and 8 GPU chips can be found in Table 2. The values ​​in the middle indicate the wavelength numbers. When the first column represents the chip number transmitting data and the first row represents the chip number receiving data, taking CPU1 to CPU2 as an example, the wavelength for transmitting optical signals between CPU1 and CPU2 is w2. For another example, if CPU1 wants to communicate with GPU1 and GPU2, the CPU1-GPU1 data stream is modulated into an optical signal of wavelength w3, and the CPU1-GPU2 data stream is modulated into an optical signal of wavelength w4. These two signals of different wavelengths are then multiplexed into a WDM signal. This WDM signal includes the data streams sent to different nodes. When passing through AWGR1, the wavelengths w3 and w4 are exchanged to GPU1 and GPU2 respectively, achieving data exchange. Other similar methods are not elaborated upon.

[0051] Table 2

[0052]

[0053]

[0054] To enable aggregated communication between GPU chips, according to one embodiment of the present invention, the system is configured such that: a second arrayed waveguide grating router is used for one-to-many or many-to-many aggregated communication among all GPU chips, and optical signals transmitted to the second arrayed waveguide grating router are modulated and demodulated by corresponding silicon photonic transceivers in a wavelength division multiplexing manner. For the electro-optical communication port connected to the second arrayed waveguide grating router, it can be configured in a fully connected mode for one-to-many or many-to-many aggregated communication among all GPU chips.

[0055] To enable diverse communication through on-demand connections between GPU chips, according to one embodiment of the present invention, the system is configured such that: the 3rd-Qth arrayed waveguide grating router is used for one-to-many, many-to-many, or one-to-one aggregated communication among all GPU chips, and the optical signals transmitted to the 3rd-Qth arrayed waveguide grating router are modulated and demodulated by corresponding silicon photonic transceivers in a wavelength division multiplexing or single-wavelength manner. For the electro-optical communication ports connected to the 3rd-Qth arrayed waveguide grating router, a selective connection mode can be configured for many-to-one, one-to-one, one-to-many, or many-to-many aggregated communication among all GPU chips.

[0056] See you again Figure 2 Each GPU chip also has two electro-optical communication ports for selective connection to AWGR3 and AWGR4. See illustrative example. Figure 4 This diagram illustrates an Allreduce interconnection, using eight GPU chips as an example. The selective connection mode allows for selective connections based on the corresponding ensemble communication operators in the deep learning task. For instance, GPU1 communicates with GPU2, GPU3, and GPU5 via wavelengths w1, w2, and w3, respectively. In selective connection mode, the wavelengths used for communication between GPU chips can be dynamically changed. Selective connection mode supports reconfigurable topologies, which can be used to accelerate many-to-one, one-to-one, one-to-many, or many-to-many ensemble communication during distributed training in deep learning tasks. For example, it can accelerate Allreduce ensemble communication. For Allreduce ensemble communication, this system supports several mainstream algorithms, including the Ring algorithm, Recursive Doubling (RD) algorithm, and Having and Doubling (HD) algorithm. These algorithms are widely used in distributed training. This system can utilize dynamic topology reconstruction to optimize the bandwidth of the entire link. For the Ring algorithm, the topology can be configured as a ring; while for the RD and HD algorithms, a highly dynamic reconstruction method can be used for topology matching.

[0057] like Figure 5 As shown, the RD algorithm has a total of log2X steps (X is the total number of GPU chips). In each step, a chip communicates only with another chip, and the destination node for communication differs between steps. Therefore, wavelength reconstruction can be used to reconstruct the wavelength between each step, thereby reconstructing the direct link matching RD algorithm. Table 3 illustrates the correspondence between the wavelength and routing of the electro-optical communication ports in the system.

[0058] Table 3

[0059] GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7 GPU8 Step 1 w1 w1 w5 w5 w1 w1 w5 w5 Step 2 w2 w4 w2 w4 w2 w4 w2 w4 Step 3 w4 w6 w0 w2 w4 w6 w0 w2

[0060] For example, in selective connection mode, GPU1 communicates with GPU2 in step 1 of the algorithm. At this time, GPU1's electro-optical communication port modulates data onto an optical signal of wavelength w1. This optical signal is a single-wavelength signal, and through AWGR, the w1 wavelength signal is exchanged to GPU2's electro-optical communication port, completing the data exchange. In step 2, GPU1 needs to communicate with GPU3. At this time, the corresponding electro-optical communication port modulates data onto an optical signal of wavelength w2, and the optical signal is exchanged to GPU3's corresponding electro-optical communication port through AWGR. In step 3, GPU1 needs to communicate with GPU5. At this time, the corresponding electro-optical communication port modulates data onto an optical signal of wavelength w4, and the optical signal is exchanged to GPU5's corresponding electro-optical communication port through AWGR.

[0061] According to one embodiment of the present invention, the silicon photonic transceiver is configured to modulate and demodulate at least O different wavelengths of optical signals, where O = N + M, where N represents the total number of CPU chips and M represents the total number of GPU chips. See also Figure 6 Silicon optical transceivers can use microring modulators (MRs) as both optical modulators and demodulators. In the first type of mode, at the data transmission end, the electrical signal is transmitted through the CU to the silicon optical chip layer where the transceiver is located, and is then modulated by O microrings operating at different wavelengths (w1, w2, w3). N+M The MR modulation of the signal is converted into a WDM signal, which is transmitted to the receiving end through the optical waveguide and AWGR. The receiving end receives the WDM signal, demultiplexes it through the MR, and performs photoelectric conversion using a photodetector to demodulate it into an O-channel electrical signal.

[0062] For the second type of pattern, see Figure 7 Assuming one-to-one communication, the electrical signal is transmitted to the silicon photonics chip layer. One of the O MRs (red in the figure) operating at different wavelengths modulates the optical signal. The signal is transmitted to the receiver via an optical waveguide and AWGR. The receiver receives a single-wavelength optical signal and demodulates it into a single-channel electrical signal through the corresponding MR and photodetector.

[0063] The encapsulation of each component in the system can be set up in the following manner.

[0064] According to one embodiment of the present invention, the arrayed waveguide grating router and the chip are packaged on different chip layers using 3D stacking technology. For example, each chip is packaged on a first chip layer, and each arrayed waveguide grating router is packaged on a second chip layer, with the first and second chip layers packaged adjacent to each other. Alternatively, all CPU chips and GPU chips can be packaged on the first chip layer. Or, all CPU chips can be packaged on the first chip layer, the arrayed waveguide grating router on the second chip layer, and all GPU chips on a third chip layer, with the second chip layer packaged between the first and third chip layers. All of these embodiments can achieve at least the following beneficial technical effects: by packaging the arrayed waveguide grating router and the chip on different chip layers, especially adjacent chip layers, using 3D stacking technology, the communication distance between them can be reduced, and the communication latency can be reduced. To illustrate, the system employs a 3D integration scheme, in which all chips and silicon photonic devices (silicon photonic transceivers and AWGRs) are placed on different chip layers. These different chip layers are integrated using 3D stacking packaging technologies, such as those based on TSV (Through Silicon Via) or Micro Bump 3D, thereby reducing the transmission distance of electrical signals and ensuring maximum communication bandwidth.

[0065] According to one embodiment of the present invention, the silicon photonic transceiver and the arrayed waveguide grating router are packaged on the same chip layer. For example, each silicon photonic transceiver is packaged on the aforementioned second chip layer. The technical solution of this embodiment can achieve at least the following beneficial technical effects: packaging each silicon photonic transceiver and the arrayed waveguide grating router on the same chip layer can reduce the layout design and manufacturing difficulty of the optical waveguides between them, making the system performance more reliable.

[0066] According to one embodiment of the present invention, the system is configured to: provide a silicon via for transmitting data via electrical signals between each silicon photonic transceiver and its connected electro-optical communication port; and provide an optical waveguide for transmitting data via optical signals between each silicon photonic transceiver and its connected arrayed waveguide grating router.

[0067] According to one embodiment of the present invention, the system further includes: dedicated memory for storing data of CPU chips or GPU chips, wherein each CPU chip is connected to at least one dedicated memory and each GPU chip is connected to at least one dedicated memory, wherein the dedicated memory is high-bandwidth memory.

[0068] According to one embodiment of the present invention, the present invention also provides a deep learning method based on a chip-particle interconnect system for deep learning, the method comprising: acquiring deep learning task-related data, the deep learning task-related data including an initial deep learning model and training data for training the model; the CPU chip distributing the model parameters of the initial deep learning model and the training data to each GPU chip performing training through the electro-optical communication port, silicon photonic transceiver, and arrayed waveguide grating router of the CPU chip; each GPU chip executing the deep learning task of the deep learning model and updating the model parameters, and interacting with the model parameters, training data, and intermediate feature data during deep learning through the electro-optical communication port, silicon photonic transceiver, and arrayed waveguide grating router of the GPU chip; and after deep learning is completed, each GPU chip transmitting the trained model parameters to the CPU chip through the electro-optical communication port, silicon photonic transceiver, and arrayed waveguide grating router. The technical solution of this embodiment can achieve at least the following beneficial technical effects: Through this method, the communication latency of data interaction between different chips in deep learning tasks can be greatly reduced, chips can be prevented from waiting for data for a long time, the deep learning process can be accelerated, and the efficiency of deep learning tasks can be improved.

[0069] In summary, because the packaged chip is based on a combination of electrical communication and silicon photonics communication and uses wavelength tuning, its speed can reach the nanosecond level. Considering the physical layer link reconstruction overhead and control overhead, the total latency of each reconstruction is at the microsecond or even nanosecond level. In contrast, the transmission time of the load within the reconstruction interval can reach tens or even hundreds of milliseconds. The extremely low reconstruction latency ensures that the reconstruction mechanism will not cause network performance penalties, greatly reducing the data transmission latency of deep learning and improving the efficiency of deep learning.

[0070] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0071] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0072] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0073] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A core particle interconnection system for deep learning, each core particle of the core particle interconnection system is provided with an electro-optical communication port, the core particles are divided into CPU core particles and GPU core particles, the system comprises: at least one CPU core particle for controlling the transmission and reception of deep learning task related data and the task execution process, and the CPU core particle is provided with an electro-optical communication port for transmitting deep learning task related data; a plurality of GPU core particles for executing deep learning tasks according to deep learning task related data, and the GPU core particle is provided with an electro-optical communication port for transmitting deep learning task related data; a plurality of silicon optical transceivers, wherein each electro-optical communication port is connected with a corresponding silicon optical transceiver, each silicon optical transceiver transmits data between the electro-optical communication port connected therewith by electrical signals, and each silicon optical transceiver converts electrical signals from the electro-optical communication port into optical signals and sends them out; a plurality of arrayed waveguide grating routers for routing the optical signals sent by the core particles through the silicon optical transceivers to the silicon optical transceivers corresponding to the core particles receiving data, so as to demodulate the optical signals into electrical signals through the silicon optical transceivers corresponding to the core particles receiving data and transmit them to the core particles receiving data, wherein the arrayed waveguide grating routers and the core particles are packaged in different chip layers by 3D stacking technology; the system is configured to: total number of CPU cores total number of GPU cores wherein , and Total number of electro-optical communication ports of a single CPU die Number of electro-optical communication ports of a single GPU die , , , The total number of arrayed waveguide grating routers is greater than or equal to the number of electrical optical communication ports on a single GPU chip ; the number of ports of a single arrayed waveguide grating router is greater than or equal to the total number of GPU core particles and CPU core particles; the system is further configured to: the 1st to Lth electro-optical communication ports of each CPU core particle are respectively connected with the ports of the 1st to Lth arrayed waveguide grating routers; the 1st to Qth electro-optical communication ports of each GPU core particle are respectively connected with the ports of the 1st to Qth arrayed waveguide grating routers; wherein the 1st arrayed waveguide grating router is used for one-to-many or many-to-many collective communication among all CPU core particles and all GPU core particles, and the optical signals sent to and received from the 1st arrayed waveguide grating router are modulated and demodulated by the corresponding silicon optical transceiver in a wavelength division multiplexing manner; the system is further configured to: the 2nd arrayed waveguide grating router is used for one-to-many or many-to-many collective communication among all GPU core particles, and the optical signals sent to and received from the 2nd arrayed waveguide grating router are modulated and demodulated by the corresponding silicon optical transceiver in a wavelength division multiplexing manner.

2. The system of claim 1, wherein, the system is configured to: a through silicon via is arranged between each silicon optical transceiver and the electro-optical communication port connected therewith for transmitting data by electrical signals; and an optical waveguide is arranged between each silicon optical transceiver and the arrayed waveguide grating router connected therewith for transmitting data by optical signals.

3. The system of claim 2, wherein, the system is configured to: according to the task execution process corresponding to the deep learning task, the CPU core particle controls the plurality of GPU core particles, the plurality of silicon optical transceivers and the plurality of arrayed waveguide grating routers to complete the data interaction between the GPU core particles or between the GPU core particles and the CPU core particles according to the collective communication topology algorithm required by the current process in the task execution process.

4. The system of claim 3, wherein, The system is configured that the third to Qth array waveguide grating routers are used for one-to-many, many-to-many or one-to-one collective communication among all GPU chiplets, and the corresponding silicon optical transceiver devices modulate and send optical signals to the third to Qth array waveguide grating routers and demodulate optical signals from the third to Qth array waveguide grating routers in a wavelength division multiplexing or single wavelength manner.

5. The system of claim 3 or 4, wherein, The silicon optical transceiver device is configured to at least modulate and demodulate A kind of different wavelength optical signal, wherein, , Indicates the total number of CPU chiplets, Indicates the total number of GPU chiplets.

6. The system according to one of claims 1-4, characterized in that, The system further comprises a dedicated memory for storing data of the CPU chiplets or the GPU chiplets, each CPU chiplet being connected to at least one dedicated memory, each GPU chiplet being connected to at least one dedicated memory, and the dedicated memory being a high-bandwidth memory.

7. A deep learning method based on the system of any one of claims 1-6, characterized in that, The method comprises: obtaining deep learning task related data, the deep learning task related data comprising an initial deep learning model and training data for training the model; downloading, by the CPU chiplet, model parameters of the initial deep learning model and the training data to each GPU chiplet performing training through an electro-optical communication port of the CPU chiplet, a silicon optical transceiver and an array waveguide grating router; performing, by each GPU chiplet, a deep learning task of the deep learning model and updating the model parameters, and interacting the model parameters, the training data and intermediate feature data during the deep learning through the electro-optical communication port of the GPU chiplet, the silicon optical transceiver and the array waveguide grating router; after the deep learning is completed, transmitting, by each GPU chiplet, the trained model parameters to the CPU chiplet through the electro-optical communication port, the silicon optical transceiver and the array waveguide grating router. The system is configured that the third to Qth array waveguide grating routers are used for one-to-many, many-to-many or one-to-one collective communication among all GPU chiplets, and the corresponding silicon optical transceiver devices modulate and send optical signals to the third to Qth array waveguide grating routers and demodulate optical signals from the third to Qth array waveguide grating routers in a wavelength division multiplexing or single wavelength manner. The system further comprises a dedicated memory for storing data of the CPU chiplets or the GPU chiplets, each CPU chiplet being connected to at least one dedicated memory, each GPU chiplet being connected to at least one dedicated memory, and the dedicated memory being a high-bandwidth memory. The method comprises: obtaining deep learning task related data, the deep learning task related data comprising an initial deep learning model and training data for training the model; downloading, by the CPU chiplet, model parameters of the initial deep learning model and the training data to each GPU chiplet performing training through an electro-optical communication port of the CPU chiplet, a silicon optical transceiver and an array waveguide grating router; performing, by each GPU chiplet, a deep learning task of the deep learning model and updating the model parameters, and interacting the model parameters, the training data and intermediate feature data during the deep learning through the electro-optical communication port of the GPU chiplet, the silicon optical transceiver and the array waveguide grating router; after the deep learning is completed, transmitting, by each GPU chiplet, the trained model parameters to the CPU chiplet through the electro-optical communication port, the silicon optical transceiver and the array waveguide grating router.

Citation Information

Patent Citations

  • Data center network system and data communication method based on software definition

    CN103441942A

  • Optical calculation device and optical calculation method

    CN116502689A