Topological structure and transmission system based on optical interconnection

By using sparse optical interconnect topology and hybrid connectivity architecture, the problems of insufficient bandwidth of copper interconnect and waste of resources of dense optical interconnect are solved, achieving efficient utilization of GPU computing resources and reduction of hardware costs, and improving model inference speed and parallel efficiency.

CN121815132APending Publication Date: 2026-04-07PHOTON ARITHMETIC (NANJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing copper interconnect technology has insufficient bandwidth in multi-GPU parallel computing, and the number of data transmission links is limited, resulting in high hardware costs, large size, high power consumption, and serious waste of dense optical interconnect resources, which cannot meet the high bandwidth requirements.

Method used

Employing a sparse optical interconnect topology, this architecture directly connects GPUs with the same number via optical fibers between computing nodes. Combined with a hybrid connection architecture of PCIe, QPI, and NIC, it optimizes data transmission paths and reduces communication latency and resource waste.

Benefits of technology

While improving model inference speed, it reduces hardware cost, size and power consumption, maintains efficient resource utilization, and adapts to the improvement of GPU computing power and the parallel characteristics of model inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815132A_ABST
    Figure CN121815132A_ABST
Patent Text Reader

Abstract

The invention provides a topological structure based on optical interconnection and a transmission system. The topological structure based on optical interconnection comprises a plurality of computing nodes, wherein each computing node comprises two central processing units (CPU) and eight graphics processing units (GPU); in one computing node, the two CPUs are electrically connected, and each CPU is electrically connected with the four GPUs respectively; determining one CPU from each computing node as a target CPU, wherein the target CPUs of the plurality of computing nodes are connected through a network; the eight GPUs in each computing node are numbered, and the GPUs, with the same number, of the multiple computing nodes are in optical connection. In the mode, the model reasoning speed can be improved, and meanwhile, the hardware cost, the size and the power consumption are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a topology and transmission system based on optical interconnection. Background Technology

[0002] Currently, existing cluster interconnect technologies are mainly divided into two categories: copper interconnect technology and optical interconnect technology.

[0003] Copper interconnect technology relies on traditional electrical signal transmission. Limited by hardware physical characteristics (such as signal attenuation and transmission rate limits), the number of data transmission links is limited, making it difficult to meet the high bandwidth demands of multi-GPU (Graphics Processing Unit) parallel computing. As GPU computing power continues to improve, with single-card computing power exceeding PFLOPS (1 quadrillion floating-point operations per second), the bandwidth bottleneck of copper interconnects becomes increasingly significant, leading to underutilization of GPU computing resources and increased model inference latency.

[0004] Optical interconnect technology enables communication between nodes via optical fibers, increasing the number of data transmission links in a cluster and overcoming the bandwidth limitations of copper interconnects. However, current optical interconnects often employ dense optical interconnect technologies, requiring a large number of fiber optic transmission ports and optical switches to achieve full connectivity or high redundancy. For example, if a dual-machine 16-card cluster uses all-optical interconnect, each card requires 15 optical ports, and the entire cluster needs multiple high-port-count optical switches. This not only leads to a surge in hardware costs but also increases the cluster's size, power consumption, and heat dissipation pressure, reducing system stability. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a topology and transmission system based on optical interconnects, so as to provide a cluster multi-node GPU collaborative connection topology based on optical interconnects, overcome the shortcomings of the prior art such as insufficient copper interconnect bandwidth, limited number of data transmission links, waste of dense optical interconnect resources, and inability of hardware technology to support it, and improve the model inference speed while reducing hardware cost, size and power consumption.

[0006] In a first aspect, embodiments of the present invention provide a topology based on optical interconnection, the topology based on optical interconnection comprising: multiple computing nodes, each computing node comprising: 2 central processing units (CPUs) and 8 graphics processing units (GPUs); in one computing node, the 2 CPUs are electrically connected to each other, and each CPU is electrically connected to each of the 4 GPUs; one CPU from each computing node is designated as a target CPU, and the target CPUs of the multiple computing nodes are connected via a network; the 8 GPUs in each computing node are numbered, and GPUs with the same number in the multiple computing nodes are connected via optical interconnection.

[0007] In an optional embodiment of this application, the number of computing nodes is determined based on the tensor parallelism used for model inference.

[0008] In an optional embodiment of this application, the CPU is used for resource scheduling and control within the computing node.

[0009] In an optional embodiment of this application, the GPU described above is used to perform model inference computation tasks.

[0010] In an optional embodiment of this application, each GPU includes: 1 electrical interface and N-1 optical interfaces; wherein, N is the number of computing nodes and N is an integer greater than or equal to 2; the electrical interface is used to connect to the CPU in the same computing node; the optical interface is used to connect to GPUs with the same number in different computing nodes.

[0011] In an optional embodiment of this application, the two CPUs in one computing node are electrically connected using a Fast Link Interconnect (QPI).

[0012] In an optional embodiment of this application, in the above-mentioned computing node, each CPU is electrically connected to four GPUs via a peripheral component interconnect high-speed bus PCIe.

[0013] In an optional embodiment of this application, the target CPUs of the above-mentioned plurality of computing nodes are connected via a network interface controller (NIC).

[0014] In an optional embodiment of this application, GPUs with the same number among the aforementioned multiple computing nodes are optically connected using optical fibers.

[0015] Secondly, embodiments of the present invention also provide a transmission system based on optical interconnection, the transmission system based on optical interconnection including the above-described optical interconnection-based topology.

[0016] The embodiments of the present invention bring the following beneficial effects: This invention provides a topology and transmission system based on optical interconnects. The optical interconnect-based topology includes multiple computing nodes, each node comprising two central processing units (CPUs) and eight graphics processing units (GPUs). Within each computing node, the two CPUs are electrically connected, and each CPU is electrically connected to each of the four GPUs. One CPU from each computing node is designated as the target CPU, and the target CPUs across multiple computing nodes are connected via a network. The eight GPUs in each computing node are numbered, and GPUs with the same number across multiple computing nodes are optically connected. This approach can improve model inference speed while reducing hardware cost, size, and power consumption.

[0017] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.

[0018] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 A schematic diagram of a topology based on optical interconnect provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a full interconnection structure of GPUs with the same node number provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating a GPU data transmission process between copper interconnect nodes, provided as an embodiment of the present invention. Figure 4 This is a schematic diagram of a GPU data transmission process between optical interconnect nodes provided in an embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Currently, some topologies are not optimized for the actual communication modes of model inference. Most mainstream model inference methods employ parallel strategies such as Tensor Parallelism (TP) and Pipeline Parallelism (PP). Among these, the TP mode has extremely high requirements for direct communication between GPUs within the group, primarily manifested in the following operational stages: 1. Cross-GPU splitting and concatenation of input and output features: When input data enters a layer segmented by the TP (Transformer Module), it needs to be split according to feature dimensions and distributed to the GPUs within the group. This process requires data communication. After network computation, the output features of the GPUs within the group need to be concatenated back to complete dimensions before entering the next layer. This step requires data exchange and concatenation between GPUs, resulting in significant communication.

[0023] 2. Cross-GPU communication in multi-head attention layers: In multi-head attention within Transformer models (a deep learning architecture based on self-attention), the attention layer (TP) typically splits the "multi-head" or "feature dimension" across multiple GPUs. If the attention head is split, each GPU needs to obtain the K (key) and V (value) information from other GPUs to calculate the complete attention weights when calculating the attention score, resulting in cross-GPU K and V exchanges. The linear layer of the attention output is also usually split by the TP, and the outputs of each GPU need to be concatenated to form complete features, requiring further communication.

[0024] 3. Communication in large-dimensional linear layers: For large-dimensional linear layers in the model, TP will split the weights according to the input or output dimensions. When the input features pass through the split linear layers, the output of each GPU is only a partial result, which needs to be concatenated into a complete dimension through communication. The amount of communication in this process is proportional to the dimension.

[0025] 4. Cross-GPU synchronization before layer normalization: If the normalization layer is after the TP segmentation layer and the mean and variance need to be calculated based on the complete features, the two GPUs need to exchange their respective feature statistics for synchronization before performing the normalization operation, resulting in additional communication.

[0026] During model inference, a large amount of data transfer is required between GPUs within a tensor parallel group. However, in traditional copper interconnect structures, data interaction between GPUs within the group must pass through the PCIe (Peripheral Component Interconnect Express) bus, CPU (Central Processing Unit) forwarding, or switch relay, increasing communication latency and limiting parallel efficiency. In dense optical interconnect structures, redundant GPU connections between tensor parallel groups occur, resulting in significant resource waste and increased costs.

[0027] Therefore, there is an urgent need for a cluster connection topology that balances high bandwidth, low latency, and resource conservation to adapt to the increased computing power of GPUs and the parallel characteristics of model inference.

[0028] Based on this, embodiments of the present invention provide a topology and transmission system based on optical interconnects, specifically providing a cluster multi-node GPU collaborative connection topology based on optical interconnects and a novel connection method between cluster multi-nodes based on optical interconnects, mainly including: 1. The concept of sparse optical interconnect was proposed, which does not pursue extremely dense optical interconnect in the interconnection method of the cluster, but adopts high-efficiency sparse optical interconnect.

[0029] 2. The idea of ​​developing connection methods based on reasoning patterns is proposed. The connection principle is not a fixed, dense interconnection, but is tailored to the reasoning pattern to improve efficiency.

[0030] 3. Based on the parallel reasoning mode of TP, a novel interconnection method based on optical interconnection was developed.

[0031] To facilitate understanding of this embodiment, a detailed description of an optical interconnect topology disclosed in this embodiment of the invention will be provided first.

[0032] Example 1: This invention provides a topology based on optical interconnects, comprising: multiple computing nodes, each computing node including: 2 central processing units (CPUs) and 8 graphics processing units (GPUs); in one computing node, the 2 CPUs are electrically connected, and each CPU is electrically connected to each of the 4 GPUs; one CPU from each computing node is designated as the target CPU, and the target CPUs of the multiple computing nodes are connected via a network; the 8 GPUs in each computing node are numbered, and GPUs with the same number from multiple computing nodes are optically connected.

[0033] See Figure 1 The diagram shows a topology based on optical interconnect. Figure 1 The diagram shows two computing nodes (i.e. Figure 1 The topology of nodes A and B in the network. Both node A and node B include: 2 CPUs (i.e., ... Figure 1 CPU-0 and CPU-1) and 8 GPUs (i.e. Figure 1 GPU-0 to GPU-7 (in the GPU-0 to GPU-7).

[0034] In some embodiments, the number of computing nodes is determined based on the tensor parallelism used for model inference.

[0035] In this embodiment, the number of computing nodes N (N≥2) can be determined according to the TP method. Figure 1 The diagram shows the case where N=2.

[0036] In some embodiments, the CPU is used for resource scheduling and control within a compute node.

[0037] In some embodiments, the GPU is used to perform model inference computation tasks.

[0038] In this embodiment, the CPU can perform resource scheduling and control within the computing node, while the GPU can perform inference computing tasks for executing models (e.g., Transformer models).

[0039] In some embodiments, each GPU includes: 1 electrical interface and N-1 optical interfaces; wherein, N is the number of compute nodes, and N is an integer greater than or equal to 2; the electrical interface is used to connect to CPUs in the same compute node; the optical interface is used to connect to GPUs with the same number in different compute nodes.

[0040] like Figure 1 As shown, each GPU includes: 1 electrical interface (copper port) and N-1 optical interfaces. The electrical interface is used to connect to the CPU, and the optical interfaces are used for full interconnection with other GPUs of the same number on other nodes.

[0041] I. Intra-node communication links: In some embodiments, in one compute node, two CPUs are electrically connected via a Fast Link Interconnect (QPI).

[0042] In some embodiments, in one computing node, each CPU is electrically connected to four GPUs via a peripheral component interconnect high-speed bus (PCIe).

[0043] like Figure 1 As shown, in a single computing node, the CPU and GPU are connected via a PCIe bus, and the CPUs interact with each other at high speed through Quick Path Interconnect (QPI).

[0044] QPI is a high-speed serial point-to-point connection technology primarily used for data transmission between processors and chipsets. It replaces the traditional front-side bus, solving the bandwidth bottleneck problem. QPI features high bandwidth, bidirectional transmission support, and low latency, making it widely applicable in processors.

[0045] PCIe bus is a high-speed serial computer expansion bus standard, primarily used to connect external devices on the motherboard, such as graphics cards, solid-state drives, and network cards. It uses point-to-point full-duplex transmission, with each device having its own dedicated bandwidth, avoiding the performance bottlenecks of traditional parallel buses.

[0046] II. Inter-node communication links: In some embodiments, target CPUs of multiple computing nodes are networked together using a Network Interface Controller (NIC).

[0047] In some embodiments, GPUs with the same number on multiple computing nodes are optically connected using optical fibers.

[0048] like Figure 1 As shown, among the target CPUs of multiple computing nodes (i.e. Figure 1 The CPU-0 in the compute node uses a NIC (Network Interface Card) for network connectivity. Multiple compute nodes share the same GPU number (i.e.,...). Figure 1 Optical fiber is used to connect the GPU-0 in the system.

[0049] The NIC is a key piece of hardware that enables computers to connect to the internet. It is responsible for converting data into a format that the network can recognize and then transmitting it through a network cable.

[0050] Optical fiber is a thin, long fiber made of glass or plastic that transmits optical signals using the principle of total internal reflection. It is a core transmission medium in the field of communications. It mainly consists of a core, cladding, and outer sheath, and achieves efficient information transmission through optical pulses.

[0051] See also Figure 2 The diagram shows a full interconnection structure of GPUs with the same node number. Taking GPU-0 as the target CPU as an example, the GPU-0 of each node are fully interconnected via optical fibers, as shown below. Figure 2 As shown, when the total number of nodes is 4, the connection method between cluster machines is as shown by the solid line, and each GPU needs 3 optical interfaces; when the total number of nodes is N, the connection method between cluster machines adds dashed lines to the solid lines, and each GPU needs N-1 optical interfaces.

[0052] GPUs with the same number on different computing nodes are directly connected via optical fiber (i.e., GPU-i on node A and GPU-i on node B are connected point-to-point via optical interface, i=0-7). GPUs with different numbers between nodes are transmitted via PCIe bus, CPU forwarding, network interface controller (NIC) and traditional network (such as Ethernet).

[0053] In summary, the embodiments of the present invention mainly provide the following: (1) Targeted optical interconnect design: Direct optical fiber connection is configured only for GPUs with the same number between nodes to meet the communication requirements in TP parallel mode. When the model is split in TP parallel mode, GPUs with the same number are logical cooperative pairs. Direct optical interconnect can avoid PCIe bus, CPU or switch relay, reduce communication hops and increase transmission bandwidth.

[0054] (2) Change in model allocation method: In the past, the TP parallel mode was mostly completed on the machine. This invention changes the TP parallel mode to the machine through the direct optical interconnection of GPUs with the same number between machines, which balances the load of each node and makes the cluster run more evenly.

[0055] (3) Hybrid connection architecture: Combining the advantages of PCIe (high-efficiency communication within a node), QPI (inter-CPU collaboration), optical interconnect (high-speed direct connection of GPUs with the same signal) and NIC (general data and control signaling transmission), a hierarchical communication system is formed.

[0056] (4) Resource optimization configuration: When the cluster contains N nodes, each GPU only needs N-1 optical interfaces (instead of the 8×(N-1) required for full interconnection between machines). The number of optical interfaces and optical fibers are 1 / 8 of the full interconnection, and no optical switch connection is required, which reduces hardware costs, while reducing the cluster size and power consumption.

[0057] This invention provides a topology based on optical interconnects. The topology includes multiple computing nodes, each node comprising two central processing units (CPUs) and eight graphics processing units (GPUs). Within each computing node, the two CPUs are electrically connected, and each CPU is electrically connected to each of the four GPUs. One CPU from each computing node is designated as the target CPU, and the target CPUs across multiple computing nodes are connected via a network. The eight GPUs in each computing node are numbered, and GPUs with the same number across multiple computing nodes are optically connected. This approach can improve model inference speed while reducing hardware cost, size, and power consumption.

[0058] Example 2: This invention provides another optical interconnect-based topology, which is implemented based on the above embodiments. Taking a dual-machine 16-card cluster as an example, the specific implementation of the optical interconnect-based topology is described in detail.

[0059] I. Hardware configuration: Node A and Node B each contain 2 CPUs (supporting QPI) and 8 GPUs; II. The specific connection method is as follows: (1) Each CPU connects to 4 GPUs via a PCIe 4.0 x8 bus (e.g., CPU-0 connects to GPU-0 to GPU-3, and CPU-1 connects to GPU-4 to GPU-7). (2) The CPUs of node A and node B are connected through the QPI channel to achieve CPU collaboration within the node; (3) Each GPU uses one PCIe electrical interface (for intra-node communication) and one optical interface (for inter-node same-signal connection); (4) GPU-i (i=0-7) of node A and GPU-i (i=0-7) of node B are connected point-to-point via optical fiber, without the need for an optical switch; (5) The CPU-0 of node A and the CPU-0 of node B are connected through NIC for non-same GPU data transmission and control signaling interaction.

[0060] The optical interconnect-based topology provided in this embodiment of the invention has the following main advantages: (1) Improves inference speed compared to traditional copper interconnect.

[0061] Compared to traditional copper interconnects, optical interconnects between GPUs of the same architecture significantly improve data transmission speed, especially for model inference performance in TP mode.

[0062] See Figure 3 The diagram illustrates a GPU data transfer process between copper interconnect nodes. Figure 3 This illustrates a scenario with four copper interconnect computing nodes. In a traditional copper interconnect topology, data exchange between nodes using the same GPU number first requires the GPU to transmit data to the CPU via the PCIe bus, the CPU to transmit the data to the switch via the NIC, the switch to the next node, and finally, the data to the target GPU via the PCIe bus.

[0063] Figure 4 The diagram illustrates a GPU data transmission process between optical interconnect nodes. Figure 4 The diagram illustrates a scenario with four optical interconnect computing nodes. Compared to copper interconnects, optical interconnects, through optical interconnects of GPUs with the same node number, achieve point-to-point data transmission, eliminating the need for transmission to the CPU and switch, thus reducing resource consumption and connection hop count. When eight GPUs transmit data simultaneously, the data shares the bandwidth of the NIC and switch, increasing transmission latency by more than eight times and slowing down model inference speed by more than eight times.

[0064] (2) Compared with dense optical interconnects, resources are used more efficiently.

[0065] To avoid over-configuration of dense optical interconnects, assuming N is the number of cluster nodes, dense inter-machine optical interconnects require 8×(N-1) optical ports. However, in this connection method, only N-1 optical ports are needed, reducing the number of optical ports and optical paths to 1 / 8. The direct connection method does not require optical switches, significantly reducing hardware costs and energy consumption. Furthermore, it can maintain the same inference efficiency in the TP's parallel mode.

[0066] (3) Strong compatibility.

[0067] It retains traditional interfaces such as PCIe, QPI, and NIC, and can be seamlessly adapted to existing GPU and CPU hardware and model deployment frameworks (such as TensorFlow and PyTorch) without requiring large-scale modifications to the software stack.

[0068] (4) Good scalability.

[0069] The topology can be expanded to multi-node clusters (such as 4 machines with 32 cards or 8 machines with 64 cards) by simply adding fiber optic connections between GPUs of the same number, thus maintaining the simplicity of the architecture.

[0070] III. The model reasoning process is as follows: Taking a dual-machine 16-card cluster as an example, when deploying a model with TP=2, the model layer is split into two parts and deployed on the same GPUs of node A and node B respectively. During inference, feature tensor exchange between GPUs of the same number is directly transmitted through optical fiber. Asymmetric data (such as input / output and control commands) is transmitted through NIC. PP parallelism is used within the node, and data interaction is minimal. Interaction between CPUs is completed through PCIe, and collaboration between CPUs is completed through QPI, which does not cause a large amount of latency. It maintains the same push speed as dense optical connections, thereby significantly improving the inference efficiency during TP inference.

[0071] The aforementioned co-signal GPU interconnection method based on optical interconnection provided in this invention can leverage the cluster advantage in tensor parallel (TP) parallel mode, achieving faster inference speeds with fewer resources. Furthermore, building upon the existing connection method, a sparse optical interconnection concept is proposed, allowing for the development of specific interconnection structures tailored to particular inference methods.

[0072] Example 3: This invention provides a transmission system based on optical interconnection, which is implemented based on the above embodiments. The transmission system based on optical interconnection includes the optical interconnection-based topology provided in the foregoing embodiments.

[0073] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the optical interconnect-based transmission system described above can be referred to the corresponding process in the foregoing embodiments, and will not be repeated here.

[0074] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0075] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0076] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A topology based on optical interconnect, characterized in that, The optical interconnect-based topology includes: multiple computing nodes, each computing node including: 2 central processing units (CPUs) and 8 graphics processing units (GPUs); In one computing node, two CPUs are electrically connected to each other, and each CPU is electrically connected to four GPUs respectively. One CPU from each computing node is selected as the target CPU, and the target CPUs of multiple computing nodes are connected by a network. Each of the eight GPUs in the computing node is numbered, and the GPUs with the same number in multiple computing nodes are connected by optical connections.

2. The optical interconnect-based topology according to claim 1, characterized in that, The number of computing nodes is determined based on the tensor parallelism used for model inference.

3. The optical interconnect-based topology according to claim 1, characterized in that, The CPU is used for resource scheduling and control within the computing node.

4. The optical interconnect-based topology according to claim 1, characterized in that, The GPU is used to perform model inference computation tasks.

5. The optical interconnect-based topology according to claim 1, characterized in that, Each GPU includes: 1 electrical interface and N-1 optical interfaces; where N is the number of computing nodes and is an integer greater than or equal to 2; The electrical interface is used to connect the CPU in the same computing node; The optical interface is used to connect GPUs with the same number in different computing nodes.

6. The optical interconnect-based topology according to claim 1, characterized in that, In one of the computing nodes, the two CPUs are electrically connected via a Fast Link Interconnect (QPI).

7. The optical interconnect-based topology according to claim 1, characterized in that, In one of the computing nodes, each CPU is electrically connected to the four GPUs via a PCIe high-speed bus for peripheral component interconnection.

8. The optical interconnect-based topology according to claim 1, characterized in that, The target CPUs of the multiple computing nodes are connected via a network interface controller (NIC).

9. The optical interconnect-based topology according to claim 1, characterized in that, The GPUs with the same number on multiple computing nodes are optically connected by optical fibers.

10. A transmission system based on optical interconnect, characterized in that, The optical interconnect-based transmission system includes: the optical interconnect-based topology as described in any one of claims 1-9.