Torus network scale expansion method based on optical switch, medium and equipment
By introducing optical switches to connect supernodes into the Torus network and building a hierarchical structure, the congestion, port pressure and resource fragmentation problems of the Torus network during the expansion process are solved, efficient resource utilization and communication optimization are achieved, and the needs of large-scale distributed training clusters are met.
Patent Information
- Application Number
- CN202510443864.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
AI Technical Summary
During the expansion process, the existing Torus network has problems caused by dimensional expansion, port pressure, resource fragmentation caused by node heterogeneity, and single-node failures affecting the entire network, which is difficult to meet the needs of large-scale distributed training clusters.
Optical switches are used to connect supernodes, build a hierarchical Torus structure, optimize local communications, and dynamically interconnect the links between supernodes through optical switches, combining direct path priority and indirect load balancing strategies to reduce node fragmentation and meet the communication needs of different distributed training tasks.
It realizes flexible expansion of large-scale Torus networks, improve resource utilization, optimize communication performance, reduce hardware costs and management overhead, and ensures that latency and throughput meet performance requirements.
Smart Images

Figure CN120281727A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data center networks, and particularly relates to a method, medium, and device for expanding the scale of a Torus network based on an optical switch. Background Art
[0002] Compared with data center network architectures such as Dragonfly, BCube, and FatTree, the Torus network has significant hardware cost advantages and high scalability. However, when constructing a single Torus network for a large-scale distributed training task cluster, the following bottlenecks exist: (1) Congestion bottleneck caused by dimension expansion: One-dimensional linear expansion will cause the network diameter to increase with the growth of the node scale. (2) Port pressure brought by high-dimensionalization: Constructing a k-dimensional Torus requires each node to be configured with 2k physical ports for adjacent connections. When k≥3, the port density requirement of a single node exceeds the integration capacity of mainstream switching chips, forcing the adoption of a multi-level switching structure, resulting in non-linear growth of hardware costs. (3) Resource fragmentation caused by node heterogeneity: When expanding the network by enhancing the switching capacity of a single node, heterogeneous hardware (such as differences in NVLink versions of different generations of GPUs) will destroy the regular topological characteristics of the Torus, increasing routing complexity and reducing bandwidth utilization.
[0003] In addition, all of the above expansion schemes will face the problem that if a single node fails, it will affect the entire network traffic, which will lead to an increase in the cost of network upgrade and maintenance. At the same time, the physical deployment of ultra-large-scale Torus is limited by the computer room space and wiring complexity, and it is difficult to meet the requirements of the distributed training cluster for geographical distribution flexibility. Summary of the Invention
[0004] Aiming at the scale limitations of the existing Torus and the performance defects of the multi-dimensional Torus, the present invention provides a method, medium, and device for expanding the scale of a Torus network based on an optical switch, which uses a hierarchical Torus structure with a single Torus network as a super node. Inside each super node, a small-scale Torus is used to optimize local communication for compute-intensive tasks, while a dynamic interconnection layer is constructed between super nodes through an OCS.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for expanding the scale of a Torus network based on an optical switch, including: using a k×k Torus structure as a supernode, and connecting the supernodes through an optical switch, where each supernode has k ports connected to the optical switch; for the number of nodes required for a task, task nodes are preferentially placed within the same supernode, and the unused nodes within the supernode whose number is less than the number of nodes contained in the supernode are regarded as fragments and placed in the cluster; if there is a set of fragments available for the task in the cluster, select appropriate fragments from the set of fragments for partitioning; if there is no set of fragments available for the task in the cluster, partition the complete supernode.
[0007] Optionally, the placement of the task nodes aims to minimize the use of fragments.
[0008] Optionally, when there is a set of fragments available for the task in the cluster, the fragments in the set of fragments are arranged in ascending order and divided into the following three types: the fragment size is less than the number of nodes required for the task; the fragment size is greater than or equal to the number of nodes required for the task and the fragment is not directly connected to the optical switch; the fragment size is greater than or equal to the number of nodes required for the task and there are nodes within the fragment directly connected to the optical switch.
[0009] Optionally, if the utilization rate of the cluster is less than or equal to the threshold value, select and partition the fragments to meet the node quantity requirement and communication requirement of the task; if the utilization rate of the cluster is greater than the threshold value, preferentially select the fragments that meet the node quantity requirement of the task, without ensuring that there are nodes within the fragments directly connected to the optical switch.
[0010] Optionally, for tasks that do not require external communication, preferentially select fragments whose size is greater than or equal to the number of nodes required for the task and the fragment is not directly connected to the optical switch; for tasks that require external communication, preferentially select fragments whose size is greater than or equal to the number of nodes required for the task and there are nodes within the fragment directly connected to the optical switch.
[0011] Optionally, when partitioning the complete supernode, for tasks that do not require external communication, place the task nodes on the side far from the optical switch.
[0012] Optionally, the optical switch establishes connections between ports according to traffic requirements, and does not interfere with the running tasks during reconfiguration and ensures communication between task nodes placed in different supernodes.
[0013] Optionally, when task nodes cannot directly communicate through the optical switch, forward the traffic through other task nodes directly connected to the optical switch.
[0014] In a second aspect, the present invention provides a computer-readable storage medium storing a computer program, which causes a computer to execute the method for expanding the scale of a Torus network based on an optical switch as described in the first aspect.
[0015] In a third aspect, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for expanding the scale of a Torus network based on an optical switch as described in the first aspect is implemented.
[0016] The beneficial effects of the present invention are as follows:
[0017] (1) The present invention adopts corresponding placement strategies for different distributed training tasks, reduces node fragmentation in the cluster, and improves resource utilization.
[0018] (2) The present invention effectively meets the communication requirements of different distributed training tasks by dynamically configuring the links between supernodes.
[0019] (3) The present invention combines the direct path priority and indirect load balancing strategies to optimize the Torus communication performance.
[0020] The present invention can achieve flexible scale expansion of a large-scale Torus network without introducing excessive hardware costs and management overheads; for each distributed training task, its connectivity can be effectively satisfied during reconfiguration, and it is ensured that the latency and throughput meet the performance requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is the physical topology of the Scale Expansion of Torus Network (STON) based on an optical switch.
[0022] Figure 2a is the placement (least fragmentation) of a task with 0 nodes on STON(4, 5).
[0023] Figure 2b is the placement (fragment splitting) of a task with 20 nodes on STON(4, 5).
[0024] Figure 3 are different forms of fragmentation existence.
[0025] Figure 4 is the physical topology of STON(4, 5).
[0026] Figure 5 is the logical topology of STON(4, 5).
[0027] Figure 6a is the communication requirement between supernodes when there are 4 tasks.
[0028] Figure 6b It is the communication requirement of the super node after the arrival of a new task.
[0029] Figure 7 It is the consideration of introducing task connectivity by multiple optical switches.
[0030] Figure 8a It is the logical topology of Case 1.
[0031] Figure 8b It is the logical topology of Case 2.
[0032] Figure 8c It is the logical topology of Case 3.
[0033] Figure 9a It is the CDF graph of the flow completion time of Case 1.
[0034] Figure 9b It is the average flow completion time and tail delay distribution of Case 1.
[0035] Figure 10a It is the CDF graph of the flow completion time of Case 2.
[0036] Figure 10b It is the average flow completion time and tail delay distribution of Case 2.
[0037] Figure 11a It is the CDF graph of the flow completion time of Case 3.
[0038] Figure 11b It is the average flow completion time and tail delay distribution of Case 3.
[0039] Figure 12a It is the CDF graph of the flow completion time of the Kalos dataset.
[0040] Figure 12b It is the average flow completion time and tail delay distribution of the Kalos dataset. Specific implementation mode
[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0042] The present invention introduces an optical circuit switch (OCS) as a bridge between supernodes. Compared with the traditional electrical circuit switch (EPS), the OCS has advantages in terms of architecture and performance. First, based on the passive optical transmission characteristics, the OCS is completely transparent to signal rates and communication protocols, and there is no need to replace the hardware as the server network card rate is upgraded. For example, the same OCS can adapt to optical modules of different generations from 10G to 800G, and its single-port cost remains constant in cross-generation deployment. Moreover, as an integral part of the data center infrastructure, its capital expenditure can be amortized over a 5-10-year cycle. Second, the OCS does not require optoelectronic conversion and packet processing logic. The typical power consumption of its 400G port is less than 1W (the power consumption of the corresponding EPS port is greater than 10W), and the transmission delay is on the order of dozens of nanoseconds (the EPS delay is hundreds of microseconds). In addition, the OCS uses miniaturized devices (such as a 3D-MEMS system with a hundred ports only requires two 2D MEMS chips), and its system-level failure probability is lower than that of the EPS. And if a faulty link occurs, the OCS can achieve the recovery of the faulty link through redundant ports and software-defined optical path switching, while the EPS requires manual optical module replacement. However, the OCS lacks packet processing capabilities, so in terms of actual functions, the OCS is only equivalent to pairing and linking two input-output ports. To maximize the potential of OCS reconfiguration, accurate traffic prediction needs to be implemented, while the traditional traffic in the DCN is difficult to predict. Fortunately, the distributed training task is based on a known dataset, and its traffic pattern has weak predictability and periodicity, which just makes up for the defect of OCS reconfiguration. Therefore, it is feasible to use the OCS to connect the supernodes constructed by the Torus network to form a distributed training cluster network.
[0043] Here, the link utilization rate of the (2n + 1)×(2n + 1) Torus in the AlltoAll traffic pattern is analyzed, where Rate injection is the rate at which each node injects data, and Rate forwarding is the forwarding rate of each node's single port. For a single node in the (2n + 1)×(2n + 1) Torus, there are four ports, namely up, down, left, and right. Then the forwarding capacity of a single node is:
[0044] C forwarding = 4Rate forwarding
[0045] The forwarding capacity required for a single node is:
[0046]
[0047] Assuming that the node divides the bandwidth equally for each flow, then the above formula should be:
[0048]
[0049] In the AlltoAll communication mode, by analyzing the path lengths from any node to other nodes, it can be found that when the path length is less than or equal to n, the corresponding number of nodes is 4l; when the path length is greater than or equal to n + 1, the corresponding number of nodes is 4(2n - l + 1). Combining the path information, the above formula can be written as:
[0050]
[0051] Then, the link utilization rate of the (2n + 1)×(2n + 1) Torus in the AlltoAll traffic mode can be obtained as:
[0052]
[0053] Similarly, the link utilization rate of the (2n + 1)(2n + 1)(2n + 1) Torus in the AlltoAll traffic mode can be obtained as:
[0054]
[0055] When the link utilization rate is greater than 1, congestion will occur in the network. For a 2D Torus, if the injection rate and the forwarding rate are equal, congestion will occur at o = 3.5, that is, congestion will occur when the topology scale is greater than 8×8. That is to say, if you want to expand the Torus scale by increasing the number of nodes in a single dimension, congestion problems will be introduced into the network.
[0056] Based on the above analysis of the Torus scale expansion scheme, in one embodiment, the present invention proposes a method for expanding the Torus network scale based on optical switches (Scalable Torus Network Architecture Based on Optical Switches, STON) to construct a distributed training task cluster, which mainly includes three parts: task placement, logical topology mapping, and traffic forwarding.
[0057] I. Task Placement
[0058] STON has different placement strategies for tasks of different sizes. Tasks that require more nodes than the number of nodes in a single supernode are regarded as large tasks, while tasks that require fewer nodes than the number of nodes in a single supernode are regarded as small tasks. As shown in lines 10 to 20 of the algorithm in Table 1, for the number of nodes required by a task, STON tries to place the nodes within the same supernode. The number of nodes less than one supernode is treated as fragments and placed in the cluster. If there is an existing set of available fragments in the cluster, then a suitable fragment is searched for and partitioned within that set of fragments; if not, the algorithm in Table 3 is used to split a complete supernode. This design is to reduce the traffic requirements between supernodes, Figure 2a and 2b shows two ways of placing fragments, Figure 2a for minimizing the number of fragments, i.e., there is only one fragment; Figure 2b for splitting fragments, for a fragment of size y nodes, without considering the cluster capacity, the number of placement ways can be calculated by the function Here only one is shown. Assuming an AlltoAll communication is carried out within the task, then Figure 2a the communication requirement between supernodes is 16×8×2 = 256 streams. Figure 2b the communication requirement between supernodes is 16×4×4 + 4×4×2 = 288 streams. Keeping the number of fragments to a minimum in placement can not only reduce the communication requirements between supernodes, but also minimize the generation of fragments within the cluster, making it easier for other tasks to be placed.
[0059] When there is already a set of fragments in the cluster, STON selects a suitable fragment from the set of fragments through the algorithm in Table 2 for placement. For the convenience of decision-making, the set of fragments is arranged in ascending order. There may be several ways in which fragments appear. First, the fragment size is less than the size required by the task; second, the fragment size is greater than or equal to the size required by the task and the fragment is not directly connected to the OCS; third, the fragment size is greater than or equal to the size required by the task and there are nodes within the fragment directly connected to the OCS. The selection of fragments is related to the current state of the cluster and the requirements of the task. If the utilization rate of the cluster is less than or equal to the threshold value, then STON will tend to meet the requirements of the task to obtain better performance; if the utilization rate of the cluster is greater than the threshold value, STON will tend to select a fragment of a suitable size rather than split a large fragment, even if it may not be able to meet the direct connection communication between the fragment and the supernode. Figure 3Shows different ways of fragment existence. For an 8-node task that does not need to communicate externally, the best choice is the fragments within Group C; for an 8-node task that needs to communicate externally, the best choice is the fragments within Group B. If such fragments do not exist, the fragments within Group D can also meet its requirements. The partitioning of the complete supernodes is similar to that of the fragments. The main consideration factor is the external communication requirement. As shown in the algorithm in Table 3, for tasks that do not require OCS connection, STON places them on the side farther from the OCS.
[0060] Table 1 Task Placement Algorithm
[0061]
[0062] II. Logical Topology Mapping
[0063] For a supernode composed of 4×4 Torus, each supernode has 4 ports connected to the optical switch, as Figure 4 shown. The optical switch establishes connections between ports according to different traffic demands, Figure 4 and the corresponding logical topology is as Figure 5 shown.
[0064] The mapping from the physical topology to the logical topology needs to meet two requirements: reconfiguration should not interfere with the running tasks and each task should ensure connectivity and performance. In Figure 6a there are only two tasks with communication demands between supernodes, while Figure 6b at a certain moment, new tasks and corresponding communication demands appear. At this time, the red task nodes cannot directly communicate through the OCS. To avoid interfering with the running gray tasks, the red task nodes will achieve their communication demands through traffic forwarding, which will be introduced in detail in the subsequent traffic forwarding section of this work.
[0065] Commercially available MEMS-OCS already has 320×320 ports. If the scale of the supernode is k×k, then using a single optical switch can connect at most supernodes, that is, 320×k nodes. When k = 4, the number of nodes is 1280. Therefore, to form a larger-scale cluster, multiple OCSs need to be introduced for connection. When there are multiple OCSs, the connectivity of the supernodes needs to be considered. Figure 7 Shows that if a task is placed on supernode C and supernode G, connectivity problems may occur.
[0066] Table 2 Fragment Selection Algorithm
[0067]
[0068] Table 3 Supernode Splitting Algorithm
[0069]
[0070] 3. Traffic forwarding
[0071] In the previous design, we can find that if the cluster load is high, tasks will be placed on multiple supernodes, and there will be communication needs between them but no direct links. At this time, it is necessary to introduce a traffic forwarding strategy. The communication of distributed training is regular and intermittent, and is not as unpredictable as the data center network. Figure 6b In the figure, without changing the existing logical topology, the gray task node needs to forward traffic to the red task node, which is feasible for distributed training traffic and will not affect the communication of the gray task node. The communication regularity of distributed training tasks is obvious, and there is a long computing time in the middle. During this period, traffic forwarding will not affect the tasks of the current node.
[0072] Next, the STON proposed in this embodiment is evaluated in combination with specific experiments.
[0073] The NS3 network simulator is used in the experiment to evaluate the performance of STON and the baseline scalability solution (hereinafter referred to as the baseline solution). All experiments use DCQCN as the congestion control algorithm and PFC as the flow control algorithm. The bandwidth of all links is 100Gbps and the propagation delay is lus; the data injection rate of each node is 400Gbps. The traffic within the task is simulated using AlltoAll communication.
[0074] The following indicators reflecting network performance are used in the experiment: average flow completion time (AvgFCT), flow tail delay (95th percentile FCT, 99th percentile FCT, 99.9th percentile FCT); and the performance is evaluated in different scenarios.
[0075] The baseline scale expansion scheme is set to direct connection between Torus supernodes. If the scale of supernodes is k×k, there are two cases. The first is that the number of supernodes is less than or equal to k+1, then any two supernodes can be directly connected; the second is that the number of supernodes is greater than k+1, then some supernodes cannot be directly connected. For this case, the design of this embodiment is also adopted.
[0076] This experiment analyzes several case studies. All case studies are conducted under 5 4×4Torus supernode topologies without special instructions. The physical topology is as follows Figure 1 The task is described in Json format, as shown in the algorithm in Table 4.
[0077] Table 4 Task description format
[0078]
[0079] The first case includes two tasks that arrive simultaneously. The experimental results are as Figure 9a and 9b shown. Compared with the baseline scheme, the avgFCT of STON is reduced by 42.2%, the 95th percentile FCT is reduced by 28.6%, the 99th percentile FCT is reduced by 28.5%, and the 99.9th percentile FCT is reduced by 27.1%. In this case, the logical topology corresponding to STON is as Figure 8a shown; while there is no reconfiguration in the baseline scheme, which is the reason for the performance improvement of STON compared with the baseline scheme.
[0080] The second case includes three tasks that arrive simultaneously. The experimental results are as Figure 10a and 10b shown. Compared with the baseline scheme, the avgFCT of STON is reduced by 61.1%, the 95th percentile FCT is reduced by 71.5%, the 99th percentile FCT is reduced by 60.4%, and the 99.9th percentile FCT is reduced by 60.2%. In this case, the logical topology corresponding to STON is as Figure 8b shown. There are two fragments in this case. STON divides the links according to their fragment sizes because the links from a supernode to the OCS are limited; in fact, if resource limitations are not considered, placing the fragment of Task 0 on a new node can achieve better performance.
[0081] The third case includes five tasks that arrive simultaneously. The experimental results are as Figure 11a and 11b shown. Compared with the baseline scheme, the avgFCT of STON is reduced by 42.9%, the 95th percentile FCT is reduced by 60.6%, the 99th percentile FCT is reduced by 51.2%, and the 99.9th percentile FCT is reduced by 50.7%. In this case, the logical topology corresponding to STON is as Figure 8c shown. This case includes five fragments. When placing Task 1 and Task 5, it is necessary to consider minimizing the communication resources occupied between supernodes; when placing Task 1, since no supernode communication is required, it is placed at the end of Supernode A far from the communication. When placing Task 0, since it is the first division of Supernode A and the fragment of Task 0 needs to communicate externally, the fragment of Task 0 obtains a set of nodes in the same column in Supernode A, that is, 4 nodes and 1 external link. When placing Task 3 in Supernode A, the remaining nodes in Supernode A fit the fragment size of Task 3, so 4 nodes and 3 external links can be obtained.
[0082] This experiment also selected the dataset publicly disclosed by the Shanghai AI Laboratory in the paper. The dataset includes distributed training task data published on the Kalos cluster from March 2023 to August 2023. Since the original dataset included 470,497 GPU tasks, task data from a randomly selected day was used in the experiment to simulate the real task arrival situation. The experimental results are as Figure 12a and 12b shown. Compared with the baseline scheme, the avgFCT of STON decreased by 52.6%, the 95th percentile FCT decreased by 74.5%, the 99th percentile FCT decreased by 73.4%, and the 99.9th percentile FCT decreased by 72.8%.
[0083] In another embodiment, the present invention proposes a computer-readable storage medium storing a computer program, which causes a computer to execute the method for expanding the scale of the Torus network based on an optical switch in the foregoing embodiment.
[0084] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for expanding the scale of the Torus network based on an optical switch in the foregoing embodiment is implemented.
[0085] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CDROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0087] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. Method for expanding the scale of a Torus network based on an optical switch, characterized in that, Take the k×k Torus structure as a supernode, and connect the supernodes through optical switches. Each supernode has k ports connected to the optical switches. For the number of nodes required for a task, preferentially place the task nodes within the same supernode. The unused nodes within the supernode whose number is less than the number of nodes contained in the supernode are regarded as fragments and placed in the cluster. If there is a set of fragments available for the task in the cluster, select a suitable fragment from the set of fragments for partitioning. If there is no set of fragments available for the task in the cluster, then partition the complete supernode.
2. The method for expanding the scale of the Torus network based on an optical switch as claimed in claim 1, wherein: The placement of the task nodes aims to minimize the use of fragments.
3. The method for expanding the scale of the Torus network based on an optical switch as claimed in claim 1, wherein: When there is a set of fragments available for the task in the cluster, arrange the fragments in the set of fragments in ascending order and divide them into the following three types: the fragment size is less than the number of nodes required for the task; the fragment size is greater than or equal to the number of nodes required for the task and the fragment is not directly connected to the optical switch; the fragment size is greater than or equal to the number of nodes required for the task and there are nodes in the fragment directly connected to the optical switch.
4. The method for expanding the scale of the Torus network based on an optical switch as described in claim 3, characterized in that: If the utilization rate of the cluster is less than or equal to the threshold value, select and partition the fragments to meet the node quantity requirement and communication requirement of the task. If the utilization rate of the cluster is greater than the threshold value, preferentially select the fragments that meet the node quantity requirement of the task, without guaranteeing that there are nodes in the fragments directly connected to the optical switch.
5. The method for expanding the scale of the Torus network based on an optical switch as described in claim 3, wherein: For tasks that do not need to communicate externally, preferentially select the fragments whose size is greater than or equal to the number of nodes required for the task and the fragments are not directly connected to the optical switch. For tasks that need to communicate externally, preferentially select the fragments whose size is greater than or equal to the number of nodes required for the task and there are nodes in the fragments directly connected to the optical switch.
6. The method for expanding the scale of the Torus network based on an optical switch as claimed in claim 1, wherein: When partitioning the complete supernode, for tasks that do not need to communicate externally, place the task nodes on the side far from the optical switch.
7. The method for expanding the scale of the Torus network based on an optical switch as described in claim 1, wherein: The optical switch establishes connections between ports according to traffic requirements, does not interfere with the running tasks during reconfiguration, and ensures communication between task nodes placed in different supernodes.
8. The method for expanding the scale of the Torus network based on an optical switch as claimed in claim 7, wherein: When task nodes cannot directly communicate through the optical switch, forward the traffic through other task nodes directly connected to the optical switch.
9. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute the method for expanding the scale of the Torus network based on the optical switch according to any one of claims 1-8.
10. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for expanding the scale of the Torus network based on the optical switch according to any one of claims 1-8.