Arithmetic processing unit, arithmetic processing program, and arithmetic processing method
Patent Information
- Application Number
- JP2025036824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2026-09-17
AI Technical Summary
【0013】 1つの側面では、ネットワークトポロジのスケーラビリティを向上させることができる。
Smart Images

Figure 2026148318000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an arithmetic processing device, an arithmetic processing program and an arithmetic processing method.
Background Art
[0002] Parallel computers such as massively parallel computers and GPGPU (General-Purpose computing on Graphics Processing Units) clusters are used as conventional computing infrastructure for scientific computation and simulation. Interconnect technology is essential for parallel computers to build large-scale systems and achieve high system performance. In recent years, AI (Artificial Intelligence) dedicated systems have been developed into parallel computers as their scale expands.
[0003] The configuration scale of a parallel computer depends on the node configuration, and the network configuration in the system also varies accordingly. In a system where large-scale fat-nodes are used to improve individual computing performance and the number of nodes is reduced accordingly, the scale of the network is also reduced. Conversely, in a system where a large number of thin-nodes with reduced individual configuration are used, the network scale becomes large. The number of parameters for AI has been increasing rapidly year by year, and in order to store these parameters in memory, individual nodes of AI dedicated systems tend to be further fat-noded. Fat-nodes require a wide bandwidth, and it is estimated that as bandwidth continues to increase, CPO (Co-Packaged Optics) rather than AOC (Active Optical Cable) will be required.
Prior Art Literature
Non-Patent Literature
[0004]
Non-Patent Literature 1
[0005] Because CPO has a high signal density, it has the advantage of fewer implementation constraints compared to using AOC. However, when CPO is adopted in an indirect network, CPO integration is also required on the switch side, which increases network costs.
[0006] Figure 1 shows Slim Fly, a network topology of a direct network in a conventional example.
[0007] Slim Fly, as shown in Figure 1, is proposed as a topology that enables direct networking and reduces costs, latency, and energy consumption by reducing the network diameter (e.g., Non-Patent Document 1). However, it imposes significant implementation constraints because it requires the complex wiring between local groups to be housed within the rack.
[0008] Figure 2 shows PolarFly, a network topology of a direct network in a conventional example.
[0009] PolarFly, as shown in Figure 2, is proposed as a topology with an even shorter diameter while improving on the implementation constraints compared to Slim Fly (for example, Non-Patent Document 2).
[0010] Each node uses a left-normalized vector, composed of three elements on a finite field, as its address, and connects nodes whose inner product of the vectors is zero. Due to the properties of left-normalized vectors, it is guaranteed that any two vectors are either orthogonal or share a common orthogonal vector. Therefore, any two nodes are connected within two hops. However, there is a problem of insufficient scalability. PolarFly constructed using left-normalized vectors on a finite field F7, as shown in Figure 2, connects 57 nodes. The required number of ports is 8 per node.
[0011] One aspect of this is the aim to improve the scalability of the network topology. [Means for solving the problem]
[0012] In one aspect, the arithmetic processing unit is a single arithmetic processing unit included in the network topology, comprising an arithmetic processing unit that performs arithmetic processing, and an input / output unit that performs optical communication with a plurality of other arithmetic processing units included in the network topology via a plurality of ports, wherein the arithmetic processing unit switches the network topology between a first network topology having a local topology within a global topology and a second network topology having only the local topology by switching the connection destinations of at least some of the plurality of ports in the input / output unit. [Effects of the Invention]
[0013] In one aspect, the scalability of the network topology can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] [Figure 1] FIG. 1 is a diagram showing Slim Fly, which is a direct network topology in a conventional example. [Figure 2] FIG. 2 is a diagram showing PolarFly, which is a direct network topology in a conventional example. [Figure 3] FIG. 3 is a diagram showing a first configuration example of an arithmetic processing device in a related example. [Figure 4] FIG. 4 is a diagram showing a second configuration example of an arithmetic processing device in a related example. [Figure 5] FIG. 5 is a diagram showing a third configuration example of an arithmetic processing device in a related example. [Figure 6] FIG. 6 is a diagram showing a configuration example of a first network topology in an arithmetic processing device according to an embodiment. [Figure 7] FIG. 7 is a diagram showing a configuration example of a second network topology in an arithmetic processing device according to an embodiment. [Figure 8] FIG. 8 is a diagram showing a configuration example of virtual channels for PolarFly and Hypercube according to an embodiment. [Figure 9] FIG. 9 is a flowchart explaining fault detour processing according to an embodiment. [Figure 10] FIG. 10 is a diagram explaining a detour example of fault detour processing according to an embodiment. [Figure 11] FIG. 11 is a table exemplifying costs for each transmission medium in a network topology. [Figure 12] FIG. 12 is a table exemplifying network costs and injection bandwidths in each of the network topologies shown in FIGS. 3 to 7. [Figure 13] FIG. 13 is a graph exemplifying the relationship between the number of nodes and bisection bandwidth in each of the network topologies shown in FIGS. 3 to 7. MODE FOR CARRYING OUT THE INVENTION
[0015] (A) Related Example Figure 3 is a diagram showing a first configuration example of the arithmetic processing device 6 in the related example.
[0016] The network topology to which the arithmetic processing device 6 shown in Figure 3 is applied may be referred to as a 6D Torus, and as indicated by reference sign A1, the number of coordinates in the X-axis direction is 24, the number of coordinates in the Y-axis direction is 24, and the number of coordinates in the Z-axis direction is 24.
[0017] Furthermore, each node indicated by reference sign A1, as in the network topology shown by reference sign A2, has 2 coordinates in the A-axis direction, 3 coordinates in the B-axis direction, and 2 coordinates in the C-axis direction.
[0018] Therefore, the 6D Torus shown in Figure 3 has dimensions of 2 × 3 × 2 × 24 × 24 × 24, has high scalability, and is suitable for HPC (High Performance Computing).
[0019] The arithmetic processing device 6 includes an xPU (x Processing Unit) 11 and a CPO 12 connected to the xPU 11.
[0020] Two ports in the A / C axis direction (in other words, electrical connections) are provided from the xPU 11. Furthermore, eight ports in the B+ / B-, X+ / X-, Y+ / Y-, and Z+ / Z- axis directions are provided from the CPO 12. Hereinafter, Figures 3 to 7 show an example where the CPO 12 has 8 optical fibers.
[0021] Figure 4 is a diagram showing a second configuration example of arithmetic processing devices 6a and 6b in the related example.
[0022] The network topology to which the arithmetic processing devices 6a and 6b shown in Figure 4 are applied may be referred to as Quad-rail 4D-Torus, and as indicated by reference sign B1, the number of coordinates in the X-axis direction is 24, the number of coordinates in the Y-axis direction is 24, and the number of coordinates in the Z-axis direction is 24.
[0023] Furthermore, each node shown by symbol B1 has 3 coordinates in the B-axis direction, as shown in the network topology shown by symbol B2.
[0024] Therefore, the Quad-rail 4D-Torus shown in Figure 4 has 3 x 24 x 24 x 24 dimensions, integrating the equivalent of four 6D-Torus nodes, resulting in a configuration with large memory capacity and suitable for AI.
[0025] The arithmetic processing unit 6a is comprised of four xPUs 11. Each xPU 11 is equipped with one CPO 12 and is connected to the other xPUs 11 by ports in the A / C axis direction.
[0026] The arithmetic processing unit 6a may have the configuration of the arithmetic processing unit 6b.
[0027] The arithmetic processing unit 6b comprises one xPU (x Processing Unit) 11 and four CPOs 12 connected to the xPU 11.
[0028] Each CPO12 is configured with eight ports in the B+ / B-, X+ / X-, Y+ / Y-, and Z+ / Z- axes. Therefore, the entire arithmetic processing unit 6b has 8 x 4 = 32 ports configured.
[0029] Figure 5 shows a third example configuration of the arithmetic processing unit in the related example.
[0030] The network topology to which the arithmetic processing unit 6c shown in Figure 5 is applied may be called an Oct-rail Fat-tree. The network topology shown by symbol E1 has 32 x 64 = 2048 nodes, and a total of 768 switches (SWs) in addition to 8 x 64 = 512 SWs directly connected to each node, and 8 x 32 = 256 SWs above it.
[0031] Therefore, the Oct-rail Fat-tree shown in Figure 5 is a 64-port switch, two-stage Fat-tree, with eight pins of Tera-PHY branched to form an eight-tiered Fat-tree. Note that 64 ports were deemed appropriate when the bandwidth of one port of the switch was set to 512 Gbps to match CPO12 (which is not enough to reach 128 ports).
[0032] The arithmetic processing unit 6c comprises an xPU11 and a CPO12 connected to the xPU11. Eight ports are configured on the CPO12.
[0033] [B] Embodiment An embodiment will be described below with reference to the drawings. However, the embodiment shown below is merely illustrative, and there is no intention to exclude various modifications or applications of techniques not explicitly shown in the embodiment. In other words, this embodiment can be implemented in various ways without departing from its spirit. Furthermore, each figure is not intended to represent only the components shown in the figure, but may include other functions, etc.
[0034] In the following diagrams, the same symbols indicate the same parts, so their explanations are omitted.
[0035] Figure 6 shows an example of the configuration of the first network topology in the arithmetic processing unit 1 in the embodiment.
[0036] PolarFly, shown in Figure 2, is a direct network with a short diameter and few implementation constraints, but it lacks scalability.
[0037] Therefore, let's consider a topology that combines PolarFly and Hypercube. That is, we treat Hypercube as a local group and connect them with PolarFly.
[0038] For HPC applications, it's possible to configure a POD that leverages PolarFly's short diameter, while for AI applications, it's possible to utilize Hypercube's tight coupling. The POD is envisioned to have a configuration of 64 nodes or less using Fat-nodes.
[0039] However, considering scalability for HPC, the entire system is envisioned to be on the scale of several thousand nodes. Therefore, we consider combining PolarFly on F7 with a 7D Hypercube. Since 57 local groups of 128 nodes are connected, the entire system will consist of 7296 nodes.
[0040] The first network topology shown in Figure 6 is often referred to as PolarFly+, and as shown by symbol C1, an 8D-Hypercube (symbol C12) is configured at each node of PolarFly (symbol C11).
[0041] The arithmetic processing unit 1 comprises an xPU 11 and two CPOs 12 connected to the xPU 11.
[0042] xPU11 is an example of a processing unit and performs calculations.
[0043] CPO12 is an example of an input / output unit, and it performs optical communication with multiple other xPU11 units included in the network topology via multiple ports.
[0044] In the first network topology shown in Figure 6, of the two CPO12s, one is configured with eight ports for PolarFly and the other with eight ports for Hypercube.
[0045] While PolarFly on F7 requires 8 ports per node, a 7D Hypercube uses 7 ports, for a total of 15 ports. However, considering that Tera-PHY has 8 cores, it would be sufficient to provide 8 ports for the Hypercube as well, with 1 port reserved as an expansion spare. In other words, two Tera-PHYs per node would suffice, allocating 16 fibers to each of the 16 ports.
[0046] Figure 7 shows an example of the configuration of the second network topology in the arithmetic processing unit 1 in the embodiment.
[0047] The second network topology shown by symbol D1 in Figure 7 is often referred to as PolarFly+ AI POD, and as shown by symbol D11, it consists of tightly coupled PODs using multiple Hypercubes that utilize all ports.
[0048] In the second network topology shown in Figure 7, both CPO12s are configured as ports for the Hypercube. That is, 8 x 2 = 16 ports are configured for the Hypercube.
[0049] xPU11 may be switchable between the first network topology (PolarFly+) shown in Figure 6 and the second network topology (PolarFly+ AI POD) shown in Figure 7. xPU11 switches one of the two CPO12s for PolarFly or Hypercube, and fixes the other for Hypercube.
[0050] In other words, xPU11 switches the destination of at least some of the multiple ports in CPO12. This causes xPU11 to switch the network topology between a first network topology having a local topology (Hypercube) within a global topology (PolarFly+) and a second network topology having only a local topology (Hypercube).
[0051] xPU11 may be, for example, one of the following: CPU, MPU, DSP, ASIC, PLD, or FPGA. Alternatively, xPU11 may be a combination of two or more of the following: CPU, MPU, DSP, ASIC, PLD, and FPGA. Note that CPU is an abbreviation for Central Processing Unit, MPU is an abbreviation for Micro Processing Unit, DSP is an abbreviation for Digital Signal Processor, and ASIC is an abbreviation for Application Specific Integrated Circuit. Also, PLD is an abbreviation for Programmable Logic Device, and FPGA is an abbreviation for Field Programmable Gate Array.
[0052] Figure 8 shows an example of the configuration of the virtual channels of PolarFly and Hypercube in the embodiment.
[0053] Routing on PolarFly+ is divided into movement on PolarFly and movement on Hypercube.
[0054] Movement on PolarFly can occur up to two times. In Figure 8, code E1 represents how the local groups composed of Hypercube are connected in a ring shape, and deadlocks can be avoided by separating the virtual channels (VCs) for the first and second hops on PolarFly. However, since a response is returned for each request, two VCs are required for both the request and the response.
[0055] Movement on the Hypercube can occur up to three times. PolarFly+ can also be considered as PolarFly connected via the Hypercube. In Figure 8, code E2 represents how PolarFly connects in a ring shape via the Hypercube. To avoid deadlocks, the VC should be divided into before the movement on the PolarFly, after one hop on the PolarFly, and after two hops on the PolarFly; therefore, three VCs are needed for both the Request and the Response.
[0056] In other words, when the minimum hop in the global topology is n (where n is a natural number), (n+1) × 2 virtual channels are required for routing in the local topology and n × 2 virtual channels are required for routing in the global topology. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with the multiple processing units 1.
[0057] The fault bypass process in the embodiment configured as described above will be explained with reference to Figure 10, following the flowchart (steps S1 to S11) shown in Figure 9. Figure 10 is a diagram illustrating an example of fault bypass in the embodiment.
[0058] In the failure handling process (steps S1-S2), when a failure occurs (step S1), the management server (for example, xPU11 on any node) notifies all nodes of the location of the failure (step S2).
[0059] During the transmission process (steps S3 to S5), the xPU11 of the source node calculates the route to the destination (step S3).
[0060] The xPU11 on the source node determines if there is a fault on the route (step S4).
[0061] If there are no faults along the route (see No Route in step S4), the source node's xPU11 writes routing information to the packet (step S5), executes transmission, and the fault bypass process is completed.
[0062] On the other hand, if there is a fault in the route (see Yes route in step S4), the process proceeds to the detour process (steps S6-S11).
[0063] The source node's xPU11 performs a brute-force search of movement on the Hypercube, with PolarFly as the minimum hop (step S6). The brute-force search of movement on the Hypercube includes minimum detours, VC switching detours, and non-minimum detours, as shown by the symbol F1 in Figure 10. In Figure 10, "X" represents a fault location.
[0064] In other words, if an xPU11 fails on the route to the destination xPU11 among multiple xPU11s, the xPU11 searches for a first detour route that passes through the local topology once each.
[0065] The xPU11 of the source node determines whether bypassing is possible (step S7).
[0066] If a detour is possible (see the Yes route in step S7), the process proceeds to step S11.
[0067] If a detour is not possible (see No Route in step S7), the xPU11 on the source node determines whether the source node and the destination node are on the same Hypercube (step S8).
[0068] If the source node and destination node are not on the same Hypercube (see No route in step S8), the fault bypass process stops and terminates.
[0069] On the other hand, if the source node and destination node are on the same Hypercube (see the Yes route in step S8), the xPU11 of the source node searches all possible paths to return to PolarFly by moving across the Hypercube, moving PolarFly, and then moving across the Hypercube again (step S9). This search for all possible paths to return to PolarFly by moving across the Hypercube, moving PolarFly, and then moving across the Hypercube again includes the global detour shown as indicated by symbol F2 in Figure 10.
[0070] In other words, if the first detour route is unavailable and one xPU11 and the destination xPU11 are located in the same local topology, the xPU11 searches for a second detour route that returns to the same local topology via other local topologies.
[0071] The xPU11 of the source node determines whether bypassing is possible (step S10).
[0072] If a detour is not possible (see Route No. in step S10), the fault detour process stops and terminates.
[0073] On the other hand, if a detour is possible (see Yes route in step S10), the xPU11 of the source node selects the shortest hop from the routes (step S11), and the process proceeds to step S5.
[0074] [C] Effect Figure 11 is a table illustrating the cost for each transmission medium in a network topology. Figure 12 is a table illustrating the network cost and injection bandwidth in each network topology shown in Figures 3 to 7.
[0075] As shown in Figure 11, assuming the cost of electrical cables is 1, the cost of switches per port is 3.3, the cost of NICs (Network Interface Cards) per port is 3.3, and the cost of CPOs is 32.5.
[0076] Figure 12 shows the network topology costs for each of the network topologies shown in Figures 3 to 7, when the costs of each transmission medium shown in Figure 11 are applied. The injection bandwidth shown in Figure 12 is an example where the CPO12 bandwidth is 8 cores and 4 Tbps.
[0077] The network cost of 6D-Torus shown in Figure 3 is 52.3, and the injection bandwidth is 3 Tbps (virtual 3D 6 directions x 512 Gbps).
[0078] The network cost of the Quad-rail 4D-Torus shown in Figure 4 is 209.2, and the injection bandwidth is 12 Tbps (four times that of 6D-Torus).
[0079] The network cost of the Oct-rail Fat-tree shown in Figure 5 is 235.6, and the injection bandwidth is 4 Tbps (total bandwidth of one CPO).
[0080] The network cost of PolarFly+ shown in Figures 6 and 7 is 91.4, and the injection bandwidth is 4 Tbps (8-way x 512 Gbps for 8D-Hypercube).
[0081] Thus, when comparing the ratio of injection bandwidth to cost, 6D-Torus, Quad-rail 4D-Torus, and PolarFly+ are comparable, while Oct-rail Fat-tree is more expensive.
[0082] Figure 13 is a graph illustrating the relationship between the number of nodes and bisection bandwidth in each network topology shown in Figures 3 to 7.
[0083] In Figure 13, the BBW (Bisection Bandwidth) for 6D-Torus and QR 4D-Torus (Quad-rail 4D-Torus) is plotted with configurations where the X, Y, and Z axes have 24, 16, and 8 coordinates, respectively.
[0084] PolarFly+ differs depending on whether it's for HPC (PolarFly+) or AI (PolarFly+ AI POD). For HPC, PolarFly is used, but for AI, the 8 ports reserved for PolarFly are switched to Hypercube ports, and all 16 ports on each node are used to configure the Hypercube POD.
[0085] For PolarFly+, the BBW is plotted for configurations with Hypercube in 8, 6, and 4 dimensions. For PolarFly+ AI POD, the BBW is plotted for a 256-node POD with 2 ports allocated to 1 dimension, a 16-node POD with 4 ports allocated to 1 dimension, a 4-node POD with 8 ports allocated to 1 dimension, and a 2-node POD with all 16 ports allocated to 1 dimension.
[0086] For OR Fat-trees (Oct-rail Fat-trees), the BBW (Base-by-Way) is plotted for configurations of 2048 nodes, 512 nodes, and 128 nodes. For comparison, the BBW of Fat-trees with the same number of nodes is also plotted.
[0087] While 6D-Torus and QR 4D-Torus have a proportional relationship between cost and IBW (Injection Bandwidth), 6D-Torus is suitable for HPC applications due to its thin-node multi-node configuration, whereas 4D-Torus is suitable for AI POD configurations because it combines four 6D-Torus nodes into one fat-node configuration connected with high-bandwidth quad-rails. Quad-plane 4D-Torus uses four Tera-PHYs per node, resulting in a higher BBW for the same number of nodes compared to 6D-Torus.
[0088] While BBW offers the highest network cost for OR Fat-tree architectures, its network costs are excessively high. A standard Fat-tree without redundancy reduces network costs to 1 / 8, but BBW also achieves this reduction. Furthermore, it can only connect up to 2048 nodes, lacking sufficient scalability for HPC applications.
[0089] PolarFly+ offers a BBW (Base-by-Web) equivalent to QR 4D-Torus and Fat-tree. However, it has half the network cost of QR 4D-Torus and higher scalability than Fat-tree. Because PolarFly+ AI POD is a multiplexed Hypercube, it can achieve a higher BBW than Fat-tree at the same node level without changing network costs.
[0090] In other words, PolarFly+ AI POD is considered appropriate for AI applications with 256 nodes or less, while PolarFly+ is considered appropriate for HPC applications with 14,592 nodes or less.
[0091] According to the arithmetic processing unit 1 in the above-described embodiment, the following effects can be achieved, for example.
[0092] xPU11 switches the network topology between a first network topology, which has a local topology within the global topology, and a second network topology, which has only a local topology, by switching the destination of at least some of the ports in CPO12.
[0093] This improves the scalability of the network topology and also improves the cost-to-injection bandwidth ratio.
[0094] When the minimum hop in the global topology is n (where n is a natural number), (n+1) × 2 virtual channels are required for routing in the local topology, and n × 2 virtual channels are required for routing in the global topology. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with multiple processing units.
[0095] This makes it possible to avoid deadlocks caused by routing.
[0096] If a failure occurs in the route to the destination xPU11 among multiple xPU11s, xPU11 searches for a first detour route that passes through the local topology once each.
[0097] This allows the system to search for the shortest possible detour route in the event of a malfunction or other failure.
[0098] If the first detour route is unavailable and both the xPU11 and the destination xPU11 are located in the same local topology, the xPU11 searches for a second detour route that returns to the same local topology via other local topologies.
[0099] This improves the likelihood of finding an alternative route even when a shorter alternative route is unavailable.
[0100] [D] Other The disclosed technology is not limited to the embodiments described above and can be implemented in various modifications without departing from the spirit of this embodiment. Each configuration and process of this embodiment can be selected or combined as needed.
[0101] [E] Note In addition to the embodiments described above, the following additional information is disclosed.
[0102] (Note 1) A single processing unit included in the network topology, A processing unit that performs calculations, An input / output unit that performs optical communication via multiple ports with multiple other processing units included in the network topology, Equipped with, The arithmetic processing unit switches the network topology between a first network topology having a local topology within a global topology and a second network topology having only the local topology by switching the connection destinations of at least some of the multiple ports in the input / output unit. Processing unit.
[0103] (Note 2) When the minimum hop within the global topology is n (where n is a natural number), the routing of the local topology requires (n+1) × 2 virtual channels, and the routing of the global topology requires n × 2 virtual channels. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with the multiple processing units. The arithmetic processing unit described in Appendix 1.
[0104] (Note 3) The arithmetic processing unit searches for a first detour route that passes through the local topology once each, in the event that a failure occurs in the route to the destination arithmetic processing unit among the plurality of arithmetic processing units. The arithmetic processing unit described in Appendix 1 or 2.
[0105] (Note 4) The arithmetic processing unit searches for a second detour route that returns to the same local topology from the same local topology via other local topologies, when the first detour route cannot be used and the first arithmetic processing unit and the destination arithmetic processing unit are located in the same local topology. The arithmetic processing unit described in Appendix 3.
[0106] (Note 5) On a computer in one node included in the network topology, The network topology is switched between a first network topology having a local topology within a global topology and a second network topology having only the local topology by switching the connection destination of at least some of the ports that perform optical communication with multiple other nodes included in the network topology. A program that performs calculations and processes.
[0107] (Note 6) When the minimum hop within the global topology is n (where n is a natural number), the routing of the local topology requires (n+1) × 2 virtual channels, and the routing of the global topology requires n × 2 virtual channels. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with the multiple nodes. The arithmetic processing program described in Appendix 5, which causes the computer to perform the processing.
[0108] (Note 7) If a failure occurs in the route to the destination node among the aforementioned multiple nodes, a first detour route is searched for that passes through the local topology once each. A calculation processing program as described in Appendix 5 or 6, which causes the computer to perform the processing.
[0109] (Note 8) If the first detour route cannot be used and the first node and the destination node are located in the same local topology, a second detour route is searched for that returns to the same local topology via other local topologies. The arithmetic processing program described in Appendix 7, which causes the computer to perform the processing.
[0110] (Note 9) A computer on one node included in the network topology, The network topology is switched between a first network topology having a local topology within a global topology and a second network topology having only the local topology by switching the connection destination of at least some of the ports that perform optical communication with multiple other nodes included in the network topology. A method of performing calculations or processing.
[0111] (Note 10) When the minimum hop within the global topology is n (where n is a natural number), the routing of the local topology requires (n+1) × 2 virtual channels, and the routing of the global topology requires n × 2 virtual channels. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with the multiple nodes. The calculation processing method described in Appendix 9, wherein the computer performs the processing.
[0112] (Note 11) If a failure occurs in the route to the destination node among the aforementioned multiple nodes, a first detour route is searched for that passes through the local topology once each. The arithmetic processing method described in Appendix 9 or 10, wherein the computer performs the processing.
[0113] (Note 12) If the first detour route cannot be used and the first node and the destination node are located in the same local topology, a second detour route is searched for that returns to the same local topology via other local topologies. The calculation processing method described in Appendix 11, wherein the computer performs the processing. [Explanation of Symbols]
[0114] 1,6,6a,6b,6c: Arithmetic processing units 11: xPU 12: CPO
Claims
1. A single processing unit included in the network topology, A processing unit that performs calculations, An input / output unit that performs optical communication via multiple ports with multiple other processing units included in the network topology, Equipped with, The arithmetic processing unit switches the network topology between a first network topology having a local topology within a global topology and a second network topology having only the local topology by switching the connection destinations of at least some of the multiple ports in the input / output unit. Processing unit.
2. When the minimum hop within the global topology is n (where n is a natural number), the routing of the local topology requires (n+1) × 2 virtual channels, and the routing of the global topology requires n × 2 virtual channels. Therefore, each of the multiple ports has (n+1) × 2 virtual channels and performs optical communication with the multiple processing units. The arithmetic processing device according to claim 1.
3. The arithmetic processing unit searches for a first detour route that passes through the local topology once each, in the event that a failure occurs in the route to the destination arithmetic processing unit among the plurality of arithmetic processing units. The arithmetic processing apparatus according to claim 1 or 2.
4. The arithmetic processing unit searches for a second detour route that returns to the same local topology from the same local topology via other local topologies, when the first detour route cannot be used and the first arithmetic processing unit and the destination arithmetic processing unit are located in the same local topology. The arithmetic processing device according to claim 3.
5. On a computer in one node included in the network topology, By switching the connection destination of at least some of the ports that perform optical communication with multiple other nodes included in the network topology, the network topology is switched between a first network topology having a local topology within the global topology and a second network topology having only the local topology. A program that performs calculations and processes.
6. A computer on one node included in the network topology, By switching the connection destination of at least some of the ports that perform optical communication with multiple other nodes included in the network topology, the network topology is switched between a first network topology having a local topology within the global topology and a second network topology having only the local topology. A method of performing calculations or processing.