Computer system
Patent Information
- Application Number
- JP2025556414
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-11-06
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-15
AI Technical Summary
Conventional computer systems face a trade-off between the number of connections of Non-Transparent Bridges (NTBs) and bandwidth, where increasing the number of devices connected reduces the bandwidth per device.
The computer system incorporates an optical line switch, a data flow monitoring unit, a data flow analysis unit, an optimization decision unit, and an optimization execution unit to dynamically manage connections and optimize data flow between PCIe devices, thereby reducing the trade-off between the number of connections and bandwidth.
This approach allows for improved latency and lower power consumption compared to using only NTB switches, while also enabling dynamic expansion and reduction of PCIe bandwidth for each application, optimizing data flow, and enhancing power efficiency.
Abstract
Description
Computer Systems
[0001] The present invention relates to a computer system, and more particularly to a bus connection technique in a computer's internal architecture.
[0002] In computer systems, a Non-Transparent Bridge (NTB) is a technology that allows two or more independent systems to communicate with each other while maintaining their independence. Simply put, it allows two different computer systems to share resources and communicate without understanding the details of each other's systems.
[0003] The NTB concept is primarily used in PCIe (Peripheral Component Interconnect Express) technology, a high-speed serial computer expansion bus standard designed to replace older bus standards such as PCI.
[0004] In the context of PCIe, NTB enables communication between two separate PCIe hierarchies or domains. A standard PCIe bridge is called "opaque" because neither system can directly see or interfere with the other's memory or resources. NTB achieves this by providing separate address spaces on each side of the bridge and managing the translation and routing of transactions between these address spaces, allowing direct memory access (DMA) transfers and interrupts between the systems, enabling efficient communication and resource sharing.
[0005] As mentioned above, NTB isolates the systems on both sides and allows each system to see only what it is authorized to see. An NTB switch extends the NTB concept to multiple systems. Each system is connected to the NTB switch via an NTB adapter. The NTB switch manages communication between the systems. This allows each system to communicate with other systems without being aware of the details of the other systems or having direct access to their resources. The NTB switch ensures that communication between systems is managed effectively and that each system can only access the resources that are expected of it. The core of an NTB switch is a PCIe switch, and it can be implemented by operating a PCIe switch in NTB mode.
[0006] The configuration of the system disclosed in Non-Patent Document 1 is shown in Fig. 9. In the example of Fig. 9, 12 PCIe devices 1-1 to 1-12 are connected to an NTB switch 2. The PCIe devices 1-1 to 1-12 include host devices equipped with a CPU (Central Processing Unit) and an OS (Operating System), and GPU (Graphics Processing Unit) devices used by the host devices for processing.
[0007] Although simplified in FIG. 9, each of the devices 1-1 to 1-12 has an NTB adapter. The NTB adapter can communicate with the NTB switch 2 at a bandwidth equal to or less than the PCIe lanes of the devices 1-1 to 1-12. The NTB adapter has multiple ports. A PCIe lane is assigned to each port. A lane is a data transfer path, and more specifically, a bundle of data transfer paths for transmission and reception. In the following description, the number of lanes will be represented by "x number." For example, x16 indicates that the number of lanes is 16.
[0008] When the CPUs or GPUs in Devices 1-1 to 1-12 are ×16 lanes and the NTB adapter has four ports P1 to P4, a PCIe4 port is assigned to one port. The NTB switch 2 has a plurality of ports. By connecting the ports of the NTB switch 2 and the ports of the NTB adapters of Devices 1-1 to 1-12, PCIe communication becomes possible. At this time, PCIe communication can be performed without connecting all four ports to the NTB switch 2. The example in FIG. 9 shows an example in which Devices 1-1 to 1-12 are communicating using two ports, P1 and P2, of the NTB adapter.
[0009] The problem of the conventional system is that there is a trade-off between the connection bandwidth and the number of connections. For example, in the example of FIG. 9, although 12 devices 1-1 to 1-12 are connected to the NTB switch 2, in order to increase the number of devices to 24, it is necessary to reduce the number of connection ports per device (NTB adapter). Therefore, when a large number of devices are connected to the NTB switch 2, the bandwidth per device becomes narrow.
[0010] Jonas Markussen, Lars Bjorlykke Kristiansen, Pal Halvorsen, Halvor Kielland-Gyrud, Hakon Kvale Stensland, and Carsten Griwodz, “SmartIO: Zero-overhead Device Sharing through PCIe Networking”, ACM Transactions on Computer Systems, Vol. 38, No. 1-2, Article. 2, pp 1-78, 2021, <https: / / doi.org / 10.1145 / <3462545>
[0011] The present invention has been made to solve the above problems, and an object thereof is to provide a computer system capable of alleviating the trade-off between the number of NTB connections and the bandwidth.
[0012] The computer system of the present invention is characterized by comprising: a plurality of PCIe devices each having an NTB adapter for communication with another device; an NTB switch capable of selectively connecting the plurality of PCIe devices; an optical line switch capable of selectively connecting the plurality of PCIe devices; a data flow monitoring unit configured to monitor a data flow between the PCIe devices via at least one of the NTB switch and the optical line switch; a data flow analysis unit configured to identify communication between the PCIe devices that is a bottleneck in the data flow based on data flow information collected by the data flow monitoring unit; an optimization determination unit configured to determine a connection plan for optimizing the data flow based on analysis results by the data flow analysis unit; and an optimization execution unit configured to instruct the PCIe devices and the optical line switch to change the connection between the PCIe devices in accordance with the connection plan determined by the optimization determination unit.
[0013] According to the present invention, by providing an optical line switch, a data flow monitor, a data flow analyzer, an optimization decision unit, and an optimization execution unit, it is possible to ease the trade-off between the number of NTB connections and bandwidth. Also, by using an optical line switch for data transfer, it is possible to achieve improved latency and lower power consumption compared to when an NTB switch is used.
[0014] FIG. 1 is a block diagram showing the configuration of a computer system according to a first embodiment of the present invention. FIG. 2 is a block diagram showing the configuration of a host device according to the first embodiment of the present invention. FIG. 3 is a block diagram showing the configuration of a GPU device according to the first embodiment of the present invention. FIG. 4 is a diagram explaining a PCIe cable used in the first embodiment of the present invention. FIG. 5 is a functional block diagram of an OCS controller according to the first embodiment of the present invention. FIG. 6 is a flowchart explaining the operation of the computer system according to the first embodiment of the present invention. FIG. 7 is a block diagram showing the configuration of a computer system according to a second embodiment of the present invention. FIG. 8 is a diagram explaining a PCIe cable used in the second embodiment of the present invention. FIG. 9 is a block diagram showing the configuration of a conventional computer system.
[0015] [First Embodiment] An embodiment of the present invention will now be described with reference to the drawings. Fig. 1 is a block diagram showing the configuration of a computer system according to a first embodiment of the present invention. The computer system includes PCIe devices 1-1 to 1-12, an NTB switch 2, an optical circuit switch (OCS) 3, and an OCS controller 4 that controls the OCS 3.
[0016] As described above, the PCIe devices 1-1 to 1-12 include a host device, a GPU device, and the like. The configuration of a host device is shown in Fig. 2. The host device includes a CPU 10, a memory 11, an NTB adapter 12, and a NIC (Network Interface Card) 13. The CPU 10 executes the following processes in accordance with a program stored in the memory 11. The CPU 10 and the NTB adapter 12 are connected by a x16 lane PCIe bus.
[0017] The configuration of a GPU device is shown in Fig. 3. The GPU device includes a CPU 10, a memory 11, an NTB adapter 12, a NIC 13, and a GPU 14. The CPU 10 and the GPU 14, and the CPU 10 and the NTB adapter 12 are connected by a x16 lane PCIe bus.
[0018] As in the past, the NTB adapter 12 of each of the devices 1-1 to 1-12 has four ports P1 to P4. To simplify the description, in Fig. 1, PCIe cables 5 are shown only for the devices 1-6 and 1-12, but in this embodiment, all of the ports P1 to P4 of each of the devices 1-1 to 1-12 are connected one-to-one to ports of the NTB switch 2 by the PCIe cables 5, and at the same time, are connected one-to-one to ports of the OCS 3 by the PCIe cables 5.
[0019] 4 , the PCIe cable 5 is a cable that branches into two from a connector 50 that is connected to a port of the NTB adapter 12, one being an electrical cable 51 and the other being an optical cable 52. A connector 53 at the end of the electrical cable 51 is connected to a port of the NTB switch 2, and a connector 54 at the end of the optical cable 52 is connected to a port of the OCS 3. A photoelectric conversion device is built into the connector 50 that is connected to the port of the NTB adapter 12. The photoelectric conversion device converts electrical signals from the NTB adapter 12 to the OCS 3 into optical signals and sends them to the optical cable 52, and converts optical signals received from the OCS 3 via the optical cable 52 into electrical signals and sends them to the NTB adapter 12.
[0020] By adopting the above-described dual connection, it is possible to utilize the advantages of both the NTB switch 2 and the OCS 3. Since dual connection is performed for each port of the NTB adapter 12, the assembly of the PCIe cable 5 for each device includes a connector 50 connected to a port of the NTB adapter 12, a connector 53 connected to a port of the NTB switch 2, a connector 54 connected to a port of the OCS 3, four electrical cables 51, four optical cables 52, and four photoelectric conversion devices.
[0021] The NTB switch 2 provides dynamic and agile PCIe communication and is suitable for control and data transfer tasks. The OCS3 is a type of switch that manages data traffic using optical signals. Although the OCS3 cannot adapt as dynamically as the NTB switch 2 for PCIe communication, it is effective in handling large amounts of data. Even if the number of ports connected to the NTB switch 2 decreases, the OCS3 can be used to supplement the bandwidth.
[0022] Therefore, to manage communications between NTBs, an OCS controller 4 is employed, which aggregates connection information between NTBs. The OCS controller 4 dynamically adjusts the configuration of the OCS 3 to optimize data flow. The OCS controller 4 includes a CPU 40, a memory 41, and a NIC 42. The CPU 40 executes the following processes according to a program stored in the memory 41. For simplicity's sake, only device 1-6 is connected to the OCS controller 4 in FIG. 1, but each of devices 1-1 to 1-12 can communicate with the CPU 40 of the OCS controller 4 via the NIC 13, the network, and the NIC 42 of the OCS controller 4.
[0023] 5 is a functional block diagram of the OCS controller 4. The CPU 40 of the OCS controller 4 executes the following processes in accordance with a program stored in the memory 41, and functions as a data flow monitor 400, a data flow analyzer 401, an optimization decision unit 402, and an optimization execution unit 403.
[0024] 6 is a flowchart illustrating the operation of the computer system of this embodiment. First, the CPU 10 of the host device establishes communication with another host device or with an accelerator device such as a GPU device via the NTB switch 2 (step S100 in FIG. 6).
[0025] The data flow monitor 400 of the OCS controller 4 continuously monitors the data flow between devices via at least one of the NTB switch 2 and the OCS 3 (step S101 in FIG. 6). Items monitored include the data transfer pattern (whether the amount of data transferred fluctuates instantaneously or periodically), the load (amount of data transferred) on each port of the device, the effective bandwidth of each port, and overall system performance. An example of an item indicating overall system performance is the amount of data transferred per second by application software executed by the CPU 10 of the host device. The processing of step S101 can be achieved by the data flow monitor 400 making a data flow inquiry (command transmission) to the CPU 10 of each host device.
[0026] Next, the data flow analysis unit 401 of the OCS controller 4 identifies communications between devices that are bottlenecks in the data flow and communications between devices that may become bottlenecks based on the data flow information collected by the data flow monitoring unit 400 (step S102 in FIG. 6). The data flow analysis unit 401 performs analyses such as predicting future data transfer patterns based on past data transfer patterns, identifying peak times for data transfer volume at each port of the device, predicting peak times for data transfer volume, and calculating the average and peak values of effective bandwidth usage at each port.
[0027] Through such analysis, the data flow analysis unit 401 can grasp the current state of how data is flowing within the computer system, as well as predict how data will flow within the system, and identify communications between devices that are bottlenecks in the data flow and communications between devices that have the potential to become bottlenecks.
[0028] The optimization determination unit 402 of the OCS controller 4 determines a connection plan for optimizing the data flow based on the analysis results by the data flow analysis unit 401 (step S103 in FIG. 6 ). The optimization determination unit 402 determines a connection plan to detour, via the OCS 3, a portion of the data flow between devices that is a bottleneck and a portion of the data flow between devices that may become a bottleneck.
[0029] At this time, the optimization determination unit 402 balances the load between the NTB switch 2 and the OCS 3, utilizing the advantages of both the NTB switch 2 and the OCS 3 to maximize system performance. Specifically, the amount of data transfer per NTB port of application software executed by the CPU 10 of the host device is maximized. The connection plan determined by the optimization determination unit 402 includes information on ports to be released among the ports in use of the NTB adapter 12 of each device, information on ports to be released among the ports in use of the NTB switch 2, information on ports to be released among the ports in use of the OCS 3, information on ports of the NTB adapter 12 of each device that will be newly used for connection with the OCS 3, information on ports of the OCS 3 that will be newly used for connection with the device, information on ports of the NTB adapter 12 of each device that will be newly used for connection with the NTB switch 2, and information on ports of the NTB switch 2 that will be newly used for connection with the device.
[0030] Priorities may be set in advance for data flows between devices, and data flows with priorities lower than a threshold value may be connected via the NTB switch 2, while data flows with priorities higher than the threshold value may be connected via the OCS 3. Furthermore, data flows between devices that have changing destinations may be connected via the NTB switch 2, while data flows with fixed destinations may be connected via the OCS 3. Furthermore, data flows with the same destination may be divided into two, with one connected via the NTB switch 2 and the other connected via the OCS 3.
[0031] In the above example, we have described an example in which a portion of the data flow that passes through NTB switch 2 is diverted via OCS 3, but it is also possible to determine a connection plan in which a portion of the data flow that passes through OCS 3 is diverted via NTB switch 2.
[0032] The optimization execution unit 403 of the OCS controller 4 executes optimization in accordance with the connection plan determined by the optimization determination unit 402, and changes the connection of the computer system (step S104 in FIG. 6 ). Specifically, the optimization execution unit 403 instructs the CPU 10 of each device via the NIC 42 to release the ports specified in the connection plan among the ports in use of the NTB adapter 12, release the ports specified in the connection plan among the ports in use of the NTB switch 2, and connect the ports specified in the connection plan among the ports of the NTB adapter 12 to the OCS 3. In response to this instruction, the CPU 10 of the device changes the port assignment so that the ports specified by the optimization execution unit 403 among the ports in use of the NTB adapter 12 are released and the ports specified by the optimization execution unit 403 among the ports of the NTB adapter 12 are connected to the OCS 3. The CPU 10 of the device also sets the NTB switch 2 to release the ports specified by the optimization execution unit 403 among the ports of the NTB switch 2.
[0033] Furthermore, when a connection plan is determined that diverts a portion of the data flow via the OCS 3 via the NTB switch 2, the optimization execution unit 403 instructs the CPU 10 of the device to release the ports specified in the connection plan among the ports in use of the NTB adapter 12 of each device, connect the ports specified in the connection plan among the ports of the NTB adapter 12 of each device to the NTB switch 2, and connect the ports specified in the connection plan among the ports of the NTB switch 2 to the device. In response to this instruction, the CPU 10 of the device changes the port assignment so that the port specified by the optimization execution unit 403 among the ports of the NTB adapter 12 is released and the port specified by the optimization execution unit 403 among the ports of the NTB adapter 12 is connected to the NTB switch 2. The CPU 10 of the device also sets the NTB switch 2 so that the port specified by the optimization execution unit 403 among the ports of the NTB switch 2 is connected to its own device.
[0034] Furthermore, the optimization execution unit 403 outputs a control signal to the OCS 3 so as to connect a port of the OCS 3 specified in the connection plan to the device. In response to this control signal, the OCS 3 changes its own configuration so as to connect a port of the OCS 3 specified by the optimization execution unit 403 to the specified device. Furthermore, when a connection plan is determined in which a portion of the data flow passing through the OCS 3 is detoured via the NTB switch 2, the optimization execution unit 403 changes its own configuration so as to release a port specified by the optimization execution unit 403 from among the ports of the OCS 3 that are in use.
[0035] In this way, the data flow path and the bandwidth usage of the port of the device's NTB adapter 12 are changed. The processes of steps S101 to S104 are repeated until the computer system is stopped (YES in step S105 in FIG. 6). In this way, the system connection is always optimized.
[0036] As described above, the computer system of this embodiment dynamically adapts to changes in data flow and optimizes performance by utilizing the unique features of both the NTB switch 2 and the OCS 3. This allows the computer system to efficiently and effectively handle both control and large-volume data communications.
[0037] In this embodiment, by using the OCS 3 and the OCS controller 4, it is possible to alleviate the trade-off between the number of NTB connections and bandwidth. Furthermore, by using the OCS 3 for data transfer, it is possible to achieve improved latency and lower power consumption compared to when the NTB switch 2 is used. Furthermore, in this embodiment, by diverting at least a portion of the data flow of an application via the OCS 3, it is possible to dynamically expand or contract the PCIe bandwidth for each application, thereby optimizing the PCIe flow and improving power efficiency. Furthermore, in this embodiment, contention that occurs in the NTB switch 2 is less likely to occur.
[0038] 7 is a block diagram showing the configuration of a computer system according to a second embodiment of the present invention. In the computer system of this embodiment, an OCS 3-1 is arranged between PCIe devices 1-1 to 1-6 and an NTB switch 2, and an OCS 3-2 is arranged between PCIe devices 1-7 to 1-12 and the NTB switch 2.
[0039] The configurations of the PCIe devices 1-1 to 1-12 and the OCS controller 4 are the same as those in the first embodiment. For the sake of simplicity, only the device 1-6 is connected to the OCS controller 4 in Fig. 7, but each of the devices 1-1 to 1-12 can communicate with the CPU 40 of the OCS controller 4 via the NIC 13, the network, and the NIC 42 of the OCS controller 4.
[0040] In this embodiment, all of the ports P1 to P4 of the devices 1-1 to 1-6 are connected one-to-one to ports of the OCS 3-1 by PCIe cables 6-1, and all of the ports P1 to P4 of the devices 1-1 to 1-12 are connected one-to-one to ports of the OCS 3-2 by PCIe cables 6-2. As shown in Figure 8, the PCIe cables 6-1 and 6-2 have optical cables 61 connected to connectors 60 that are connected to ports of the NTB adapter 12, and connectors 62 at the ends of the optical cables 61 are connected to ports of the OCSs 3-1 and 3-2. The connectors 60 that are connected to ports of the NTB adapter 12 have photoelectric conversion devices built in. The photoelectric conversion device converts electrical signals from the NTB adapter 12 to the OCSs 3-1 and 3-2 into optical signals and sends them to the optical cable 61, and converts optical signals received from the OCSs 3-1 and 3-2 via the optical cable 61 into electrical signals and sends them to the NTB adapter 12.
[0041] OCS 3-1 and NTB switch 2 are connected by a PCIe cable 7-1, and OCS 3-2 and NTB switch 2 are connected by a PCIe cable 7-2. PCIe cables 7-1 and 7-2 are similar to PCIe cables 6-1 and 6-2 shown in Fig. 8, with connector 60 connected to a port of NTB switch 2 and connector 62 connected to ports of OCSs 3-1 and 3-2. OCS 3-1 and OCS 3-2 are connected by an optical cable 8.
[0042] The processing flow of the computer system is the same as in the first embodiment, and will be described using the reference numerals in FIG. 6 . The processing in steps S101 to S103 is the same as in the first embodiment. As in the first embodiment, the optimization determination unit 402 of the OCS controller 4 determines a connection plan for optimizing the data flow (step S103 in FIG. 6 ). In this embodiment, PCIe communication between devices can be performed only via the OCSs 3-1 and 3-2, without the NTB switch 2. That is, if, as a result of continuous optimization, there is a pair of devices that communicate independently without communicating with other devices, it is possible to determine a connection plan that shifts all communication between the pair to communication via only the OCSs 3-1 and 3-2.
[0043] As in the first embodiment, the optimization execution unit 403 of the OCS controller 4 executes optimization in accordance with the connection plan determined by the optimization determination unit 402 and changes the connection of the computer system (step S104 in FIG. 6 ). Specifically, the optimization execution unit 403 instructs the CPU 10 of each device via the NIC 42 to release the ports in use of the NTB adapter 12 specified in the connection plan, release the ports in use of the NTB switch 2 specified in the connection plan, and connect the ports of the NTB adapter 12 specified in the connection plan to OCS 3-1 or 3-2. In response to this instruction, the CPU 10 of the device releases the ports in use of the NTB adapter 12 specified by the optimization execution unit 403 and changes the port assignment so that the ports in the NTB adapter 12 specified by the optimization execution unit 403 are connected to OCS 3-1 or 3-2. Furthermore, the CPU 10 of the device sets the NTB switch 2 so as to release a port designated by the optimization execution unit 403 from among the ports of the NTB switch 2 .
[0044] Furthermore, the optimization execution unit 403 outputs control signals to OCSs 3-1 and 3-2 so that the ports of OCSs 3-1 and 3-2 specified in the connection plan are connected to the devices, the ports of OCS 3-1 specified in the connection plan are connected to OCS 3-2, and the ports of OCS 3-2 specified in the connection plan are connected to OCS 3-1. In response to these control signals, OCS 3-1 changes its own configuration so that the ports of OCS 3-1 specified by the optimization execution unit 403 are connected to the specified devices, and the ports of OCS 3-1 specified by the optimization execution unit 403 are connected to the specified ports of OCS 3-2. In addition, in response to the control signal, OCS3-2 changes its configuration so that the port of OCS3-2 specified by the optimization execution unit 403 is connected to the specified device, and the port of OCS3-2 specified by the optimization execution unit 403 is connected to the specified port of OCS3-1.
[0045] In the case of this embodiment, when communication is between devices connected to the same OCS, it goes without saying that the communication between the devices is performed via this OCS. Other operations are the same as those in the first embodiment.
[0046] In this embodiment, communication between devices is performed only via the OCSs 3-1 and 3-2, without going through the NTB switch 2, thereby making it possible to further reduce power consumption.
[0047] [Third Example] The first and second examples may be applied to resource disaggregation. In the case of disaggregation where each server in a rack has different devices, the CPU box has a small number of NICs and a large number of host devices. The GPU (accelerator) box has a small number of NICs, a small number of host devices, and a large number of GPU devices. The OCS controller box has a large number of NICs and a small number of host devices.
[0048] An NTB switch is used to connect between a CPU box and another CPU box, or between a CPU box and an accelerator box, and the configurations described in the first and second embodiments can be applied to this connection.
[0049] In this embodiment, storing the servers in a rack makes management and operation easier. Also, since the CPU box and accelerator box are closer to each other, it is easier to handle electrical cables between the devices and the NTB switch.
[0050] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes.
[0051] (Supplementary Note 1) A computer system of the present invention includes a plurality of PCIe devices each having an NTB adapter for communicating with another device, an NTB switch capable of selectively connecting the plurality of PCIe devices together, an optical line switch capable of selectively connecting the plurality of PCIe devices together, a data flow monitoring unit configured to monitor a data flow between the PCIe devices via at least one of the NTB switch and the optical line switch, a data flow analysis unit configured to identify a communication between the PCIe devices that is a bottleneck in the data flow based on information on the data flow collected by the data flow monitoring unit, an optimization determination unit configured to determine a connection plan for optimizing the data flow based on an analysis result by the data flow analysis unit, and an optimization execution unit configured to instruct the PCIe devices and the optical line switch to change the connection between the PCIe devices in accordance with the connection plan determined by the optimization determination unit.
[0052] (Supplementary Note 2) In the computer system described in Supplementary Note 1, the data flow analysis unit identifies, based on the data flow information collected by the data flow monitoring unit, communications between PCIe devices that have become a bottleneck in the data flow, as well as communications between PCIe devices that have the potential to become a bottleneck.
[0053] (Appendix 3) In the computer system described in Appendix 1, the NTB switch and the optical line switch are arranged in parallel, the PCIe device is connected to both the NTB switch and the optical line switch, and the optimization determination unit determines the connection plan so that at least one of communication between devices via the NTB switch and communication between devices via the optical line switch is performed.
[0054] (Appendix 4) The computer system described in Appendix 1 is configured such that a plurality of optical line switches are arranged, each PCIe device is connected to one of the plurality of optical line switches, the NTB switch is arranged to connect between the plurality of optical line switches, and the optimization determination unit determines the connection plan so that at least one of communication between devices via one or more of the optical line switches and the NTB switch and communication between devices via one or more of the optical line switches is performed.
[0055] (Supplementary Note 5) In the computer system according to Supplementary Note 1, the optimization determination unit determines the connection plan so as to maximize the amount of data transfer per NTB port of the application executed by the PCIe device.
[0056] The present invention can be applied to a computer system that uses PCIe devices.
[0057] 1-1 to 1-12... PCIe devices, 2... NTB switches 2, 3, 3-1, 3-2... optical line switches, 4... OCS controllers 4, 5, 6-1, 6-2, 7-1, 7-2, 8... cables, 10, 40... CPUs, 11, 41... memories, 12... NTB adapters, 13, 42... NICs, 14... GPUs, 400... data flow monitoring units, 401... data flow analysis units, 402... optimization decision units, 403... optimization execution units.
Claims
1. A computer system comprising: a plurality of PCIe devices each having an NTB adapter for communication with another device; an NTB switch capable of selectively connecting between the plurality of PCIe devices; an optical line switch capable of selectively connecting between the plurality of PCIe devices; a data flow monitoring unit configured to monitor a data flow between the PCIe devices via at least one of the NTB switch and the optical line switch; a data flow analysis unit configured to identify a communication between the PCIe devices that is a bottleneck in the data flow based on information on the data flow collected by the data flow monitoring unit; an optimization determination unit configured to determine a connection plan for optimizing the data flow based on a result of analysis by the data flow analysis unit; and an optimization execution unit configured to instruct the PCIe devices and the optical line switch to change a connection between the PCIe devices in accordance with the connection plan determined by the optimization determination unit.
2. A computer system according to claim 1, characterized in that the data flow analysis unit identifies communications between PCIe devices that have the potential to become a bottleneck in addition to communications between PCIe devices that have become a bottleneck in the data flow based on the data flow information collected by the data flow monitoring unit.
3. A computer system according to claim 1, characterized in that the NTB switch and the optical line switch are arranged in parallel, the PCIe device is connected to both the NTB switch and the optical line switch, and the optimization determination unit determines the connection plan so that at least one of communication between devices via the NTB switch and communication between devices via the optical line switch is performed.
4. A computer system according to claim 1, characterized in that a plurality of the optical line switches are arranged, each PCIe device is connected to one of the plurality of the optical line switches, the NTB switch is arranged to connect between the plurality of the optical line switches, and the optimization determination unit determines the connection plan so that at least one of communication between devices via one or more of the optical line switches and the NTB switch and communication between devices via one or more of the optical line switches is performed.
5. A computer system according to claim 1, wherein the optimization determination unit determines the connection plan so as to maximize the amount of data transfer per NTB port of the application being executed by the PCIe device.