Computer system

JPWO2025100436A1Pending Publication Date: 2025-05-15
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025556415
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2023-11-06
Filing Date
2024-11-06
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

In existing computer systems, PCIe switches are prone to conflicts, resulting in data transmission delays and performance degradation, especially under data-intensive workloads.

Method used

A computer system is designed, including a root complex body, a CPU, a PCIe device, a fiber optic circuit switch and a controller. Through data flow monitoring, conflict detection, optimization of decision-making and optimization of execution units, the PCIe communication path is optimized to avoid conflicts and improve system performance.

Benefits of technology

Effectively manage and reduce conflicts in PCIe switches, improve data transmission efficiency and system performance, and reduce power consumption.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This computer system is provided with route complexes (1-1 - 1-n), CPUs (2-1 - 2-n), a PCIe device, a circuit switch (7), and a controller (8). The controller (8) commands the CPUs (2-1 – 2-n) and the circuit switch (7) to: monitor data flows between the CPUs (2-1 - 2-n) and the PCIe device; identify, on the basis of information about the data flows, communication in which contention occurs in the circuit switch (7); determine a communication path optimization plan on the basis of the identified result; and change communication paths between the CPUs (2-1 - 2-n) and the PCIe device according to the communication path optimization plan.
Need to check novelty before this filing date? Find Prior Art

Description

Computer Systems

[0001] The present invention relates to computer systems.

[0002] 6 shows a typical bus architecture of a computer system. A CPU 101, memory 102, a PCIe (Peripheral Component Interconnect Express) switch 103, and a PCH (Platform Controller Hub) 104 are connected to a root complex 100. A GPU (Graphics Processing Unit) 105, an RDMA NIC (Remote Direct Memory Access Network Interface Card) 106, and the like are connected to the PCIe switch 103. An NVMe SSD (Non-Volatile Memory Express Solid State Drive) 107, a NIC (Network Interface Card) 108, and the like are connected to the PCH 104. Such an architecture is disclosed in Non-Patent Document 1.

[0003] PCIe is a high-speed serial computer expansion bus standard used to connect hardware devices to a computer. CPU 101 connects to the PCIe bus and communicates with other devices via the PCIe bus. Root complex 100 is the primary device connecting CPU 101 and memory 102 to the PCIe bus and is part of the chipset of CPU 101. Root complex 100 initiates transactions on the PCIe bus and routes PCIe traffic to the appropriate device.

[0004] The PCIe switch 103 is used when there are more devices than PCIe lanes available from the root complex 100. The PCIe switch 103 takes one or more incoming connections and splits them into multiple outgoing connections, allowing more devices to be connected to the bus. In the PCIe bus architecture, all of these components are interconnected. The CPU 101 and memory 102 communicate with other devices, such as the GPU 105, NVMe SSD 107, and NICs 106 and 108, via the PCIe bus. The PCIe bus is controlled by the root complex 100 and can also be expanded using the PCIe switch 103. This architecture enables high-speed data transfer between all devices, improving overall system performance.

[0005] Meanwhile, resource disaggregation within a rack is another known example of a computer system. Rack disaggregation refers to the concept of separating and pooling different resources, such as root complexes 200-1 to 200-n (n is an integer equal to or greater than 2), CPUs 201-1 to 201-n, memories 202-1 to 202-n, GPUs 203, NVMe SSDs 204, and NICs 205, as shown in FIG. 7 . These pooled resources are interconnected via a fabric 207, which is a collection of PCIe switches 206. The root complexes 200-1 to 200-n function as an interface between the CPUs 201-1 to 201-n and PCIe devices.

[0006] In such a partitioned setup, contention may occur for the PCIe switch 206. The PCIe switch 206 is responsible for routing data between the CPUs 201-1 to 201-n, the memories 202-1 to 202-n, the GPU 203, the NVMe SSD 204, and the NIC 205. If there is a large amount of data to be routed, the PCIe switch 206 may become a bottleneck in the system, resulting in degradation of data transfer and overall performance.

[0007] For example, both the GPU 203 and the NVMe SSD 204 are high-performance devices capable of processing large amounts of data at high speed. In data-intensive workloads such as machine learning and high-performance computing, the GPU 203 and the NVMe SSD 204 may compete for access to the PCIe switch 206 to transfer data between the CPUs 201-1 to 201-n and the memories 202-1 to 202-n.

[0008] Meanwhile, the NIC 205 manages the network communications of the system. When network traffic is heavy, the NIC 205 may compete with other devices (CPUs 201-1 to 201-n, memories 202-1 to 202-n, GPU 203, and NVMe SSD 204) for access to the PCIe switch 206. This competition may degrade the performance of network communications and the entire system. Recently, methods for high-speed transfer between devices have been established. For example, there are GPU Direct RDMA, which connects a NIC and a GPU with high throughput, and GPU Direct Storage, which connects an RDMA NIC and storage with high throughput.

[0009] Although contention in the PCIe switch 206 in Fig. 7 has been described, contention similarly occurs in the configuration in Fig. 6 . According to Non-Patent Document 1, contention in the PCIe switch 103 easily occurs in the configuration in Fig. 6 . When contention occurs, latency increases by about six times, which has a significant impact on performance and security. Furthermore, in the configuration in Fig. 7 , the load on the PCIe switch 206 increases, making the impact of contention even greater.

[0010] M. Tan, J. Wan, Z. Zhou and Z. Li, “Invisible Probe: Timing Attacks with PCIe Congestion Side-channel”, 2021 IEEE Symposium on Security and Privacy (SP), San Francisco, CA, USA, 2021, pp.322-338, doi:10.1109 / SP40001.2021.00059

[0011] The present invention has been made to solve the above-mentioned problems, and has as its object to provide a computer system that can alleviate contention in a switch.

[0012] The computer system of the present invention comprises: a root complex; a CPU connected to the root complex; a plurality of PCIe devices; a circuit switch capable of selectively connecting the root complex and the plurality of PCIe devices; a data flow monitoring unit configured to monitor a data flow between the CPU and the PCIe devices via the root complex and the circuit switch; a conflict detection unit configured to identify communications in which conflict has occurred in the circuit switch based on data flow information collected by the data flow monitoring unit; an optimization determination unit configured to determine a communication path optimization plan based on a result of identification by the conflict detection unit; and an optimization execution unit configured to instruct the CPU and the circuit switch to change the communication path between the CPU and the PCIe device in accordance with the communication path optimization plan determined by the optimization determination unit.

[0013] According to the present invention, contention in a circuit switch can be alleviated by managing contention in the circuit switch and isolating some of the communication paths for communication in which contention is occurring. As a result, the present invention can process large-volume data communications efficiently and effectively. Furthermore, the present invention can achieve improved security and latency compared to when a PCIe switch is used. When an optical circuit switch is used as a circuit switch, power consumption can be reduced compared to when a PCIe switch is used.

[0014] FIG. 1 is a block diagram showing the configuration of a computer system according to a first embodiment of the present invention. FIG. 2 is a functional block diagram of a controller according to the first embodiment of the present invention. FIG. 3 is a flowchart explaining the operation of the computer system according to the first embodiment of the present invention. FIG. 4 is a block diagram showing the configuration of a computer system according to a second embodiment of the present invention. FIG. 5 is a block diagram showing the configuration of a computer system according to a third embodiment of the present invention. FIG. 6 is a diagram showing a general bus architecture of a conventional computer system. FIG. 7 is a diagram showing the configuration of conventional rack disaggregation.

[0015] [First Embodiment] An embodiment of the present invention will be described below with reference to the drawings. Figure 1 is a block diagram showing the configuration of a computer system according to a first embodiment of the present invention. The computer system includes root complexes 1-1 to 1-n (n is an integer equal to or greater than 2), CPUs 2-1 to 2-n (hosts), memories 3-1 to 3-n, a GPU 4, an NVMe SSD 5, a NIC 6, a circuit switch 7, and a controller 8.

[0016] The CPUs 2-1 to 2-n are connected to the root complexes 1-1 to 1-n via a system bus, and the memories 3-1 to 3-n are connected to the root complexes 1-1 to 1-n via a memory bus.

[0017] The circuit switch 7 selectively connects the root complexes 1-1 to 1-n to PCIe devices (GPU 4, NVMe SSD 5, NIC 6). The circuit switch 7 may be an optical circuit switch (OCS) or an electric circuit switch. In the case of an optical circuit switch, it may be a switch that switches signal lines or an optical wavelength selective switch.

[0018] The ports of each of the route complexes 1-1 to 1-n and the ports of the circuit switch 7 are connected by one or more cables 10, and the ports of each of the PCIe devices and the ports of the circuit switch 7 are connected by one or more cables 11. When an optical circuit switch is used as the circuit switch 7, optical cables are used for the cables 10 and 11, and in this case, photoelectric conversion devices are provided in the connector of the cable 10 connected to the port of the route complexes 1-1 to 1-n and in the connector of the cable 11 connected to the port of the PCIe device.

[0019] An optoelectronic conversion device provided at the connector of the cable 10 converts electrical signals from the route complexes 1-1 to 1-n into optical signals and sends them to the cable 10, converts optical signals received from the circuit switch 7 via the cable 10 into electrical signals and sends them to the route complexes 1-1 to 1-n. An optoelectronic conversion device provided at the connector of the cable 11 converts electrical signals from the PCIe device into optical signals and sends them to the cable 11, and converts optical signals received from the circuit switch 7 via the cable 11 into electrical signals and sends them to the PCIe device.

[0020] The controller 8 controls the circuit switch 7. The controller 8 includes a CPU 80, a memory 81, and a NIC 82. The CPUs 2-1 to 2-n can communicate with the CPU 80 of the controller 8 via the PCIe bus and the NIC 82 of the controller 8. The CPU 80 of the controller 8 communicates with the CPUs 2-1 to 2-n and adjusts the flow of data throughout the system.

[0021] The main function of the controller 8 is to detect device activity that may lead to conflicts or contentions within the system, and to instruct the circuit switch 7 to establish a dedicated communication path between the detected device and any CPUs 2-1 to 2-n that wish to interact with that device. This process separates the communication paths between the devices and the CPUs 2-1 to 2-n, ensuring smooth and uninterrupted data transfer.

[0022] 2 is a functional block diagram of the controller 8. The CPU 80 of the controller 8 executes the following processes in accordance with a program stored in the memory 81, and functions as a data flow monitor 800, a conflict detector 801, an optimization decision unit 802, and an optimization execution unit 803.

[0023] FIG. 3 is a flowchart illustrating the operation of the computer system of this embodiment. The data flow monitoring unit 800 of the controller 8 continuously monitors the data flow between the CPUs 2-1 to 2-n and the PCIe devices (GPU 4, NVMe SSD 5, NIC 6) via the root complexes 1-1 to 1-n and the circuit switch 7 (step S100 in FIG. 3). Items monitored include the amount of data transfer and the data transfer pattern (whether the amount of data transfer fluctuates instantaneously or periodically). The processing of step S100 can be achieved by the data flow monitoring unit 800 querying each of the CPUs 2-1 to 2-n about the data flow (sending a command).

[0024] The conflict detection unit 801 of the controller 8 identifies communications that are experiencing conflicts in the circuit switch 7 and communications that may experience conflicts based on the data flow information collected by the data flow monitoring unit 800 (step S101 in FIG. 3 ). The conflict detection unit 801 performs analyses such as calculating bandwidth usage and predicting future bandwidth usage based on past data transfer patterns. For example, the conflict detection unit 801 detects communications whose bandwidth usage is lower than a predetermined performance threshold as communications that are experiencing conflicts in the circuit switch 7. The conflict detection unit 801 also detects communications whose future bandwidth usage is lower than the performance threshold as communications that may experience conflicts in the circuit switch 7.

[0025] In addition to these dynamic methods, the conflict detection unit 801 also uses a static method in which it decodes a communication library (software) that defines communication between the CPUs 2-1 to 2-n and the PCIe devices, and predicts potential conflicts by predicting data transfer patterns between the CPUs 2-1 to 2-n and the PCIe devices. However, in order to use such a static method, it is necessary that the conflict detection unit 801 be able to obtain a decodeable communication library.

[0026] Next, the optimization determination unit 802 of the controller 8 determines a communication path optimization plan for optimizing PCIe communication based on the result of identification by the conflict detection unit 801 (step S102 in FIG. 3 ). The optimization determination unit 802 determines the communication path optimization plan so that part of the communication in which conflict has occurred in the circuit switch 7 and part of the communication in which conflict may occur are performed via a direct communication path via the circuit switch 7.

[0027] In this case, the optimization determination unit 802 causes the communication with the highest bandwidth usage among the communications in which contention is occurring to be performed via the direct communication path, and the other communications to be performed via the previous communication path. Alternatively, the optimization determination unit 802 causes the communication predicted to have the highest bandwidth usage among the communications in which contention may occur to be performed via the direct communication path, and the other communications to be performed via the previous communication path. Furthermore, the optimization determination unit 802 may cause the communication in which contention is occurring, whose bandwidth usage exceeds a predetermined bandwidth threshold, to be performed via the direct communication path, and the communication in which bandwidth usage is equal to or less than the bandwidth threshold, to be performed via the previous communication path. Furthermore, the optimization determination unit 802 may cause the communication in which contention is occurring, whose predicted bandwidth usage exceeds the bandwidth threshold, to be performed via the direct communication path, and the communication in which predicted bandwidth usage is equal to or less than the bandwidth threshold, to be performed via the previous communication path.

[0028] The communication path optimization plan determined by the optimization determination unit 802 includes information on lanes to be used for direct communication paths among the PCIe lanes of the root complexes 1-1 to 1-n, information on ports to be used for direct communication paths among the ports of the circuit switch 7 on the root complexes 1-1 to 1-n side, information on ports to be used for direct communication paths among the ports of the circuit switch 7 on the PCIe device side, information on lanes to be released among the PCIe lanes in use of the root complexes 1-1 to 1-n, information on ports to be released among the ports in use of the circuit switch 7 on the root complexes 1-1 to 1-n side (ports connected to the PCIe lanes of the root complexes 1-1 to 1-n), and information on ports to be released among the ports in use of the circuit switch 7 on the PCIe device side (ports connected to the PCIe lanes of the PCIe device).

[0029] Next, the optimization execution unit 803 of the controller 8 executes optimization in accordance with the communication path optimization plan determined by the optimization determination unit 802, and changes the connection of the computer system (step S103 in FIG. 3 ). Specifically, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n via the NIC 82 to connect the PCIe lanes of the root complexes 1-1 to 1-n that are determined in the communication path optimization plan to be used for the direct communication path with the ports of the circuit switch 7 on the root complexes 1-1 to 1-n side that are determined in the communication path optimization plan to be used for the direct communication path. Furthermore, when the communication path optimization plan determines that the PCIe lanes and the ports of the circuit switch 7 should be released in conjunction with the transition to the direct communication path, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n to release the lanes that are determined in the communication path optimization plan from the PCIe lanes that are in use in the root complexes 1-1 to 1-n.

[0030] In response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to connect the PCIe lanes of the root complexes 1-1 to 1-n that are determined in the communication path optimization plan to be used for the direct communication path with the ports of the circuit switch 7 on the root complexes 1-1 to 1-n side that are determined in the communication path optimization plan to be used for the direct communication path. Also, in response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to release the lanes that are determined in the communication path optimization plan among the PCIe lanes that are in use by the root complexes 1-1 to 1-n. The root complexes 1-1 to 1-n change the allocation of the PCIe lanes in response to the instruction from the CPUs 2-1 to 2-n.

[0031] Furthermore, the optimization execution unit 803 outputs a control signal to the circuit switch 7 so as to connect a port on the route complex 1-1 to 1-n side of the circuit switch 7 that is determined in the communication path optimization plan to be used for a direct communication path with a port on the PCIe device side of the circuit switch 7 that is determined in the communication path optimization plan to be used for a direct communication path. Furthermore, if the communication path optimization plan determines that a port of the circuit switch 7 should be released in conjunction with the transition to a direct communication path, the optimization execution unit 803 outputs a control signal to the circuit switch 7 so as to release a port that is in use on the circuit switch 7 that is determined in the communication path optimization plan.

[0032] In response to the control signal from the optimization execution unit 803, the circuit switch 7 connects one of its own ports on the side of the root complex 1-1 to 1-n specified by the optimization execution unit 803 to a port on the side of the PCIe device specified by the optimization execution unit 803. The circuit switch 7 also releases the connection of the port specified by the optimization execution unit 803 among its own ports that are in use.

[0033] In this way, some communication routes where conflicts are occurring and some communication routes where conflicts may occur are changed in the circuit switch 7. The processes of steps S100 to S103 are repeatedly executed until the computer system is stopped (YES in step S104 in FIG. 3), thereby constantly optimizing system communications.

[0034] As described above, in this embodiment, contention in the circuit switch 7 is managed and some communication paths where contention is occurring and some communication paths where contention may occur are separated, thereby alleviating contention in the circuit switch 7. This allows the computer system to process large-volume data communications efficiently and effectively. In this embodiment, each PCIe communication is isolated by the circuit switch 7, thereby improving security. Furthermore, in this embodiment, latency can be improved compared to when a PCIe switch is used. When an optical circuit switch is used as the circuit switch 7, power consumption can be reduced compared to when a PCIe switch is used.

[0035] 4 is a block diagram showing the configuration of a computer system according to a second embodiment of the present invention. The computer system according to this embodiment includes root complexes 1-1 to 1-n, CPUs 2-1 to 2-n, memories 3-1 to 3-n, a GPU 4, an NVMe SSD 5, a NIC 6, circuit switches 7a-1 and 7a-2, a controller 8, and a PCIe switch 9.

[0036] As in the first embodiment, the circuit switches 7a-1 and 7a-2 may be optical circuit switches or electrical circuit switches. The ports of each of the route complexes 1-1 to 1-n and the ports of the circuit switch 7a-1 are connected by one or more cables 12, and the ports of each of the PCIe devices and the ports of the circuit switch 7a-2 are connected by one or more cables 13. When optical circuit switches are used as the circuit switches 7a-1 and 7a-2, optical cables are used for the cables 12 and 13. In this case, as in the first embodiment, photoelectric conversion devices may be provided in the connectors of the cables 12 connected to the ports of the route complexes 1-1 to 1-n and the connectors of the cables 13 connected to the ports of the PCIe devices.

[0037] The ports of the circuit switch 7a-1 and the ports of the PCIe switch 9 are connected by one or more cables 14, and the ports of the PCIe switch 9 and the ports of the circuit switch 7a-2 are connected by one or more cables 15. When optical circuit switches are used as the circuit switches 7a-1 and 7a-2, optical cables are used for the cables 14 and 15. In this case, similar to the first embodiment, photoelectric conversion devices may be provided in the connectors of the cables 14 and 15 connected to the ports of the PCIe switch 9. The ports of the circuit switch 7a-1 and the ports of the circuit switch 7a-2 are connected by one or more optical cables 16.

[0038] In this way, connections in this system are realized by wiring using the PCIe protocol. PCIe communication is performed between CPUs 2-1 to 2-n and PCIe devices (GPU 4, NVMe SSD 5, NIC 6) via root complexes 1-1 to 1-n, circuit switch 7a-1, PCIe switch 9, and circuit switch 7a-2. PCIe communication may also be performed only via circuit switches 7a-1 and 7a-2, without using the PCIe switch 9.

[0039] The processing flow of the computer system is the same as in the first embodiment, and will be described using the reference numerals in Fig. 3. As in the first embodiment, the data flow monitoring unit 800 of the controller 8 continuously monitors the data flow between the CPUs 2-1 to 2-n and the PCIe devices via the root complexes 1-1 to 1-n, the circuit switch 7a-1, the PCIe switch 9, and the circuit switch 7a-2, and the data flow between the CPUs 2-1 to 2-n and the PCIe devices via the root complexes 1-1 to 1-n and the circuit switches 7a-1 and 7a-2 (step S100 in Fig. 3).

[0040] As in the first embodiment, the conflict detection unit 801 of the controller 8 identifies communications in which conflicts are occurring and communications in which conflicts may occur in the circuit switch 7a-1, the PCIe switch 9, or the circuit switch 7a-2, based on the data flow information collected by the data flow monitoring unit 800 (step S101 in FIG. 3).

[0041] As in the first embodiment, the optimization determination unit 802 of the controller 8 determines a communication path optimization plan for optimizing PCIe communication based on the identification result by the conflict detection unit 801 (step S102 in FIG. 3 ). The optimization determination unit 802 determines a communication path optimization plan that employs at least one of two methods: a method of limiting the bandwidth usage of the PCIe switch 9 used by communications in which conflict has occurred and communications in which conflict may occur, and a method of shifting part of communications in which conflict has occurred and communications in which conflict may occur to a direct communication path that does not go through the PCIe switch 9.

[0042] At this time, the optimization determination unit 802 limits the bandwidth usage of the PCIe switch 9 used by communications in which contention is occurring and whose bandwidth usage is equal to or less than a predetermined bandwidth threshold. Furthermore, the optimization determination unit 802 limits the bandwidth usage of the PCIe switch 9 used by communications in which contention is likely to occur and whose predicted bandwidth usage is equal to or less than a bandwidth threshold. The bandwidth usage can be limited by releasing the PCIe lanes of the root complexes 1-1 to 1-n used for communication via the PCIe switch 9 and reducing the number of ports of the circuit switches 7a-1 and 7a-2 used for communication via the PCIe switch 9, thereby reducing the number of PCIe lanes. The extent to which the bandwidth usage should be reduced can be determined by, for example, setting a reduction rule in advance in the memory 81 of the controller 8, and the optimization determination unit 802 determining the number of PCIe lanes to reduce in accordance with this reduction rule.

[0043] Furthermore, the optimization determination unit 802 may cause, among communications in which contention is occurring, communications in which bandwidth usage exceeds a predetermined bandwidth threshold to be performed via a direct communication path that does not go through the PCIe switch 9, and communications in which bandwidth usage is equal to or less than the bandwidth threshold to be performed via a previous communication path. Furthermore, the optimization determination unit 802 may cause, among communications in which contention is likely to occur, communications in which predicted bandwidth usage exceeds the bandwidth threshold to be performed via a direct communication path that does not go through the PCIe switch 9, and communications in which predicted bandwidth usage is equal to or less than the bandwidth threshold to be performed via a previous communication path.

[0044] By determining the communication path optimization plan as described above, data transfer can be performed smoothly while preventing contention for communication bandwidth via the PCIe switch 9. The data flow is shared in the PCIe switch 9, and the data flow can be monopolized in a direct communication path that does not go through the PCIe switch 9 but goes through only the circuit switches 7a-1 and 7a-2. For example, communication between a low-bandwidth device such as the NVMe SSD 5 and the CPUs 2-1 to 2-n is performed via the PCIe switch 9, while communication between a high-bandwidth device such as the GPU 4 and the CPUs 2-1 to 2-n is performed via only the circuit switches 7a-1 and 7a-2.

[0045] The communication path optimization plan determined by the optimization determination unit 802 includes information on PCIe lanes of the root complexes 1-1 to 1-n to be used for direct communication paths, information on ports of the circuit switch 7a-1 on the root complexes 1-1 to 1-n side to be used for direct communication paths, information on ports of the circuit switch 7a-1 on the PCIe switch 9 side to be used for direct communication paths, information on ports of the circuit switch 7a-2 on the PCIe switch 9 side to be used for direct communication paths, information on ports of the circuit switch 7a-2 on the PCIe device side to be used for direct communication paths, information on PCIe lanes of the root complexes 1-1 to 1-n to be used for free communication paths, and information on the lanes to be released, information on the ports to be released among the ports in use on the root complexes 1-1 to 1-n side of the circuit switch 7a-1 (ports connected to the PCIe lanes of the root complexes 1-1 to 1-n), information on the ports to be released among the ports in use on the PCIe switch 9 side of the circuit switch 7a-1 (ports connected to the PCIe lanes of the PCIe switch 9), information on the ports to be released among the ports in use on the PCIe switch 9 side of the circuit switch 7a-2 (ports connected to the PCIe lanes of the PCIe switch 9), and information on the ports to be released among the ports in use on the PCIe device side of the circuit switch 7a-2 (ports connected to the PCIe lanes of the PCIe device).

[0046] As in the first embodiment, the optimization execution unit 803 of the controller 8 executes optimization in accordance with the communication path optimization plan determined by the optimization determination unit 802, and changes the connection of the computer system (step S103 in FIG. 3). Specifically, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n via the NIC 82 to connect the PCIe lanes of the root complexes 1-1 to 1-n that are determined in the communication path optimization plan to be used for the direct communication path with the ports of the circuit switch 7a-1 on the root complexes 1-1 to 1-n sides that are determined in the communication path optimization plan to be used for the direct communication path. Furthermore, when the communication path optimization plan specifies the release of PCIe lanes and the release of ports of the circuit switches 7a-1 and 7a-2 in order to transition to a direct communication path or limit bandwidth usage, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n to release the lanes specified in the communication path optimization plan from among the PCIe lanes in use of the root complexes 1-1 to 1-n.

[0047] In response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to connect the PCIe lanes of the root complexes 1-1 to 1-n that are determined in the communication path optimization plan to be used for the direct communication path with the ports of the circuit switch 7a-1 on the root complex 1-1 to 1-n side that are determined in the communication path optimization plan to be used for the direct communication path. Furthermore, in response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to release the lanes that are determined in the communication path optimization plan among the PCIe lanes that are in use by the root complexes 1-1 to 1-n. In response to an instruction from the CPUs 2-1 to 2-n, the root complexes 1-1 to 1-n change the allocation of the PCIe lanes.

[0048] The optimization execution unit 803 also outputs a control signal to the circuit switch 7a-1 so that a port on the PCIe switch 9 side of the circuit switch 7a-1 that is determined in the communication path optimization plan to be used for a direct communication path among the ports on the root complexes 1-1 to 1-n side of the circuit switch 7a-1 is connected to a port on the PCIe switch 9 side of the circuit switch 7a-1 that is determined in the communication path optimization plan to be used for a direct communication path. The optimization execution unit 803 also outputs a control signal to the circuit switch 7a-2 so that a port on the PCIe switch 9 side of the circuit switch 7a-2 that is determined in the communication path optimization plan to be used for a direct communication path is connected to a port on the PCIe device side of the circuit switch 7a-2 that is determined in the communication path optimization plan to be used for a direct communication path.

[0049] Furthermore, when the communication path optimization plan specifies that ports of the circuit switches 7a-1 and 7a-2 should be released in order to transition to a direct communication path or to limit bandwidth usage, the optimization execution unit 803 outputs a control signal to the circuit switches 7a-1 and 7a-2 to release the ports specified in the communication path optimization plan from among the ports in use of the circuit switches 7a-1 and 7a-2.

[0050] In response to the control signal from the optimization execution unit 803, the circuit switch 7a-1 connects one of its own ports on the side of the root complexes 1-1 to 1-n specified by the optimization execution unit 803 to a port on the side of the PCIe switch 9 specified by the optimization execution unit 803. In response to the control signal from the optimization execution unit 803, the circuit switch 7a-2 connects one of its own ports on the side of the PCIe switch 9 specified by the optimization execution unit 803 to a port on the side of the PCIe device specified by the optimization execution unit 803. Furthermore, the circuit switches 7a-1 and 7a-2 release the connection of the port specified by the optimization execution unit 803 from their own ports that are in use.

[0051] The processes of steps S100 to S103 are repeated until the computer system is shut down (YES in step S104 in FIG. 3), thereby constantly optimizing system communications.

[0052] In this embodiment, it is possible to obtain the same effects as in the first embodiment. Furthermore, in this embodiment, it becomes possible for multiple CPUs to share communications with multiple devices. In this embodiment, it is possible to provide users with a direct communication path specialized in performance and security, and a shared path specialized in flexibility and device utilization efficiency.

[0053] 5 is a block diagram showing the configuration of a computer system according to a third embodiment of the present invention. The computer system of this embodiment is the same as the configuration of the second embodiment except that the circuit switch 7a-1 is removed.

[0054] The ports of each of the root complexes 1-1 to 1-n and the ports of the circuit switch 7a-2 are connected by one or more cables 17, and the ports of each of the root complexes 1-1 to 1-n and the ports of the PCIe switch 9 are connected by one or more cables 18.

[0055] The CPUs 2-1 to 2-n and the PCIe devices (GPU 4, NVMe SSD 5, NIC 6) perform PCIe communication via the root complexes 1-1 to 1-n, the PCIe switch 9, and the circuit switch 7a-2. There are also cases where PCIe communication is performed only via the circuit switch 7a-2 without via the PCIe switch 9.

[0056] The processing flow of the computer system is the same as in the first and second embodiments, and will be described using the reference numerals in Fig. 3. As in the first and second embodiments, the data flow monitoring unit 800 of the controller 8 continuously monitors the data flow between the CPUs 2-1 to 2-n and the PCIe devices via the root complexes 1-1 to 1-n, the PCIe switch 9, and the circuit switch 7a-2, and the data flow between the CPUs 2-1 to 2-n and the PCIe devices via the root complexes 1-1 to 1-n and the circuit switch 7a-2 (step S100 in Fig. 3).

[0057] Similar to the first and second embodiments, the conflict detection unit 801 of the controller 8 identifies communications in which conflicts have occurred or communications in which conflicts may occur in the PCIe switch 9 or the circuit switch 7a-2, based on the data flow information collected by the data flow monitoring unit 800 (step S101 in FIG. 3).

[0058] As in the first and second embodiments, the optimization determination unit 802 of the controller 8 determines a communication path optimization plan for optimizing PCIe communication based on the identification result by the conflict detection unit 801 (step S102 in FIG. 3).

[0059] At this time, the optimization determination unit 802 limits the bandwidth usage of the PCIe switch 9 used by communications in which contention is occurring and whose bandwidth usage is equal to or less than a predetermined bandwidth threshold. Furthermore, the optimization determination unit 802 limits the bandwidth usage of the PCIe switch 9 used by communications in which contention is likely to occur and whose predicted bandwidth usage is equal to or less than a bandwidth threshold. The bandwidth usage can be limited by releasing the PCIe lanes of the root complexes 1-1 to 1-n used for communication via the PCIe switch 9 and reducing the number of ports of the circuit switch 7a-2 used for communication via the PCIe switch 9, thereby reducing the number of PCIe lanes. The extent to which the bandwidth usage should be reduced can be determined by, for example, setting a reduction rule in advance in the memory 81 of the controller 8, and the optimization determination unit 802 determining the number of PCIe lanes to reduce in accordance with this reduction rule.

[0060] Furthermore, the optimization determination unit 802 may cause, among communications in which contention is occurring, communications in which bandwidth usage exceeds a predetermined bandwidth threshold to be performed via a direct communication path that does not go through the PCIe switch 9, and communications in which bandwidth usage is equal to or less than the bandwidth threshold to be performed via a previous communication path. Furthermore, the optimization determination unit 802 may cause, among communications in which contention is likely to occur, communications in which predicted bandwidth usage exceeds the bandwidth threshold to be performed via a direct communication path that does not go through the PCIe switch 9, and communications in which predicted bandwidth usage is equal to or less than the bandwidth threshold to be performed via a previous communication path.

[0061] By determining the communication path optimization plan as described above, data transfer can be performed smoothly while preventing contention for communication bandwidth via the PCIe switch 9. The data flow is shared in the PCIe switch 9, and the data flow can be monopolized in a direct communication path that goes only through the circuit switch 7a-2 without going through the PCIe switch 9. For example, communication between a low-bandwidth device such as the NVMe SSD 5 and the CPUs 2-1 to 2-n is performed via the PCIe switch 9, while communication between a high-bandwidth device such as the GPU 4 and the CPUs 2-1 to 2-n is performed only via the circuit switch 7a-2.

[0062] The communication path optimization plan determined by the optimization determination unit 802 includes information on PCIe lanes of the root complexes 1-1 to 1-n to be used for direct communication paths, information on ports on the PCIe switch 9 side of the circuit switch 7a-2 to be used for direct communication paths, information on ports on the PCIe device side of the circuit switch 7a-2 to be used for direct communication paths, information on lanes to be released among PCIe lanes in use of the root complexes 1-1 to 1-n, information on ports to be released among ports in use on the PCIe switch 9 side of the circuit switch 7a-2 (ports connected to the PCIe lanes of the PCIe switch 9), and information on ports to be released among ports in use on the PCIe device side of the circuit switch 7a-2 (ports connected to the PCIe lanes of the PCIe device).

[0063] As in the first and second embodiments, the optimization execution unit 803 of the controller 8 executes optimization in accordance with the communication path optimization plan determined by the optimization determination unit 802, and changes the connection of the computer system (step S103 in FIG. 3). Specifically, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n via the NIC 82 to connect the PCIe lanes of the root complexes 1-1 to 1-n, which are determined in the communication path optimization plan to be used for the direct communication path, to the ports of the circuit switch 7a-2 on the root complexes 1-1 to 1-n sides, which are determined in the communication path optimization plan to be used for the direct communication path. Furthermore, when the communication path optimization plan specifies the release of PCIe lanes and the release of ports of the circuit switch 7a-2 in order to transition to a direct communication path or limit bandwidth usage, the optimization execution unit 803 instructs the CPUs 2-1 to 2-n to release the lanes specified in the communication path optimization plan from among the PCIe lanes in use of the root complexes 1-1 to 1-n.

[0064] In response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to connect the PCIe lanes of the root complexes 1-1 to 1-n that are determined in the communication path optimization plan to be used for the direct communication path with the ports of the circuit switch 7a-2 on the root complex 1-1 to 1-n side that are determined in the communication path optimization plan to be used for the direct communication path. Furthermore, in response to an instruction from the optimization execution unit 803, the CPUs 2-1 to 2-n instruct the root complexes 1-1 to 1-n to release the lanes that are determined in the communication path optimization plan among the PCIe lanes that are in use by the root complexes 1-1 to 1-n. In response to an instruction from the CPUs 2-1 to 2-n, the root complexes 1-1 to 1-n change the allocation of the PCIe lanes.

[0065] In addition, the optimization execution unit 803 outputs a control signal to the circuit switch 7a-2 so that a port on the root complex 1-1 to 1-n side of the circuit switch 7a-2 that is determined in the communication path optimization plan to be used for a direct communication path is connected to a port on the PCIe device side of the circuit switch 7a-2 that is determined in the communication path optimization plan to be used for a direct communication path.

[0066] Furthermore, if the communication path optimization plan specifies that a port on the circuit switch 7a-2 be released in order to transition to a direct communication path or to limit bandwidth usage, the optimization execution unit 803 outputs a control signal to the circuit switch 7a-2 to release the port specified in the communication path optimization plan from among the ports in use on the circuit switch 7a-2.

[0067] In response to a control signal from the optimization execution unit 803, the circuit switch 7a-2 connects one of its own ports on the side of the root complex 1-1 to 1-n specified by the optimization execution unit 803 to a port on the side of the PCIe device specified by the optimization execution unit 803. In addition, the circuit switch 7a-2 releases the connection of the port specified by the optimization execution unit 803 among its own ports that are in use.

[0068] The processes of steps S100 to S103 are repeated until the computer system is shut down (YES in step S104 in FIG. 3). This ensures that system communications are always optimized. This embodiment achieves the same effects as the second embodiment, but with lower power consumption and lower cost.

[0069] [Fourth Example] There are various approaches to resource disaggregation, such as a method of preparing different resources for each rack and connecting the racks (rack disaggregation), and a method of preparing multiple racks with different resources for each server (server disaggregation).

[0070] According to the present invention, the first to third embodiments can be applied to the above-described resource disaggregation. Specifically, in server disaggregation, the amount of PCIe communication between hosts and devices within a rack is greater than the amount of communication between racks, so using circuit switches 7, 7a-1, and 7a-2 as switches within the rack offers significant benefits in terms of performance and security. Needless to say, in rack disaggregation, the switches within the rack may also be circuit switches 7, 7a-1, and 7a-2. Because the distance between the CPU and the device is short (within 1 to 2 m), the circuit switches 7, 7a-1, and 7a-2 may be electrical circuit switches or optical circuit switches.

[0071] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes.

[0072] (Supplementary Note 1) A computer system of the present invention includes a root complex, a CPU connected to the root complex, a plurality of PCIe devices, a circuit switch capable of selectively connecting the root complex and the plurality of PCIe devices, a data flow monitoring unit configured to monitor a data flow between the CPU and the PCIe devices via the root complex and the circuit switch, a conflict detection unit configured to identify communications in which conflict has occurred in the circuit switch based on data flow information collected by the data flow monitoring unit, an optimization determination unit configured to determine a communication path optimization plan based on a result of identification by the conflict detection unit, and an optimization execution unit configured to instruct the CPU and the circuit switch to change a communication path between the CPU and the PCIe device in accordance with the communication path optimization plan determined by the optimization determination unit.

[0073] (Appendix 2) In the computer system described in Appendix 1, the conflict detection unit identifies communications that are experiencing conflicts in the circuit switch, as well as communications that may be experiencing conflicts, based on the data flow information collected by the data flow monitoring unit.

[0074] (Supplementary Note 3) In the computer system described in Supplementary Note 2, the conflict detection unit decodes a communication library that defines communication between the CPU and the PCIe device, and identifies communication that may cause a conflict in the circuit switch.

[0075] (Appendix 4) In the computer system described in Appendix 1, the optimization determination unit determines the communication path optimization plan so that, among the communications identified by the conflict detection unit, communications whose bandwidth usage exceeds a predetermined bandwidth threshold are carried out via a direct communication path that does not share a communication path with other communications.

[0076] (Supplementary Note 5) The computer system described in Supplementary Note 1 further includes a PCIe switch between the root complex and the circuit switch, the root complex is connected to the PCIe switch and the circuit switch, and the optimization determination unit determines the communication path optimization plan so that at least one of communication via the PCIe switch and the circuit switch and communication via only the circuit switch is performed.

[0077] (Supplementary Note 6) The computer system described in Supplementary Note 1 further includes a PCIe switch, wherein a plurality of the circuit switches are arranged and connected to each other, the root complex is connected to one of the plurality of circuit switches, each PCIe device is connected to a circuit switch different from the circuit switch to which the root complex is connected, the PCIe switch is arranged to connect the plurality of circuit switches, and the optimization determination unit determines the communication path optimization plan so that at least one of communication via the plurality of circuit switches and the PCIe switch and communication via only the plurality of circuit switches is performed.

[0078] (Appendix 7) In the computer system described in Appendix 5 or 6, the optimization determination unit determines the communication path optimization plan so that, among the communications identified by the conflict detection unit, communications whose bandwidth usage exceeds a predetermined bandwidth threshold are performed via a direct communication path that does not go through the PCIe switch.

[0079] (Supplementary Note 8) In the computer system described in Supplementary Note 5 or 6, the optimization determination unit determines the communication path optimization plan so as to limit the bandwidth usage of the PCIe switch used by communications identified by the conflict detection unit that have bandwidth usage equal to or less than a predetermined bandwidth threshold.

[0080] The present invention can be applied to a computer system that uses PCIe devices.

[0081] 1-1 to 1-n...Root complex, 2-1 to 2-n, 80...CPU, 3-1 to 3-n, 81...Memory, 4...GPU, 5...NVMe SSD, 6, 82...NIC, 7, 7a-1, 7a-2...Circuit switch, 8...Controller, 9...PCIe switch, 10 to 18...Cable, 800...Data flow monitoring unit, 801...Conflict detection unit, 802...Optimization determination unit, 803...Optimization execution unit.

Claims

1. A computer system comprising: a root complex; a CPU connected to the root complex; a plurality of PCIe devices; a circuit switch capable of selectively connecting the root complex and the plurality of PCIe devices; a data flow monitoring unit configured to monitor a data flow between the CPU and the PCIe devices via the root complex and the circuit switch; a conflict detection unit configured to identify communications in which conflicts have occurred in the circuit switch based on data flow information collected by the data flow monitoring unit; an optimization determination unit configured to determine a communication path optimization plan based on a result of the identification by the conflict detection unit; and an optimization execution unit configured to instruct the CPU and the circuit switch to change a communication path between the CPU and the PCIe device in accordance with the communication path optimization plan determined by the optimization determination unit.

2. A computer system as claimed in claim 1, characterized in that the conflict detection unit identifies communications which may potentially cause conflicts in addition to communications which are causing conflicts in the circuit switch based on the data flow information collected by the data flow monitoring unit.

3. A computer system according to claim 2, wherein the conflict detection unit decodes a communication library that defines communication between the CPU and the PCIe device, and identifies communication that may cause a conflict in the circuit switch.

4. A computer system as described in claim 1, characterized in that the optimization determination unit determines the communication path optimization plan so that communications identified by the conflict detection unit, whose bandwidth usage exceeds a predetermined bandwidth threshold, are carried out via a direct communication path that does not share a communication path with other communications.

5. A computer system according to claim 1, further comprising a PCIe switch between the root complex and the circuit switch, the root complex being connected to the PCIe switch and the circuit switch, and the optimization determination unit determining the communication path optimization plan so that at least one of communication via the PCIe switch and the circuit switch and communication via only the circuit switch is performed.

6. A computer system according to claim 1, further comprising a PCIe switch, wherein a plurality of the circuit switches are arranged and connected to each other, the root complex is connected to one of the plurality of circuit switches, each PCIe device is connected to a circuit switch different from the circuit switch to which the root complex is connected, the PCIe switch is arranged to connect between the plurality of circuit switches, and the optimization determination unit determines the communication path optimization plan so that at least one of communication via the plurality of circuit switches and the PCIe switch and communication only via the plurality of circuit switches is performed.

7. A computer system according to claim 5 or 6, characterized in that the optimization determination unit determines the communication path optimization plan so that communications identified by the conflict detection unit, whose bandwidth usage exceeds a predetermined bandwidth threshold, are carried out via a direct communication path that does not go through the PCIe switch.

8. A computer system according to claim 5 or 6, characterized in that the optimization determination unit determines the communication path optimization plan so as to limit the bandwidth usage of the PCIe switch used by communications identified by the conflict detection unit that have bandwidth usage below a predetermined bandwidth threshold.