Data transfer controller and information processing apparatus

The data transfer control device addresses bandwidth limitations by routing excess data to a memory for retrieval, ensuring accurate processing in systems with multiple accelerators.

JP2025119281APending Publication Date: 2025-08-14FUJITSU LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024014079
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

In systems with multiple accelerators, transferring large amounts of data per unit time can exceed the bandwidth between the receiving device and the switch, leading to data loss and incorrect processing by accelerators.

Method used

A data transfer control device manages data transfers by routing excess data to a first memory and allowing second devices to retrieve it, ensuring data is processed without exceeding the bandwidth limit.

Benefits of technology

Normal data transfer is maintained to multiple second devices without exceeding the bandwidth, enabling accurate processing and system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025119281000001_ABST
    Figure 2025119281000001_ABST
Patent Text Reader

Abstract

To normally transfer data from a first device to each of a plurality of second devices without exceeding a bandwidth between the first device and a switch even when an hourly transfer volume of the data transferred from the first device to the second devices is more than a predetermined volume.SOLUTION: A data transfer controller that controls transfer of a plurality of data items from a first device to a plurality of second devices is installed in an information processing apparatus. The information processing apparatus includes: the first device that receives the plurality of data items; the second device that receives the plurality of data items transmitted in parallel from the first device, and processes them; one or more first memories that can hold the data; and a switch. When an hourly transfer volume of the plurality of data items transferred from the first device to the switch exceeds a bandwidth between the first device and the switch, the data transfer controller causes the data which is to be transferred to the plurality of second devices to be transferred to the first memory, and causes the plurality of second devices to obtain the data held in the first memory.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a data transfer control device and an information processing device. [Background technology]

[0002] In a computer system in which a host bus connected to a CPU (Central Processing Unit) and memory and an I / O bus connected to I / O (Input / Output) devices are interconnected via a bus bridge, the performance of the bus may be degraded due to memory access conflicts between the CPU and the I / O devices. To address this issue, a method has been proposed in which segmented memories are connected to the host bus and the I / O bus, respectively, and data transfer between the I / O devices is performed via the memory connected to the I / O bus (see, for example, Patent Document 1).

[0003] In an image transmission device capable of realizing a live distribution function, when data encoded by an encoding unit is written to a memory for live distribution and a storage medium for storage, if data writing conflicts on the bus and the bus performance deteriorates, real-time performance may not be ensured. Therefore, a method has been proposed in which a dedicated bus connected to the storage medium is provided and data is written to the storage medium without going through the bus connected to the memory (see, for example, Patent Document 2).

[0004] In an information processing device for displaying images, which has a video memory connected to a common bus and a main memory connected to a local bus, if the common bus is occupied by display data, processing using the common bus may become impossible. Therefore, a method has been proposed in which the transfer of display data from the main memory to the video memory is made less frequent than the writing of display data to the main memory (see, for example, Patent Document 3). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 09-006711 [Patent Document 2] International Publication No. 2004-093445 [Patent Document 3] Japanese Patent Application Laid-Open No. 2015-176569 Summary of the Invention [Problem to be solved by the invention]

[0006] Recently, systems have been developed that combine multiple accelerators with different processing capabilities, and transfer one or more pieces of data to each accelerator for processing, thereby enabling multiple types of data processing to be performed at high speed.

[0007] In this type of system, when multiple data received by a receiving device are transferred in parallel to multiple accelerators via a switch, the amount of data transferred per unit time from the receiving device to the switch may exceed the bandwidth between the receiving device and the switch. If the amount of data transferred per unit time exceeds the bandwidth between the receiving device and the switch, data with missing information may be transferred to each accelerator, making it difficult for each accelerator to process the data correctly.

[0008] In one aspect, the present invention aims to enable normal data to be transferred to each of multiple second devices without exceeding the bandwidth between the first device and the switch, even when the amount of data transferred per unit time from the first device to the second device is large. [Means for solving the problem]

[0009] According to one aspect, a data transfer control device is mounted on an information processing device having a first device that receives multiple data sets, multiple second devices that receive and process the multiple data sets transmitted in parallel from the first device, one or more first memories capable of holding the data, and a switch that interconnects the first device, the second device, and the first memory, and controls the transfer of the multiple data sets from the first device to the multiple second devices.When the amount of data transferred per hour from the first device to the switch exceeds the bandwidth between the first device and the switch, the data transferred from the first device to the multiple second devices is transferred to the first memory, and the data held in the first memory is retrieved by the multiple second devices. [Effects of the Invention]

[0010] Even when the amount of data transferred per unit time from the first device to the second device is large, it is possible to transfer normal data to each of the multiple second devices without exceeding the bandwidth between the first device and the switch. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram illustrating an example of an information processing device according to an embodiment. [Figure 2] 1. FIG. 4 is an explanatory diagram showing an example of an operation of transferring data without exceeding the bandwidth of a bus in the information processing device of FIG. [Figure 3] FIG. 10 is a block diagram illustrating an example of an information processing device according to another embodiment. [Figure 4] 4 is an explanatory diagram showing an example in which stream data is not transferred normally in the information processing device of FIG. 3; FIG. [Figure 5] 5A and 5B are diagrams illustrating examples of various tables provided in the transfer control unit of FIG. 4. [Figure 6] 4 is a diagram illustrating an example of an address space of a PCIe bus in the information processing device of FIG. 3. FIG. [Figure 7]5A and 5B are diagrams illustrating examples of various tables provided in a transfer control unit to solve the problem shown in FIG. 4. [Figure 8] 5 is an explanatory diagram showing an example of an operation for solving the problem shown in FIG. 4 in the information processing device of FIG. 3. [Figure 9] 9 is a flowchart showing an example of an operation for determining a transfer path for the stream data shown in FIG. 8. FIG. [Figure 10] 9 is an explanatory diagram showing an example of calculating latency until stream data 2 arrives at GPU1-GPU3 from FPGA in the data transfer path of FIG. 8. FIG. [Figure 11] 10A and 10B are diagrams illustrating examples of various tables provided in a transfer control unit to solve the problem of unsatisfied latency. [Figure 12] 4 is an explanatory diagram showing an example of an operation for solving the problem of unsatisfied latency in the information processing device of FIG. 3. FIG. [Figure 13] 13 is a diagram showing an example of settings of various tables of a transfer control unit on a transfer path of stream data shown in FIG. 12. FIG. [Figure 14] 13 is an explanatory diagram showing an example of calculating latency until stream data 2 arrives at GPU1-GPU3 from FPGA in the data transfer path of FIG. 12. FIG. [Figure 15] 13 is a flowchart showing an example of an operation for determining a transfer path for the stream data shown in FIG. 12.

[0033] FIG. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments will be described with reference to the drawings.

[0013] Fig. 1 shows an example of an information processing device according to an embodiment. The information processing device 10 shown in Fig. 1 includes a first device 20 including an internal memory 21, a data transfer control device 30, multiple second devices 40 each including an internal memory 41, a switch 50, and a memory 60. The first device 20, each of the second devices 40, and the memory 60 are connected to the switch 50 via a bus BUS. The switch 50 can transfer data between the first device 20, each of the second devices 40, and the memory 60 connected to each bus BUS. The bandwidth of each bus BUS is assumed to be 3b [MB / s].

[0014] The information processing device 100 repeatedly receives multiple data items in parallel from the outside, sequentially acquires a predetermined number of the received data items into multiple second devices, and processes the data items in each of the multiple second devices. To ensure that the data processing by the second devices is performed without failure, data is transferred from the first device 20 to each second device 40 at a rate equal to or faster than the reception rate of each data item received from the outside. For example, the transfer rate b [MB / s] of each data item transferred from the first device 20 to each second device 40 is assumed to be the same as the reception rate b [MB / s] of each data item received by the first device 20. Although not particularly limited, the data received from the outside may be stream data such as video data.

[0015] The data transfer control device 30 controls the transfer paths of multiple data transferred from the first device 20 to each second device 40. The data transfer paths may include a memory 60. In this case, any of the data is transferred from the first device 20 to the second device 40 via a switch 50 and the memory 60. The data transfer control device 30 may have a table that stores information such as the transfer speed of each data, information indicating the second device 40 that processes each data, address information of the storage area in the internal memory 21 that stores each data, and information indicating the bandwidth of each bus BUS. The data transfer control device 30 may be provided within the first device 20. In this case, the information stored in the data transfer control device 30 can be referenced by the second device 40.

[0016] The first device 20 stores each of the three data received in parallel from outside the information processing device 10 in the internal memory 21. Each second device 40 transfers the data to be processed, of the three data stored in the internal memory 21, from the internal memory 21 to its own internal memory 41 via the switch 50, according to the information stored in the table in the data transfer control device 30. In the example shown in FIG. 1, the second device 40(1) reads and processes three data from the internal memory 21, and the second device 40(2) reads and processes one data from the internal memory 21. Each second device 40 sequentially outputs the processed data.

[0017] 1 shows an example in which three pieces of data are transferred from the first device 20 to each of the second devices 40 without using memory 60. The data shown by the solid lines and processed by both the second devices 40(1) and 40(2) is shown with two transfer paths. The two pieces of data shown by the two types of dashed lines and processed only by the second device 40(1) are each shown with one transfer path. Therefore, the amount of data transferred per unit time from the first device 20 to the switch 50 is 4b [MB / s], and the data cannot be transferred over the bus BUS with a bandwidth of 3b [MB / s].

[0018] Fig. 2 shows an example of the operation of transferring data without exceeding the bandwidth of the bus BUS in the information processing device 10 of Fig. 1. In the example shown in Fig. 2, data shown by the solid line is transferred to the second devices 40(1) and 40(2) via memory 60. In this case, the transfer rate per unit time of data sent from the first device 20 to the switch 50 is 3b [MB / s], which can be kept below the 3b [MB / s] bandwidth of the bus BUS.

[0019] Data indicated by a single solid line is written to memory 60. Data indicated by two solid lines is read from memory 60 and transferred to second devices 40(1) and 40(2) via switch 50. Therefore, the transfer rate of data read from and written to memory 60 per unit time is 3b [MB / s], which can be kept below the 3b [MB / s] bandwidth of the bus BUS.

[0020] The amount of data transferred per unit time to second devices 40(1) and 40(2) is the same as in FIG. 1 and can be kept below the bandwidth of the bus BUS of 3b [MB / s]. As a result, data can be transferred from first device 20 to second devices 40(1) and 40(2) without exceeding the bandwidth of each bus BUS. Therefore, second devices 40(1) and 40(2) can receive and process data normally without any loss, allowing information processing device 10 to operate normally.

[0021] As described above, in this embodiment, data can be transferred from the first device 20 to the plurality of second devices 40 without exceeding the bandwidth of the bus BUS connecting the first device 20 and the switch 50. In other words, even when the number of parallel data transfers from the first device 20 to the switch 50 is reduced to prevent exceeding the bandwidth, a predetermined number of data can be transferred to each of the plurality of second devices 40 without dropping any data. Therefore, the plurality of second devices 40 can receive and process data normally without dropping any data, allowing the information processing device 10 to operate normally.

[0022] Fig. 3 shows an example of an information processing device according to another embodiment. The information processing device 100 shown in Fig. 3 includes a field-programmable gate array (FPGA) 200, multiple graphics processing units (GPUs) 400, and a peripheral component interconnect (PCIe) switch 500. The information processing device 100 also includes multiple compute express link (CXL) memories 600, a root complex 700, a central processing unit (CPU) 800, and a host memory 900.

[0023] The FPGA 200 is an example of a first device, and the GPU 400 is an example of a second device. The CXL memory 600 is an example of a first memory. The information processing device 100 transfers data such as stream data using a PCIe interface, but may transfer data using another interface. Instead of the CXL memory 600, a memory of another standard may be installed in the information processing device 100.

[0024] The FPGA 200, GPUs 400(1), 400(2), 400(3), and 400(4), CXL memories 600(1) and 600(2), and the root complex 700 are connected to the PCIe switch 500 via PCIe buses and can communicate with each other. Hereinafter, the GPUs 400(1), 400(2), 400(3), and 400(4) will be referred to as GPU1, GPU2, GPU3, and GPU4, respectively, and the CXL memories 600(1) and 600(2) will be referred to as CXL memory 1 and CXL memory 2, respectively. For simplicity, it is assumed that the bandwidth of all PCIe buses is the same, B [MB / s]. For example, each PCIe bus is a ×16 slot.

[0025] The host memory 900 is connected to the root complex 700 via a memory bus MBUS. The CPU 800 is connected to the root complex 700 via a system bus SBUS, and controls the entire information processing device 100. For example, a control program for the information processing device 100 executed by the CPU 800 may be stored in the host memory 900.

[0026] The FPGA 200 includes an internal memory 201 having a bandwidth equal to or greater than the bandwidth of the PCIe bus, and a transfer control unit 300. The transfer control unit 300 is an example of a data transfer control device. Note that the transfer control unit 300 may be provided outside the FPGA 200 and within the information processing device 100.

[0027] GPU1-GPU4 each have internal memories 401(1), 401(2), 401(3), and 401(4) with a bandwidth equal to or greater than the bandwidth of the PCIe bus. Therefore, the processing performance of the information processing device 100 is not limited by the transfer speed of data input / output to / from the internal memories 201, 401(1), 401(2), 401(3), and 401(4). The internal memories 401(1), 401(2), 401(3), and 401(4) are examples of second memories. Hereinafter, the internal memories 401(1), 401(2), 401(3), and 401(4) will be referred to as internal memories 1, 2, 3, and 4, respectively.

[0028] The FPGA 200 receives four stream data 1-4 in parallel from outside the information processing device 100. For example, each of the stream data 1-4 is video data from a single surveillance camera. The video data includes multiple event data, such as multiple consecutive frames. Note that the number of stream data received in parallel by the FPGA 200 is not limited to four, as long as the total reception speed is equal to or less than the bandwidth of the PCIe bus.

[0029] The FPGA 200 performs preprocessing on the received stream data 1-4, respectively, and generates preprocessed stream data 1-4. For example, the preprocessing may be decoding, filtering, or resizing of the stream data 1-4. The FPGA 200 stores the preprocessed stream data 1-4 in a pre-specified memory such as the internal memory 201. The memory in which each stream data 1-4 is stored is either the internal memory 201, the host memory 900, the internal memories 1-4, or the CXL memory 1-2, and is specified by the transfer control unit 300.

[0030] For ease of explanation, it is assumed that the reception speed of each of the multiple data streams 1-4 received by the FPGA 200 is the same as the generation speed of each of the multiple data streams 1-4 generated by the FPGA 200, which is b [MB / s]. It is also assumed that the transfer speed of each of the data streams 1-4 on each PCIe bus is the same as the generation speed of each of the data streams 1-4, which is b [MB / s].

[0031] Each PCIe bus is capable of transferring up to five streams of data in parallel. Therefore, if the bandwidth of each PCIe bus is B [MB / s], the transfer speed of one stream of data is 0.2B [MB / s]. In other words, the bandwidth of each PCIe bus is 5b [MB / s]. In the following description, the bandwidth of each PCIe bus is assumed to be 5b [MB / s].

[0032] Each of GPU1-GPU4 operates as an accelerator controlled by CPU 800, for example. Of the stream data generated by FPGA 200 in preprocessing, each of GPU1-GPU4 reads a predetermined number of stream data from the storage memory and performs data processing such as AI inference processing. For example, if stream data 1-4 is video data, the inference processing is image recognition. The result data of the data processing by each of GPU1-GPU4 is transferred to CPU 800, for example.

[0033] In order to prevent frame loss in the stream data 1-4 preprocessed by FPGA200, each of GPU1-GPU4 has the ability to process the stream data at a speed equal to or greater than the generation speed (bMB / s) of the stream data 1-4 by FPGA200.

[0034] The size of the result data from the data processing by each of GPU1-GPU4 is sufficiently small compared to the size of the stream data 1-4. For example, the transfer speed of the result data from each of GPU1-GPU4 to CPU800 is less than 1 / 100 of the generation speed of each of the stream data 1-4 by FPGA200. For this reason, the following description will be given assuming that the transfer of the result data from each of GPU1-GPU4 to CPU800 does not affect the transfer of the stream data 1-4 processed by each of GPU1-GPU4. For example, the bandwidth of each PCIe bus is actually B [MB / s] + α, where α is used for transferring the result data from each of GPU1-GPU4 to CPU800.

[0035] The CXL memory 1-2 can temporarily store stream data to be processed by GPU1-GPU4. The data input / output speed (bandwidth) of the CXL memory 1-2 is equal to or greater than the bandwidth of the PCIe bus, and in the example of Figure 3, it is assumed to be 5b [MB / s].

[0036] 3, the FPGA 200 preprocesses each of the stream data 1-4 received from outside the information processing device 100 and then stores the preprocessed data in an internal memory 201 within the FPGA 200. The GPU1-GPU4 reads one of the stream data 1-4 in parallel from the internal memory 201 and processes it. In this case, the four stream data 1-4 are transferred in parallel from the FPGA 200 (internal memory 201) to the PCIe switch 500.

[0037] Therefore, the transfer rate per unit time of the stream data 1-4 on the PCIe bus between the FPGA 200 and the PCIe switch 500 is 4b [MB / s], which is lower than the bandwidth (5b [MB / s]) of the PCIe bus. Therefore, each of the stream data 1-4 is transferred from the FPGA 200 to the PCIe switch 500 without any data dropout.

[0038] Furthermore, one of the four stream data 1-4 is transferred from the PCIe switch 500 to each of GPU1-GPU4. Therefore, the transfer rate per unit time of the stream data on the PCIe bus between the PCIe switch 500 and each of GPU1-GPU4 is b [MB / s], which is lower than the bandwidth of the PCIe bus.

[0039] 3, each of GPU1 to GPU4 can perform data processing such as inference processing without causing data loss, such as missing frames, in stream data 1 to 4 that are preprocessed in parallel by FPGA 200. As a result, each of GPU1 to GPU4 can transfer normal result data of the data processing to CPU 800, and information processing device 100 can operate normally.

[0040] 4 shows an example in which stream data 1-4 are not transferred normally in the information processing device 100 of FIG. 3. In FIG. 4, the FPGA 200 preprocesses four pieces of stream data 1-4 received from outside the information processing device 100 and then stores them in the internal memory 201 of the FPGA 200. The GPU 1 reads the four pieces of stream data 1-4 from the internal memory 201 and processes them. The GPU 2 reads three pieces of stream data 1-3 from the internal memory 201 and processes them. The GPU 3 reads two pieces of stream data 1-2 from the internal memory 201 and processes them. The GPU 4 reads one piece of stream data 1 from the internal memory 201 and processes it.

[0041] In this case, 10 stream data items are transferred in parallel from the FPGA 200 (internal memory 201) to the PCIe switch 500. Therefore, the amount of stream data transferred per hour (performance requirement) required on the PCIe bus between the FPGA 200 and the PCIe switch 500 is 10b [MB / s], which is twice the bandwidth of the PCIe bus (5b [MB / s]). The performance requirement for the amount of data transferred per hour for each of the stream data items 1-4 is indicated in parentheses on the PCIe bus connected to the FPGA 200. Hereinafter, the amount of data transferred per hour is also referred to as the transfer rate.

[0042] In reality, the transfer speed of each of the stream data 1-4 transferred from the FPGA 200 to the PCIe switch 500 is limited by the bandwidth of the PCIe bus, and is therefore half of the performance requirement. Therefore, the total transfer speed of the 10 stream data is 5b [MB / s], which is equal to the bandwidth of the PCIe bus. Furthermore, the total transfer speed of the 10 stream data from the PCIe switch 500 to GPU1-GPU4 is equal to the actual transfer speed of the 10 stream data transferred from the FPGA 200 to the PCIe switch 500.

[0043] As a result, the transfer rate of the four stream data 1-4 transferred from the PCIe switch 500 to GPU1 is 2b [MB / s]. The transfer rate of the three stream data 1-3 transferred from the PCIe switch 500 to GPU2 is 1.5b [MB / s]. The transfer rate of the two stream data 1-2 transferred from the PCIe switch 500 to GPU3 is b [MB / s]. The transfer rate of the single stream data 1 transferred from the PCIe switch 500 to GPU4 is 0.5b [MB / s]. As a result, each of GPU1-GPU4 will be processing stream data with half of the data being dropped, making it difficult to perform normal data processing.

[0044] Fig. 5 shows examples of various tables provided in the transfer control unit 300 of Fig. 4. The transfer control unit 300 has a data amount management table 301 for stream data, a transfer destination management table 302, an area management table 303, and a memory management table 304. Fig. 5 shows the initial state of each table set by the transfer control unit 300.

[0045] The data amount management table 301 has entries that hold flow rates indicating the generation speeds of the stream data 1-4 generated by the FPGA 200 through preprocessing. In the example of Fig. 5, the flow rate of each stream data 1-4 is b [MB / s]. The transfer control unit 300 adds or deletes entries to or from the data amount management table 301 when adding or deleting stream data received from outside the information processing device 100.

[0046] The transfer destination management table 302 has entries for each of stream data 1-4, each of which holds information identifying the GPU that uses the stream data. In the example of Fig. 5, stream data 1 is used by GPU1-GPU4, and stream data 2 is used by GPU1-GPU3. Stream data 3 is used by GPU1-GPU2, and stream data 4 is used by GPU1.

[0047] The transfer control unit 300 adds or deletes an entry in the transfer destination management table 302 when adding or deleting stream data received from outside the information processing device 100. Furthermore, the transfer control unit 300 updates the corresponding entry in the transfer destination management table 302 when adding or deleting a GPU workload that uses each stream data.

[0048] The area management table 303 has entries that hold information indicating a storage destination device for the stream data generated by the FPGA 200 through preprocessing for each of the stream data 1-4, and information indicating an address area of the PCIe bus where the stream data is stored. The memories of various devices connected to the PCIe bus are allocated to the address space of the PCIe bus. An example of the address space of the PCIe bus is shown in FIG. 6. The area management table 303 is set by the transfer control unit 300, copied by each of GPU1-GPU4, and held in each of GPU1-GPU4.

[0049] In the example of FIG. 4, the size of each of stream data 1-4 preprocessed by the FPGA 200 is 0x100. Preprocessed stream data 1 is stored in the address range 0x3000-0x30FF of the internal memory 201 of the FPGA 200. Preprocessed stream data 2 is stored in the address range 0x3100-0x31FF of the internal memory 201 of the FPGA 200. Preprocessed stream data 3 is stored in the address range 0x3200-0x32FF of the internal memory 201 of the FPGA 200. Preprocessed stream data 4 is stored in the address range 0x3300-0x33FF of the internal memory 201 of the FPGA 200.

[0050] The transfer control unit 300 adds or deletes entries in the area management table 303 when adding or deleting stream data received from outside the information processing device 100. Furthermore, the transfer control unit 300 updates the corresponding entries in the area management table 303 when changing the memory of the device that stores each stream data.

[0051] The memory management table 304 has entries for each memory that hold the free capacity of the memory, the free bandwidth (upstream) of the data transfer path from the PCIe switch 500 to the memory, and the free bandwidth (downstream) of the data transfer path from the memory to the PCIe switch 500. In the example of Fig. 4, the host memory 900 and the CXL memories 1 and 2 are not used for transferring each stream data, so the free bandwidth is set to B [MB / s] for both the upstream and downstream directions, which is equal to the bandwidth of the PCIe bus. When changing the memory that stores stream data, the transfer control unit 300 updates the free capacity and free bandwidth (upstream, downstream) of the memory to be stored.

[0052] Fig. 6 shows an example of an address space of the PCIe bus in the information processing device 100 of Fig. 3. Memory areas of the host memory 900, CXL memories 1 and 2, FPGA 200, and GPU1 to GPU4 are allocated to the address space of the PCIe bus. Note that the allocation of the address space shown in Fig. 6 is an example, and addresses allocated to each device are not limited to the example shown in Fig. 6.

[0053] Fig. 7 shows examples of various tables provided in the transfer control unit 300 to solve the problem shown in Fig. 4. The data amount management table 301 and the transfer destination management table 302 are the same as those in Fig. 5. To solve the problem shown in Fig. 4, the transfer control unit 300 determines that the FPGA 200 should store preprocessed stream data 1 in CXL memory 1 and preprocessed stream data 2 in CXL memory 2.

[0054] Therefore, in the entry for stream data 1 in the area management table 303, the transfer control unit 300 stores information indicating CXL memory 1 and information indicating the address range 0x0000-0x00FF. Also, in the entry for stream data 2 in the area management table 303, the transfer control unit 300 stores information indicating CXL memory 2 and information indicating the address range 0x1000-0x01FF. Note that, because the preprocessed stream data 3 and 4 are stored in the internal memory 201 of the FPGA 200, the entries for stream data 3 and 4 in the area management table 303 are maintained in the state shown in FIG.

[0055] Furthermore, the transfer control unit 300 changes the free space and the free upstream and downstream bandwidths of the entry for CXL memory 1 in the memory management table 304 in accordance with the updated area management table 303. Furthermore, the FPGA 200 changes the free space and the free upstream and downstream bandwidths of the entry for CXL memory 2 in the memory management table 304 in accordance with the updated area management table 303. The change in the free bandwidths will be described with reference to FIG. 8.

[0056] Fig. 8 shows an example of an operation for solving the problem shown in Fig. 4 in the information processing device 100 of Fig. 3. A detailed description of the same operations as those in Fig. 4 will be omitted. The operation shown in Fig. 8 is performed by the FPGA 200 and GPU1 to GPU4 based on the information in the various tables shown in Fig. 7 set by the transfer control unit 300.

[0057] The FPGA 200 stores preprocessed stream data 1 in the CXL memory 1, and stores preprocessed stream data 2 in the CXL memory 2. GPU1-GPU4 read the stream data 1 from the CXL memory 1. GPU1-GPU3 read the stream data 2 from the CXL memory 2.

[0058] Because GPU1 to GPU4 do not read stream data 1 from internal memory 201 of FPGA 200, only one stream data 1 needs to be sent from FPGA 200, to CXL memory 1. Similarly, because GPU1 to GPU3 do not read stream data 2 from internal memory 201 of FPGA 200, only one stream data 2 needs to be sent from FPGA 200, to CXL memory 2. As a result, the sum of the performance requirements of the PCIe bus between FPGA 200 and PCIe switch 500 becomes 5b [MB / s], which is the same as the bandwidth of the PCIe bus, and the bandwidth of the PCIe bus can be satisfied.

[0059] The CXL memory 1 receives one stream data 1 and transmits four stream data 1 to each of the GPUs 1 to 4 in response to transfer requests from the GPUs 1 to 4. Therefore, the sum of the performance requirements of the PCIe bus between the CXL memory 1 and the PCIe switch 500 is 5b [MB / s], which is the same as the bandwidth of the PCIe bus.

[0060] The CXL memory 2 receives one stream data 2 and transmits three stream data 2 to each of GPU1 to GPU3 in response to transfer requests from GPU1 to GPU3. Therefore, the sum of the performance requirements of the PCIe bus between the CXL memory 2 and the PCIe switch 500 is 4b [MB / s], which is smaller than the bandwidth of the PCIe bus (5b [MB / s]).

[0061] Each of GPU1 to GPU4 receives a maximum of four stream data via the PCIe switch 500. Therefore, the sum of the performance requirements of the PCIe bus between the PCIe switch 500 and each of GPU1 to GPU4 is 4b [MB / s] or less, which is smaller than the bandwidth of the PCIe bus (5b [MB / s]).

[0062] Therefore, each of GPU1 to GPU4 can receive stream data and perform data processing such as inference processing without causing data loss, such as missing frames, of the stream data that is preprocessed in parallel by FPGA 200. As a result, each of GPU1 to GPU4 can transfer normal result data of the data processing to CPU 800, and information processing device 100 can operate normally.

[0063] Fig. 9 shows an example of an operational flow for determining a transfer path for the stream data shown in Fig. 8. The operation shown in Fig. 9 is performed by the transfer control unit 300 when adding and deleting stream data, when adding and deleting GPU workloads that use each stream data, and when adding and deleting CXL memory.

[0064] First, in step S100, the transfer control unit 300 determines whether or not data loss occurs in the transfer of stream data from the FPGA 200 to the GPU 400. That is, the transfer control unit 300 determines whether or not the amount of stream data transferred per unit time from the FPGA 200 to the PCIe switch 500 is greater than the bandwidth of the PCIe bus. If data loss occurs, the transfer control unit 300 performs step S102 because it is difficult for the GPU 400 to process data normally. If no data loss occurs, the transfer control unit 300 ends the operation shown in FIG. 9 because it is possible for the GPU 400 to process data normally.

[0065] In step S102, the transfer control unit 300 refers to the data amount management table 301 and the transfer destination management table 302. Then, the transfer control unit 300 determines, from among the stream data held in the internal memory 201, the stream data with the largest flow rate (i.e., the transfer amount per unit time) to be transferred to each GPU 400. By referring to the data amount management table 301 and the transfer destination management table 302, the transfer control unit 300 can determine the stream data with the largest flow rate by a simple calculation.

[0066] Here, the transfer control unit 300 calculates the transfer flow rate using equation (1). The flow rate of stream data transferred to one GPU 400 is calculated as the product of the number of transfers of the stream data per unit time and the size of the stream data. (Stream data flow rate transferred to one GPU 400) × (number of GPUs 400 receiving the stream data) (1)

[0067] Next, in step S104, the transfer control unit 300 refers to the data amount management table 301, the transfer destination management table 302, and the memory management table 304. Then, the transfer control unit 300 checks whether the stream data determined in step S102 can be moved from the FPGA 200 to a memory outside the FPGA 200. The memories to be checked are the CXL memories 1 and 2 and the host memory 900. Here, "movable" means that the memory to be moved has free capacity to move the stream data and satisfies the performance requirements (bandwidth) of all GPUs 400 that use the stream data moved to the memory.

[0068] Next, in step S106, if there is memory to which the stream data can be moved based on the processing result of step S104, the transfer control unit 300 performs step S108. If there is no memory to which the stream data can be moved, the transfer control unit 300 ends the operation shown in Fig. 9. In this case, the transfer control unit 300 distributes and processes the multiple stream data among multiple information processing devices 100, so that data loss does not occur when transferring the stream data from the FPGA 200 to the GPU 400.

[0069] If there are multiple memories to which stream data can be transferred, the transfer control unit 300 selects one of the memories based on a preset criterion. For example, the transfer control unit 300 selects the memory with the largest free space. The transfer control unit 300 determines whether the free space of the CXL memories 1 and 2 and the host memory is C. CM1 >C CM2 >C HMIn this case, CXL memory 1 is selected as the storage destination for the stream data. By selecting the memory in descending order of free space, it is possible to reduce the variation in free space among the multiple memories to which data is transferred.

[0070] Next, in step S108, the transfer control unit 300 updates the area management table 303 with information indicating the memory selected in step S104. The transfer control unit 300 also updates the free memory capacity for transferring stream data and the free upstream and downstream bandwidths in the memory management table 304, and returns to step S100. By determining the memory to which stream data is moved in order of stream data flow rate, the number of loops in the operation of Figure 9 can be reduced, and the system can quickly converge to a state where no data drops occur.

[0071] 8, for example, the stream data 2 is transferred from the FPGA 200 to the CXL memory 2, and then read from the CXL memory 2 and processed by the GPU1 to GPU3. When the stream data 2 is transferred to the CXL memory 2, the transfer time of the stream data from the FPGA 200 to the GPU1 to GPU3 is longer than in the operation shown in FIG.

[0072] If the transfer latency of the stream data from the FPGA 200 to the GPU 400 does not satisfy the processing latency required for the processing of the GPU 400, it becomes difficult for the GPU 400 to perform normal processing. Below, we will explain a solution for when the transfer latency of the stream data 2 transferred via the CXL memory 2 in Figure 8 does not satisfy the processing latency of the GPU 3.

[0073] FIG. 10 shows an example of calculating the latency until stream data 2 arrives at GPU1-GPU3 from FPGA 200 in the data transfer path of FIG. 8. For simplicity, assume that the data size of one event, such as a frame, in each stream data is b [MB]. The transfer rate of each stream data is b [MB / s], so in this case, an event is generated every second. The bandwidth of the PCIe bus is B [MB / s] (= 5b [MB / s]). Below, the bandwidth of the PCIe bus used for transferring stream data is expressed as the number of stream data.

[0074] On the transfer path of the stream data 2 from the FPGA 200 to the CXL memory 2, five pieces of stream data are transferred from the FPGA 200 to the PCIe switch 500, and one piece of stream data is transferred from the PCIe switch 500 to the CXL memory 2. Therefore, the bandwidth per piece of stream data from the FPGA 200 to the PCIe switch 500 is B / 5, and the bandwidth per piece of stream data from the PCIe switch 500 to the CXL memory 2 is B / 1. The worst-case latency of the stream data 2 from the FPGA 200 to the CXL memory 2 is b / (B / 5)=5*b / B (the sign * is a multiplication sign), using B / 5, which has a small bandwidth and a large impact on latency.

[0075] On the transfer path of stream data 2 from the CXL memory 2 to the GPU 3, three stream data are transferred from the CXL memory 2 to the PCIe switch 500, and two stream data are transferred from the PCIe switch 500 to the GPU 3. Therefore, the bandwidth per stream data from the CXL memory 2 to the PCIe switch 500 is B / 3, and the bandwidth per stream data from the PCIe switch 500 to the GPU 3 is B / 2.

[0076] The worst-case latency of stream data 2 from CXL memory 2 to GPU 3 is b / (B / 3) = 3*b / B, using B / 3, which has a small bandwidth and a large impact on latency. Therefore, the latency of stream data 2 from FPGA 200 to GPU 3 is 5*b / B + 3*b / B = 8*b / B.

[0077] In the transfer path of the stream data 2 from the FPGA 200 to the GPU 1, the worst latency of the stream data 2 from the FPGA 200 to the CXL memory 2 is b / (B / 5)=5*b / B, which is the same as the latency calculated for the GPU 3.

[0078] In the transfer path of stream data 2 from CXL memory 2 to GPU 1, the bandwidth per stream data from CXL memory 2 to PCIe switch 500 is B / 3, the same as the bandwidth calculated for GPU 3. The bandwidth per stream data from PCIe switch 500 to GPU 1 is B / 4. Therefore, the worst-case latency of stream data 2 from CXL memory 2 to GPU 1 is b / (B / 4) = 4*b / B, using B / 4, which has a small bandwidth and a large impact on latency. Therefore, the latency of stream data 2 from FPGA 200 to GPU 1 is 5*b / B + 4*b / B = 9*b / B.

[0079] In the transfer path of the stream data 2 from the FPGA 200 to the GPU 2, the worst latency of the stream data 2 from the FPGA 200 to the CXL memory 2 is b / (B / 5)=5*b / B, which is the same as the latency calculated for the GPU 3.

[0080] In the transfer path of stream data 2 from CXL memory 2 to GPU 2, the bandwidth per stream data from CXL memory 2 to PCIe switch 500 is B / 3, the same as the bandwidth calculated for GPU 3. The bandwidth per stream data from PCIe switch 500 to GPU 2 is also B / 3. Therefore, the worst-case latency of stream data 2 from CXL memory 2 to GPU 1 is b / (B / 3)=3*b / B, using B / 3. Therefore, the latency of stream data 2 from FPGA 200 to GPU 1 is 5*b / B+3*b / B=8*b / B.

[0081] The required processing latency may differ for each of the processes performed by GPU1-GPU4. For example, when stream data is used for video processing for autonomous driving of a vehicle, the stream data processed by the GPU must be immediately fed back to the vehicle control, so low latency, such as on the order of milliseconds, is required. On the other hand, when stream data is used for processing video from a surveillance camera that is checked by a human, high latency, such as on the order of seconds or minutes, is acceptable.

[0082] For example, suppose that low latency is not required for the processing of stream data 2 by GPU1 and GPU2 (e.g., latency = 10*b / B), while low latency is required for the processing of stream data 2 by GPU3 (e.g., latency = 5*b / B). In this case, in the transfer path shown in FIG. 8, the latencies of stream data 2 from FPGA to GPU1 and GPU2, "9*b / B" and "8*b / B", satisfy the latency performance requirement of GPU1 and GPU2, "10*b / B". On the other hand, the latency of stream data 2 from FPGA to GPU3, "8*b / B", does not satisfy the latency performance requirement of GPU3, "5*b / B".

[0083] GPU1=9*b / B(≦10*b / B) GPU2=8*b / B(≦10*b / B) GPU3=8*b / B(>5*b / B) The latency of stream data 2 from FPGA to GPU 3, "8*b / B," is an example of the first latency. The latency performance requirement of GPU 3, "5*b / B," is an example of the preset second latency. Stream data 2 is an example of excess data in which the first latency exceeds the second latency.

[0084] FIG. 11 shows examples of various tables provided in the transfer control unit 300 to solve the problem of unsatisfied latency. Detailed description of elements similar to those in FIG. 7 will be omitted. Each table in FIG. 11 shows the state when the stream data transfer operation of FIG. 8 is performed. The data amount management table 301, transfer destination management table 302, and area management table 303 are the same as those in FIG. 7. The memory management table 304 is obtained by adding entries for internal memories 1-3 of GPU1-GPU4 to the memory management table 304 of FIG. 7. The latency management table 305 is newly provided with respect to FIG. 7.

[0085] In the memory management table 304, the state of the entries for the host memory and CXL memories 1 and 2 is the same as in FIG. 7. In FIG. 8, the internal memory 3 of GPU1 to GPU4 is not used instead of the internal memory 201 of FPGA 200. Therefore, in the memory management table 304, the free space C G1 , C G2 , C G3 , C G4 are set to the storage capacities of the internal memories 1-4 of GPU1-GPU4, respectively.

[0086] 8, GPU1 receives stream data of 4b (=0.8B) from the PCIe switch 500, but does not transmit stream data to the PCIe switch 500. Therefore, in the entry for GPU1 in the memory management table 304, the available upstream bandwidth is set to 0.2B (=B-0.8B), and the available downstream bandwidth is set to B (=B-0).

[0087] GPU2 receives stream data of 3b (=0.6B) from the PCIe switch 500, and does not transmit stream data to the PCIe switch 500. Therefore, in the entry for GPU2 in the memory management table 304, the available upstream bandwidth is set to 0.4B (=B-0.6B), and the available downstream bandwidth is set to B (=B-0).

[0088] GPU3 receives stream data of 2b (=0.4B) from the PCIe switch 500, and does not transmit stream data to the PCIe switch 500. Therefore, in the entry for GPU2 in the memory management table 304, the available upstream bandwidth is set to 0.6B (=B-0.4B), and the available downstream bandwidth is set to B (=B-0).

[0089] GPU4 receives stream data of b (=0.2B) from the PCIe switch 500 and does not transmit stream data to the PCIe switch 500. Therefore, in the entry for GPU2 in the memory management table 304, the available upstream bandwidth is set to 0.8B (=B-0.2B) and the available downstream bandwidth is set to B (=B-0).

[0090] The latency management table 305 has entries for GPU1 to GPU4 that store the latency performance requirements of each of stream data 1 to 4. In the latency management table 305, the entries for GPU1 to GPU3 store the latency performance requirements of "10*b / B", "10*b / B", and "5*b / B" for stream data 2 described in Fig. 10, respectively. In the latency management table 305, the latency performance requirements of stream data 1, 2, and 4 are sufficiently large, so each entry stores information indicating "no requirement".

[0091] 12 shows an example of an operation for solving the problem of unsatisfied latency in the information processing device 100 of FIG. 3. Detailed description of operations similar to those of FIGS. 4 and 8 will be omitted. In FIG. 12, the FPGA 200 stores preprocessed stream data 2 in the internal memory 3 of the GPU 3. The GPU 3 processes the stream data 2 stored in the internal memory 3. The GPU 1 and GPU 2 read the stream data 2 from the internal memory 3 of the GPU 3 and process it. The transfer paths of the stream data 1, 3, and 4 are the same as those in FIG. 8.

[0092] 12, the stream data 2 is transferred from the FPGA 200 to the internal memory 3 of the GPU 3 without passing through the CXL memory 2. Therefore, it is possible to eliminate the transfer latency of the stream data 2 between the PCIe switch 500 and the CXL memory 2 that occurs when the stream data 2 passes through the CXL memory 2 shown in FIG.

[0093] Figure 13 shows example settings of various tables in the transfer control unit 300 for the transfer path of stream data 1-4 shown in Figure 12. Figure 13 is the same as Figure 11 except that the settings of area management table 303 and memory management table 304 are different from those in Figure 11. Entries shown in shaded areas indicate changes made to Figure 11.

[0094] In the area management table 303, the device for storing the stream data 2 is set to GPU 3, and the address area for storing the stream data 2 is set to the address area 0x6000-0x6100 allocated to the internal memory 3 of GPU 3. Other settings in the area management table 303 are the same as those in FIG. 11.

[0095] In the memory management table 304, the CXL memory 2 that is not used for transferring the stream data is set to the initial state. Also, the free space of the internal memory 3 of the GPU 3 that is used for transferring the stream data 2 is set to "C G3 -0x100", and the available downstream bandwidth of GPU3 is set to 0.6B. Other settings in the memory management table 304 are the same as those in FIG.

[0096] Fig. 14 shows an example of calculating the latency until the stream data 2 arrives at GPU1-GPU3 from FPGA 200 in the data transfer path of Fig. 12. Detailed description of the same elements as in Fig. 10 will be omitted.

[0097] In the latency from FPGA 200 to GPU 3, stream data 2 is stored in internal memory 3 of GPU 3 from FPGA 200 via PCIe switch 500. Therefore, the latency from FPGA 200 to GPU 3 is b / (B / 5)=5*b / B, using only B / 5, which has a small bandwidth and a large impact on latency. Therefore, by storing stream data 2 in internal memory 3 of GPU 3, the latency performance requirement of "5*b / B" for stream data 2 of GPU 3 set in latency management table 305 can be satisfied.

[0098] In terms of latency from FPGA200 to GPU1, the worst latency of stream data 2 from FPGA200 to GPU3 along the transfer path of stream data 2 is b / (B / 5) = 5*b / B, the same as the latency calculated for GPU3. In terms of the transfer path of stream data 2 from GPU3 to GPU1, the bandwidth per stream data from GPU3 to PCIe switch 500 is B / 2. The bandwidth per stream data from PCIe switch 500 to GPU1 is B / 4.

[0099] Therefore, the worst-case latency of stream data 2 from GPU3 to GPU1 is b / (B / 4)=4*b / B, using B / 4, which has a small bandwidth and a large impact on latency. As a result, the latency of stream data 2 from FPGA 200 to GPU1 is 5*b / B+4*b / B=9*b / B. Therefore, the latency performance requirement of "10*b / B" for stream data 2 of GPU1 set in the latency management table 305 can be satisfied.

[0100] In terms of latency from FPGA200 to GPU2, the worst latency of stream data 2 from FPGA200 to GPU3 along the transfer path of stream data 2 is b / (B / 5) = 5*b / B, the same as the latency calculated for GPU3. In terms of the transfer path of stream data 2 from GPU3 to GPU2, the bandwidth per stream data from GPU3 to PCIe switch 500 is B / 2, the same as the bandwidth calculated for GPU1. The bandwidth per stream data from PCIe switch 500 to GPU2 is B / 3.

[0101] Therefore, the worst-case latency of stream data 2 from GPU3 to GPU2 is b / (B / 3)=3*b / B, using B / 3, which has a small bandwidth and a large impact on latency. As a result, the latency of stream data 2 from FPGA 200 to GPU2 is 5*b / B+3*b / B=8*b / B. Therefore, the latency performance requirement of "10*b / B" for stream data 2 of GPU2 set in the latency management table 305 can be satisfied.

[0102] Fig. 15 shows an example of an operational flow for determining a transfer path for the stream data shown in Fig. 12. Detailed description of operations similar to those in Fig. 9 will be omitted. The operations shown in Fig. 15 are performed by the transfer control unit 300 when adding and deleting stream data, when adding and deleting GPU workloads that use each stream data, and when adding and deleting CXL memory.

[0103] First, in step S200, the transfer control unit 300 determines whether or not there is a GPU that does not satisfy the latency requirements by referring to the latency management table 305. That is, the transfer control unit 300 determines whether or not there is excess data, which means that the latency of the stream data from the FPGA 200 to each GPU 400 exceeds the latency performance requirement.

[0104] If there is a GPU 400 that does not satisfy the latency requirement, it is difficult for the GPU 400 to perform normal processing of the stream data that does not satisfy the latency requirement, so the transfer control unit 300 performs step S202. If there is no GPU 400 that does not satisfy the latency requirement, the transfer control unit 300 can perform normal data processing of the stream data by the GPU 400, so it ends the operation shown in FIG.

[0105] In step S202, the transfer control unit 300 refers to the data amount management table 301, the transfer destination management table 302, the memory management table 304, and the latency management table 305. The transfer control unit 300 then checks whether it is possible to move the stream data that does not satisfy the latency requirements to the internal memory 401 of the GPU 400 that does not satisfy the latency requirements. If there are multiple GPUs 400 that do not satisfy the latency requirements, the transfer control unit 300 selects one of the multiple GPUs 400. Here, "movable" means that the memory to which the stream data is to be moved has free capacity to move the stream data and satisfies the performance requirements (bandwidth) of all GPUs 400 that use the stream data moved to the memory.

[0106] Next, in step S204, if the transfer control unit 300 is able to move the stream data to any of the internal memories 401 of the GPU 400 based on the processing result of step S202, the transfer control unit 300 performs step S206. If the latency requirement cannot be satisfied even if the stream data is moved to any of the internal memories 401 of the GPU 400, the transfer control unit 300 ends the operation shown in Fig. 15. In this case, the latency requirement is satisfied by distributing and processing the multiple stream data among multiple information processing devices 100, for example.

[0107] Next, in step S206, the transfer control unit 300 updates the area management table 303 and the memory management table 304 with respect to those in FIG. 11, as shown in FIG. 13. This allows the transfer destination of the stream data 2 that does not satisfy the latency requirement to be changed to the internal memory 3 of the GPU 3 selected in step S202. The transfer control unit 300 then returns to the operation of step S100. Note that the operation shown in FIG. 15 may be performed after the operation shown in FIG. 9, or may be performed together with the operation shown in FIG. 9.

[0108] As described above, this embodiment can also achieve the same effects as the above-described embodiments. For example, stream data can be transferred from the FPGA 200 to the multiple GPUs 400 without exceeding the bandwidth of the PCIe bus connecting the FPGA 200 and the PCIe switch 500. In other words, even when the number of parallel stream data transferred from the FPGA 200 to the PCIe switch 500 is reduced to prevent the bandwidth from being exceeded, a predetermined number of data can be transferred to each of the multiple GPUs 400. Therefore, the multiple GPUs 400 can receive and normally process the stream data without any dropouts, allowing the information processing device 100 to operate normally.

[0109] Furthermore, in this embodiment, if the transfer latency of the stream data does not satisfy the processing latency of the GPU 400, the FPGA 200 transfers the stream data from the FPGA 200 to the internal memory 401 of the GPU 400 without passing through the CXL memory 600. Because the stream data is not held in the CXL memory 600, it is possible to eliminate the transfer latency of the stream data between the PCIe switch 500 and the CXL memory 600 that occurs when the stream data passes through the CXL memory 600. As a result, it is possible to set the transfer latency of the stream data to a value equal to or less than the latency performance requirement, and the GPU 400 can normally receive and process the stream data without dropping any data, thereby enabling the information processing device 100 to operate normally.

[0110] The transfer control unit 300 is also mounted within the FPGA 200 that receives the stream data. This allows the transfer control unit 300 to set information in various tables, such as the data amount management table 301 and the transfer destination management table 302, using setting information set in the FPGA 200. In other words, the transfer control unit 300 can set information in various tables without communicating with the FPGA 200 that is located externally.

[0111] The features and advantages of the embodiments will be apparent from the above detailed description. It is intended that the claims encompass the features and advantages of the above-described embodiments without departing from the spirit and scope of the claims. Furthermore, any improvements and modifications will be readily apparent to those skilled in the art. Therefore, it is not intended that the scope of the inventive embodiments be limited to the above-described embodiments, and appropriate improvements and equivalents within the scope of the disclosed embodiments may be utilized. [Explanation of symbols]

[0112] 10. Information processing equipment 20 1st device 21 Internal Memory 30 Data transfer control device 40 Second device 41 Internal Memory 50 Switch 60 memory 100 Information processing device 200 FPGA 201 Internal Memory 300 Transfer control section 301 Data Volume Management Table 302 Forwarding destination management table 303 Area Management Table 304 Memory Management Table 305 Latency Management Table 400 GPU 401 Internal Memory 500 PCIe Switch 600 CXL memory 700 Root Complex 800 CPU 900 host memory

Claims

1. A data transfer control device is mounted on an information processing device having a first device that receives a plurality of data items, a plurality of second devices that receive and process the plurality of data items transmitted in parallel from the first device, one or more first memories that can hold the data items, and a switch that interconnects the first device, the second device, and the first memory, and controls transfer of the plurality of data items from the first device to the plurality of second devices, When the transfer amount per unit time of the plurality of data from the first device to the switch exceeds the bandwidth between the first device and the switch, the data to be transferred from the first device to the plurality of second devices is transferred to the first memory, and the data held in the first memory is acquired by the plurality of second devices. Data transfer control device.

2. When a plurality of pieces of data are transferred from the first device to a plurality of second devices, and the transfer amount per unit time exceeds the bandwidth, a process of determining the first memory to which the data is to be transferred from the plurality of first memories in descending order of the data transfer amount per unit time from the first device, is repeated until the transfer amount per unit time becomes equal to or less than the bandwidth.

2. The data transfer control device according to claim 1.

3. a memory management table that holds information indicating the free space of each of the plurality of first memories; When a plurality of the data are transferred from the first device to a plurality of the second devices, and the transfer amount per unit time exceeds the bandwidth, the memory management table is referenced to determine the first memory to which the data is to be transferred from the first device in descending order of free space.

3. The data transfer control device according to claim 2.

4. a data amount management table that stores information indicating the data amount per unit time of the plurality of pieces of data received by the first device; a transfer destination management table that holds information indicating the second device to which each of the plurality of data is to be transferred; By referring to the data amount management table and the transfer destination management table, it is determined whether or not the transfer amount per unit time exceeds the bandwidth, and the data having a larger transfer amount per unit time from the first device than the others is determined.

3. The data transfer control device according to claim 2.

5. When there is excess data such that a first latency until the data is transferred from the first device to the second device via the first memory exceeds a preset second latency, The excess data is transferred to a second memory mounted in any one of the plurality of second devices to which the excess data is transferred, without passing through the first memory.

5. The data transfer control device according to claim 1.

6. a latency management table for storing the second latency for each of the data transferred to each of the plurality of second devices; The latency management table is referenced to determine whether or not there is excess data.

6. The data transfer control device according to claim 5.

7. a first device for receiving a plurality of data; a plurality of second devices that receive and process the plurality of data transmitted in parallel from the first device; one or more first memories capable of holding the data; a switch interconnecting the first device, the second device, and the first memory; a data transfer control device that controls transfer of the plurality of data from the first device to the plurality of second devices, and when the transfer amount of the plurality of data from the first device to the switch per unit time exceeds the bandwidth between the first device and the switch, transfers the data transferred from the first device to the plurality of second devices to the first memory, and causes the plurality of second devices to acquire the data held in the first memory. Information processing device.

8. an area management table for storing information indicating a memory to which each of the plurality of data is transferred from the first device and a transfer destination address; Each of the plurality of second devices refers to the area management table to determine a memory from which to read the data to be processed. The information processing device according to claim 7 .

9. The data transfer control device is included in the first device.

9. The information processing device according to claim 7 or 8.

Citation Information

Patent Citations

  • Computer system provided with segmented memory

    JP1997006711A

  • Information processing device, drawing method, and program

    JP2015176569A

  • IP image transmitter

    WO2004093445A1