Computing System
The computing system addresses the issue of FPGA accelerators not operating normally due to incorrect arithmetic circuit writes by using two accelerators with matching reconfigurable areas, where the second accelerator tests the circuit before it is written to the first accelerator, ensuring reliable operation.
Patent Information
- Application Number
- JP2023529208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-21
AI Technical Summary
In computing systems that utilize FPGA accelerators, incorrectly writing a new arithmetic circuit can lead to the accelerator not operating normally, resulting in processing failures.
A computing system is designed with two accelerators, each having a reconfigurable area with the same circuit layout. The system includes a first computer that writes an arithmetic circuit into a reconfigurable area of a first accelerator and a second computer that writes the same arithmetic circuit into a corresponding area of a second accelerator. The second computer tests the arithmetic circuit before it is written to the first accelerator, ensuring normal operation.
This solution prevents the occurrence of accelerators not operating normally due to incorrect arithmetic circuit writes, thereby ensuring reliable processing operations in computing systems.
Smart Images

Figure 0007683695000001 
Figure 0007683695000002 
Figure 0007683695000003
Abstract
Description
[Technical field]
[0001] The present invention relates to computing systems. [Background technology]
[0002] In recent years, technology has been developed that speeds up processing by having a part of the processing performed by a computer executed by an accelerator with reconfigurable circuits, rather than by a CPU, to realize virtual reality or artificial intelligence on the Internet. Patent Document 1, which discloses such technology, discloses a technology in which a circuit written in an FPGA accelerator of a computer is appropriately rewritten according to the processing to be executed by the FPGA accelerator. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2018-206195 A Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology described in Patent Document 1, a new arithmetic circuit is directly written to the FPGA accelerator. If the new arithmetic circuit is not written correctly to the FPGA accelerator, the FPGA accelerator may not operate correctly.
[0005] An object of the present invention is to prevent the occurrence of a problem in which an accelerator to which data is written does not operate normally when data is written to an arithmetic circuit. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems, a computing system according to the present invention includes a first computer configured to write an arithmetic circuit into a reconfigurable first area provided in a first accelerator, and a second computer configured to write an arithmetic circuit into a reconfigurable second area having the same circuit layout as the first area provided in a second accelerator different from the first accelerator, wherein the second computer is configured to, when the first computer writes a new arithmetic circuit into the first area, write the new arithmetic circuit into a partial area of the second area located at the same position as an unwritten partial area of the first area, and the first computer is configured to not write the new arithmetic circuit into the first area when the new arithmetic circuit is not written normally, and to write the new arithmetic circuit into the unwritten partial area of the first area when the new arithmetic circuit is written normally. Effect of the Invention
[0007] According to the present invention, it is possible to prevent the occurrence of a problem in which an accelerator to which data is written does not operate normally when data is written to an arithmetic circuit. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing the hardware configuration of a computing system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a configuration diagram of the first accelerator and the second accelerator. [Diagram 3] FIG. 3 is a flowchart of the accelerator control management process. [Figure 4] FIG. 4 is a configuration diagram of the first accelerator and the second accelerator. [Diagram 5] FIG. 5 is a configuration diagram of the first accelerator and the second accelerator. [Figure 6] FIG. 6 is a flowchart of the write process. [Figure 7]FIG. 7 is a configuration diagram showing how the write area of the second area of the second accelerator is changed. [Figure 8] FIG. 8 is a hardware configuration diagram of a computing system according to a modified example of FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] A computing system 10 according to an embodiment of the present invention will now be described with reference to the drawings.
[0010] 1, the computing system 10 includes a first computer 20, a second computer 30, and a gateway 40 as a third computer communicatively connected to the computers 20 and 30 via a LAN (Local Area Network) (not shown) or the like. The gateway 40 is connected to a network N such as the Internet. The gateway 40 relays data, such as images in this case, exchanged between the computers 20 and 30 and a client computer C of a user who uses the computing system 10 and is connected to the network N.
[0011] The computers 20 to 40 operate as nodes of a computer network that performs various processes in response to requests from each of a plurality of client computers C. Here, each client computer C transmits an image to the computing system 10 and requests image processing for the image. The computing system 10 performs image processing on the transmitted image and returns the processed image to the client computer C that transmitted the image. As described below, image processing is mainly performed by the first computer 20.
[0012] The first computer 20 includes a central processing unit (CPU) 21, a RAM 22 such as a dynamic random access memory (DRAM) that functions as the main memory of the CPU 21, and a non-volatile storage device 23. The first computer 20 further includes a first accelerator 24 including a field programmable gate array (FPGA), and a network interface card (NIC) 25 that is a network card. The storage device 23 is an auxiliary storage device such as a hard disk or a solid state drive (SSD). When the CPU 21 exchanges data with a gateway or the like outside the first computer 20, the exchange is performed via the NIC 25. The CPU 21 executes a program stored in the storage device 23 and read into the RAM 22 to perform the process described below.
[0013] The second computer 30 is configured by the same computer as the first computer 20 (although the programs and the like are different). As shown in FIG.
[0014] For example, a SYS-4028GR-TR2 server manufactured by Super Micro Computer is used as each of the computers 20 and 30. The CPU motherboard of this server is equipped with two Intel Xeon (registered trademark) CPU processors E5-2600V4 as the CPU, and eight I-O Data Device DDR4-2400 DIMM 32GB memory cards as the RAM. The CPU motherboard is also equipped with a daughter board with a 16-lane slot of PCI Express 3.0 (Gen3), and this slot is equipped with one Xillinx ALVEO U250 as the accelerator, and one Mellanox Technologies ConnectX-4 VPI MCX455A-ECAT as the NIC.
[0015] A plurality of types of image processing are executed in the first computer 20. Some of the executable types of image processing are executed by the CPU 21 executing an image processing program stored in the storage device 23. The rest of the executable types of image processing are executed by an arithmetic circuit configured in a reconfigurable first area of the first accelerator 24 by being written in the first area. The arithmetic circuit is written in the first area by the CPU 21 applying a circuit configuration (bitstream file) representing the arithmetic circuit stored in the storage device 23 to a part of the first area.
[0016] Here, one of the image processes performed by the CPU 21 executing the image processing program stored in the storage device 23 is a pixel sorting process that sorts each pixel of an image. Furthermore, the image process performed by the arithmetic circuit is a grayscale conversion process that converts the pixel-sorted image into a grayscale image.
[0017] The second accelerator 34 of the second computer 30 has a reconfigurable second area with the same circuit arrangement (same arrangement of switch cells, LUT (Look Up Table), and wiring) as the first area of the first accelerator 24. The same arithmetic circuit is written in the second area at the same position as the first accelerator 24. Specifically, the storage device 33 stores a circuit configuration similar to that of the first computer 20. The CPU 21 of the first computer 20 notifies the gateway 40 of the type of arithmetic circuit written in the first accelerator 24 and the write position in the first area. The gateway 40 notifies the second computer 30 of the type of arithmetic circuit and the write position notified from the first computer 20. Based on this notification, the CPU 31 of the second computer 30 applies the same circuit configuration to the second area of the second accelerator 34 and writes the same arithmetic circuit in the same position. In this way, the arithmetic circuits reconfigured in the first area are the same as those in the second area. As will be described later, if the same arithmetic circuit has already been written to the second accelerator 34 at the same position, the arithmetic circuit is not written in the second computer 30.
[0018] A plurality of types of circuit configurations are registered in the storage devices 33 and 43, and here, it is assumed that one of the circuit configurations, which represents an arithmetic circuit that executes grayscale conversion processing, is applied to each of the accelerators 24 and 34. The circuit configurations registered in the storage devices 33 and 43 are also registered in the gateway 40. In other words, the gateway 40 knows the types of arithmetic circuits that can be written to each of the accelerators 24 and 34. It is also assumed that the gateway 40 knows the written areas and unwritten areas of the areas R1 and R2 based on the notification from the CPU 21 of the first computer 20.
[0019] Here, as shown in Fig. 2, the first region R1 and the second region R2 of each of the accelerators 24 and 34 each have nine blocks, 3 x 3. In each of the first region R1 and the second region R2, an arithmetic circuit that executes gray scale conversion processing is written using the four blocks on the upper left side. That is, in Fig. 2, the dotted blocks have arithmetic circuits written in, and the white blocks have arithmetic circuits not written in.
[0020] The gateway 40 is composed of a server computer or the like, and includes a CPU, a main memory, a non-volatile storage device, a NIC, etc. (not shown) The gateway 40 executes the acceleration control management process shown in FIG.
[0021] In the process shown in Fig. 3, the gateway 40 first waits until it receives an image processing request including an image and specification information specifying the type of image processing to process the image from any of the multiple client computers C (step S11). If an image processing request has been received, the gateway 40 inquires of the first computer 20 whether the type of image processing specified by the specification information is possible (step S12). The first computer 20 determines whether it is possible to perform the inquired image processing. If the specification information of the image processing request is pixel sorting processing and grayscale processing, the CPU 21 of the first computer 20 determines that it is possible to perform the image processing of the image processing request, and replies to that effect to the gateway 40.
[0022] If there is a reply indicating that the image processing of the image processing request is possible (step S13; Yes), the gateway 40 supplies the image included in the image processing request to the first computer 20 and instructs the first computer 20 to execute the image processing for which there is a reply that the image processing is possible (step S14). In this case, the CPU 21 executes a program to perform pixel sorting on the received image, and performs grayscale conversion on the sorted image using an arithmetic circuit. The CPU 21 then returns the grayscale converted image to the gateway 40. The gateway 40 returns the returned image via the network N to the client computer C that made the image processing request (step S15).
[0023] If the designated information of the image processing request is, for example, a cropping process of a video source other than pixel sorting process and grayscale process, the CPU 21 of the first computer 20 replies to the gateway 40 that the process cannot be performed. When the gateway 40 receives this reply (step S13; No), the gateway 40 determines whether it is possible to write a new arithmetic circuit that executes the image processing of the image processing request in the first accelerator 24 of the first computer 20 (step S16). In this determination, the gateway 40 determines whether a circuit configuration that represents an arithmetic circuit that performs the cropping process is stored in the storage device 23 of the first computer 20. Furthermore, the gateway 40 determines whether an unwritten area of the first area R1 of the first accelerator 24 includes an area in which this arithmetic circuit can be written.
[0024] If it is not possible to write a new arithmetic circuit, that is, if at least one of the results of the two judgments is negative (step S16; No), the gateway 40 returns a message indicating that processing is not possible to the client computer C that sent the current image processing request (step S17).
[0025] If both of the two judgment results are positive (if the circuit configuration is stored in the storage device 23 and the arithmetic circuit is writable), the gateway 40 issues a command to the second computer 30 to write an arithmetic circuit that performs the type of image processing specified by the specification information of the image processing request, that is, the cutout processing of the video source (step S18). When this command is issued, the CPU 31 of the second computer 30 reads out the circuit configuration representing the arithmetic circuit of this cutout processing from the storage device 33, and applies the read circuit configuration to an unwritten area of the second accelerator 34 to write the arithmetic circuit in this area (see FIG. 4). The write position is specified by the gateway 40. Thereafter, the CPU 31 executes a predetermined program to operate as a test data generation unit that generates test data for this arithmetic circuit, inputs the test data to the arithmetic circuit, tests the arithmetic circuit, and checks the normality of this arithmetic circuit. The test data generation unit may be configured outside the second computer 30, for example, by a CPU in the gateway 40. When the arithmetic circuit is operating normally or not operating normally, the CPU 31 notifies the gateway 40 of that fact.
[0026] When the gateway 40 receives a notification that the arithmetic circuit operates normally (step S19; Yes), the gateway 40 issues a write command to the first computer 20 to write the type of image processing specified by the designation information of the image processing request, that is, the arithmetic circuit that performs the cropping process of the video source, to the first accelerator 24 (step S20). Furthermore, the gateway 40 transmits the image included in the image processing request and an instruction for image processing in the new arithmetic circuit to the first computer 20 (step S21). In response to the write command, the CPU 21 of the first computer 20 applies the circuit configuration to the first accelerator 24 to write the arithmetic circuit (see FIG. 5), inputs the image to the arithmetic circuit, and acquires the image after image processing output by the arithmetic circuit. The CPU 21 returns the acquired image to the gateway 40. The gateway 40 returns the returned image to the client computer C that requested the image processing via the network N (step S22).
[0027] If a notification is received that the arithmetic circuit is not operating normally (step S19; No), the gateway 40 returns a message to the effect that processing is not possible to the client computer C that has sent the current image processing request (step S17).
[0028] Through the above-described series of processes, if an arithmetic circuit for performing image processing requested by the client computer C has not been written to the first accelerator 24 of the first computer 20, the arithmetic circuit is written to the first accelerator 24. When writing, the second computer 30 first writes a new arithmetic circuit to a partial area of the second area R2 at the same position as the partial area of the first area R1 where no arithmetic circuit has been written. Then, the first computer 20 writes the new arithmetic circuit to the partial area of the first area R1 where no arithmetic circuit has been written only when the new arithmetic circuit has been written successfully to the second area R2. This increases the possibility that the arithmetic circuit will be written successfully to the first accelerator 24, and makes it difficult for the first accelerator 24 to operate normally due to the arithmetic circuit not being written successfully when the arithmetic circuit is directly written to the first accelerator 24.
[0029] Furthermore, in the above, when a new arithmetic circuit is written in the second region R2, the arithmetic circuit is operated. Then, when the arithmetic circuit operates normally, the arithmetic circuit is considered to have been written normally in the second region R2, and is written in the first accelerator 24 of the first computer 20. As a result, even when, for example, a user A requests pixel sorting and grayscale conversion processing from the computing system 10, and a user B requests a video source cropping process during the pixel sorting and grayscale processing, the arithmetic circuit is written suitably. That is, even when the first accelerator 24 is operating or the CPU 21 of the first computer 20 is executing other processing, an operation for testing the new arithmetic circuit is performed in the second accelerator 34 of the second computer 30, so that the influence of the test operation on the first accelerator 24 and the influence on the first computer 20 (such as the influence on traffic) can be suppressed. Therefore, a new arithmetic circuit can be introduced into the first accelerator 24 while ensuring the reliability of the first computer 20. This can eliminate the conventional problem of a decrease in reliability of the first computer 20 when a new arithmetic circuit is introduced. Note that, without performing the above test operation, it may be determined that the new arithmetic circuit has been normally written in the second region R2 when the arithmetic circuit is written in the second accelerator 34 without any abnormality.
[0030] In the above embodiment, as shown in FIG. 5, the arithmetic circuit written in the first region R1 of the first accelerator 24 and the arithmetic circuit written in the second region R2 of the second accelerator 34 are the same including the write position, but it is sufficient that the written part and the unwritten part in both regions R1 and R2 are the same position, and for example, a dummy circuit may be written in the second region R2. However, as shown in FIG. 5, it is better to make the arithmetic circuit written in the first region R1 of the first accelerator 24 and the arithmetic circuit written in the second region R2 of the second accelerator 34 the same including the write position. This allows the second computer 30 to perform a test operation of the already written arithmetic circuit and the newly written arithmetic circuit in parallel. Then, in the second computer 30, the newly written arithmetic circuit and other arithmetic circuits are test-operated in parallel, and the new arithmetic circuit may be written in the first accelerator 24 only when no abnormality is found in each test operation. This allows a new arithmetic circuit to be written in the first accelerator 24 without adversely affecting the existing arithmetic circuits.
[0031] If the reply in step S13 is that image processing is not possible (step S13; No), and the gateway 40 notifies the client computer C of this, the client computer C may supply the gateway 40 via the network N with a write instruction to newly generate and write a circuit configuration of an arithmetic circuit that performs the image processing. The instruction is supplied to the gateway 40 together with a program for generating a circuit configuration of the arithmetic circuit. This program may include a hardware description language that is the source of the circuit configuration. The gateway 40 that has received the write instruction executes the write process shown in FIG. 6 together with the second computer 30.
[0032] In the write process shown in FIG. 6, the gateway 40 first secures a write area in which a new arithmetic circuit is written in an unwritten portion of the second area R2 of the second accelerator 34 of the second computer 30 (step S51). This securing prevents another circuit from being written in the write area. The size of the write area is specified from the scale of the program (hardware description language, etc.) supplied together with the write instruction. Here, the portion indicated by the dotted line in FIG. 7 is assumed to be secured as the write area. It is located at the same position as the unwritten portion of the first area R1 of the first accelerator 24 of the first computer 20. In order to avoid performance degradation due to bandwidth congestion, the gateway 40 may secure an input / output terminal located at the same position as an unused input / output terminal of the first accelerator 24 as an input / output terminal of the secured write area. If the write area cannot be secured, the gateway 40 returns a message to that effect to the client computer C.
[0033] Thereafter, the gateway 40 supplies the supplied program together with the write instruction to the second computer 30, and the CPU 31 of the second computer 30 executes the supplied program to generate a circuit configuration for configuring an arithmetic circuit in the write area secured in step S51 (step S52). The generation of the circuit configuration appropriately includes processes such as logic synthesis of the hardware description language included in the program, placement and wiring, etc. Thereafter, the CPU 31 applies the generated circuit configuration to the currently secured write area, and writes the arithmetic circuit in that area (step S53; see FIG. 4).
[0034] Thereafter, the CPU 31 of the second computer 30 operates the arithmetic circuit written in step S53 to test whether the arithmetic circuit operates normally (step S54). The CPU 31 inputs a test pattern (test data) to the arithmetic circuit and operates it to test whether the arithmetic circuit operates normally. In the above test, when the probability of frame loss occurrence is higher than a predetermined standard, or when an operation different from the normal operation scenario is performed, for example, when the contents of a response to a specified request input to the arithmetic circuit are different, it is determined that the arithmetic circuit is not operating normally. It is preferable that the criteria for determining whether the arithmetic circuit is operating normally are determined in advance. This allows for efficient determination. The test pattern may be generated by the CPU 31 of the second computer 30 operating as a test pattern generating device that generates the test pattern, or may be obtained from a test pattern generating device connected to the second computer 30. The test pattern may be supplied to the second computer 30 via the gateway 40 together with a write instruction. It is preferable to use an FPGA for generating the test pattern. This makes it easy to generate a high-load test pattern.
[0035] When an abnormality is found in the operation of the arithmetic circuit, the CPU 31 determines whether or not the arithmetic circuit can be modified (step S55). If the arithmetic circuit can be modified (step S55; Yes), the CPU 31 modifies the circuit configuration, applies the modified circuit configuration to the second region R2, and writes the modified arithmetic circuit at the same position in the second region R2 (step S56). When the arithmetic circuit can be modified, the CPU 31 may transmit to the client computer C via the gateway 40 or the like a message indicating that the arithmetic circuit needs to be modified. The modification of the circuit configuration may be modification of an original program that generates the circuit configuration and generation of a circuit configuration based on the modified program. The modification of the circuit configuration may be modification of a hardware description language and logic synthesis and placement and wiring based on the modified hardware description language.
[0036] If it is difficult to correct the arithmetic circuit (step S55; No), the CPU 31 reserves another unwritten portion in the second area as a new writing area (step S57; see FIG. 7 for example), and performs the processes in and after step S52.
[0037] If there is no abnormality in the operation of the arithmetic circuit (step S54; Yes), the CPU 31 transfers the circuit configuration of this arithmetic circuit together with the write position of the arithmetic circuit to the first computer 20 via the gateway 40 (step S58). The first computer 20 applies the transferred circuit configuration to the first accelerator 24, and writes the arithmetic circuit that is normal to the write position.
[0038] Through the above series of processes, the arithmetic circuit that is originally to be written to the first accelerator 24 of the first computer 20 is written once to the second accelerator 34 of the second computer 30, and the operation of this arithmetic circuit is checked (here, tested using the above test pattern) before this arithmetic circuit is written to the first accelerator 24. Therefore, even if the processing in the first accelerator 24 of the first computer 20 or another processing in which a program is executed by the CPU 21 of the first computer 20 is already being executed when the writing processing of the arithmetic circuit starts, it is possible to suppress the writing of the arithmetic circuit from affecting the processing of the first accelerator 24 or the processing of the CPU 21 of the first computer 20, thereby realizing a highly reliable computing system 10. Furthermore, in the above, since logic synthesis and the like are executed on the second computer 30 side, it is not necessary for the first computer 20 to perform logic synthesis and the like, and the processing load on the first computer 20 is reduced.
[0039] Another modified example of the computing system 10 will be described with reference to Fig. 8. Note that Fig. 8 focuses on the accelerator, and the CPU and the like are omitted.
[0040] The first computer 20 includes a 1-1 computer 20A and a 1-2 computer 20B. The 1-1 computer 20A includes a plurality of accelerators 24-1 and 24-2. The 1-2 computer 20B includes a plurality of accelerators 24-3 and 24-4. The 1-1 computer 20A and the 1-2 computer 20B are connected via a network such as a LAN or the Internet (not shown), and the accelerators 24-1 to 24-4 collectively form one first accelerator 24.
[0041] The second computer 30 includes N (here, eight) accelerators 34-1 to 34-N. The accelerators 34-1 to 34-N are interconnected by buses A1 to An (n=N*(N-1) / 2). Furthermore, the accelerators 34-1 to 34-N are interconnected by buses B1 to Bn (n=N*(N-1) / 2) that are separated from the buses A1 to An. The buses A1 to An connect the accelerators 34-1 to 34-N so that the accelerators 34-1 to 34-N form a chain. The buses B1 to Bn connect the accelerators 34-1 to 34-N so that they form the start point (input) and the end point (output) of the chain. To realize these configurations, some of the buses A1 to An and B1 to Bn may be blocked so as not to be used as appropriate. By such mutual connections, the accelerators 34-1 to 34-N as a whole become one second accelerator 34, simulating a chain path described later.
[0042] The first accelerator 24 consisting of the accelerators 24-1 to 24-4 and the second accelerator 34 consisting of the accelerators 34-1 to 34-N can be regarded as having a reconfigurable first area and a reconfigurable second area, respectively, having the same overall circuit layout.
[0043] Next, the operation of the computing system 10 of this modified example will be described. Here, it is assumed that user A has written arithmetic circuits X1 and X2 into two accelerators 24-1 and 24-2 of the 1-1 computer 20A. The arithmetic circuit X1 of the accelerator 24-1 performs preprocessing on the image to be processed. The arithmetic circuit X2 infers the image contents based on the image preprocessed by the arithmetic circuit X1. The arithmetic circuits X1 and X2 are configured in a chain. In this state, it is assumed that only processing by the arithmetic circuits X1 and X2 is being executed, and the 1-1 computer 20A is in operation.
[0044] At this time, suppose that new image processing is requested from client computer C operated by user B to gateway 40. If the requested image processing can be performed by first computer 20, gateway 40 sets client computer C's communication destination to first computer 20, but since this image processing is new, it cannot be executed by first computer 20. At this time, gateway 40 switches client computer C's communication destination to second computer 30. At this time, client computer C of user B supplies second computer 30 with a program for generating a circuit configuration representing an arithmetic circuit for the image processing requested by user B.
[0045] Assume that the image processing requested by user B is a process in which image preprocessing is performed using the same chain processing as user A, followed by inference processing of the image. The CPU 31 of the second computer 30 determines that the processing amount of the preprocessing is smaller than the processing amount of inference from the contents of the program from the client computer C. From this estimate, for example, the write area Y1 in FIG. 8 in the unwritten portion of the first computer 20 is provisionally set as a write area in which an arithmetic circuit that performs image preprocessing is written, and the unwritten area Y2 is provisionally set as a write area in which an arithmetic circuit that performs inference processing after preprocessing is written. The written and unwritten portions of the first accelerator 24 and the second accelerator 34 are shared. Based on the result of the provisional setting, the CPU 31 writes the arithmetic circuit of user B into the corresponding areas of the unwritten areas Y1 and Y2 of the second accelerator 34. After writing, the normal operation (including long-term stable operation) of the arithmetic circuit is confirmed. After sufficient inspection, the arithmetic circuits are written into the write areas Y1 and Y2, and the second accelerator 34 can operate normally. In addition, since the arithmetic circuit of user B and the arithmetic circuit of user A are written to the second accelerator 34, for example, the system operator can transfer data actually used in the first computer 20 to the second computer 30 and input it to the second accelerator 34, thereby checking whether the writing of a new arithmetic circuit on the user B side affects the operation of the arithmetic circuit on the user A side that has already been written.
[0046] Here, buses A1 to An are buses that are configured as a whole with one optical transmission path and multiple optical filters. Similarly, buses B1 to Bn may also be buses that are configured as a single optical transmission path and multiple optical filters. Optical wavelength multiplexing communication is performed within the bus, using multiple different optical wavelengths. In this way, multiple accelerators may be connected to one transmission path so that each pair of accelerators that communicate with each other can communicate with light of different wavelengths. This configuration has the advantage of enabling low latency, since it is not necessary to assign a distribution ID, such as an electrical switch, to each bus in the chain between connections.
[0047] The hardware configuration of the computing system 10 is arbitrary. For example, at least two of the first computer 20, the second computer 30, and the gateway 40 may be realized by the same computer. For example, the CPU 21 of the first computer 20 may function as the gateway 40 and execute the processing of the gateway 40. The processing performed by the first computer 20 and the second computer 30 is not limited to image processing, and may be other processing. In addition to or instead of the second accelerator 34, the first accelerator 24 may also include multiple accelerators interconnected in the same manner as the second accelerator 34 in FIG. 8. In this way, the first accelerator 24 may simulate a chain path.
[0048] (Scope of the present invention) Although the present invention has been described above with reference to the embodiment and the modified examples, the present invention is not limited to the above embodiment and the modified examples. For example, the present invention includes various modifications to the above embodiment and the modified examples that can be understood by a person skilled in the art within the scope of the technical idea of the present invention. The configurations listed in the above embodiment and the modified examples can be appropriately combined within a range without contradiction. [Explanation of symbols]
[0049] 10...computing system, 20, 20A, 20B...first computer, 22...main memory, 24...first accelerator, 24-1 to 4...accelerators, 30...second computer, 34...second accelerator, 34-1 to N...accelerators, 40...gateway, A1 to An...buses, B1 to Bn...buses, C...client computer, N...network, R1...first area, R2...second area, X1, X2...arithmetic circuit, Y1, Y2...areas.
Claims
1. A first computer configured to write an arithmetic circuit into a reconfigurable first area provided in a first accelerator, a second computer configured to write an arithmetic circuit into a reconfigurable second area having the same circuit layout as the first area and provided in a second accelerator different from the first accelerator, and the second computer is configured to write a new second arithmetic circuit into a partial area at the same position as an unwritten partial area of the first area in the second area before the first computer writes a new first arithmetic circuit into the first area, the first computer is configured not to write the first arithmetic circuit into the first area when the second arithmetic circuit cannot be normally written into the partial area of the second area, and to write the first arithmetic circuit into the partial area of the first area when the second arithmetic circuit is normally written into the partial area of the second area, A computing system.
2. The second computer operates the second arithmetic circuit written into the partial area of the second area, when the second arithmetic circuit is normally written into the partial area of the second area, including when the second arithmetic circuit operates normally, The computing system according to claim 1.
3. The second computer further includes a generation unit that generates test data to be input to the second arithmetic circuit when operating the second arithmetic circuit, The computing system according to claim 2.
4. The second computer is configured to perform a test operation in parallel on the arithmetic circuit already written into the second area and the second arithmetic circuit when the same arithmetic circuit is already written at the same position in the first area and the second area, when the second arithmetic circuit is normally written into the partial area of the second area, including when the already written arithmetic circuit and the second arithmetic circuit perform a normal test operation, The computing system according to any one of claims 1 to 3.
5. The computing system further includes a third computer that controls writing of the arithmetic circuit into the first area by the first computer and writing of the arithmetic circuit into the second area by the second computer, the third computer, causes the second computer to write the second arithmetic circuit, When the second arithmetic circuit is successfully written, cause the first computer to write to the first arithmetic circuit. The computing system according to any one of claims 1 to 4.
6. At least one of the first accelerator and the second accelerator includes a plurality of interconnected FPGA accelerators that constitute the first region or the second region. The computing system according to any one of claims 1 to 5.
7. Multiple accelerators are connected to a single transmission line so that they can communicate with different wavelengths of light in each set of accelerators that communicate with each other. The computing system according to claim 6.
8. The first arithmetic circuit and the second arithmetic circuit are the same arithmetic circuit. The computing system according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processor and test method for data processor
JP2005115566A
Electronic apparatus
JP2017045318A
Information processing apparatus, information processing method, and program
JP2017120966A
Information processor and information processing method and program
JP2018165908A
Calculation system, and control method and program of calculation system
JP2018206195A