An accelerated computing system
By introducing the control end of the onboard accelerator and FPGA cluster in the acceleration computing system, data is processed and transmitted directly within the FPGA cluster, the problems of low data transmission efficiency and insufficient security in the existing system are solved, and a more efficient, secure and general acceleration computing system is achieved.
Patent Information
- Application Number
- CN202111150949.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-09-29
AI Technical Summary
In the existing acceleration system, it is necessary to design FPGA equipment factory classes, drivers, etc. in the server to start the FPGA cluster, and transmit the pending data to the FPGA cluster through the on-board FPGA on the server, resulting in low data transmission efficiency and insufficient security.
It provides an accelerated computing system, including a server and an FPGA cluster, and the server is equipped with an on-board accelerator, a TensorFlow architecture end and a control end of the FPGA cluster. The TensorFlow architecture side obtains the pending data, the control side determines the list of cluster accelerators for processing data in the FPGA cluster, and starts the cluster accelerator through the onboard accelerator, and directly performs data processing and transmission within the FPGA cluster.
The efficiency and security of the machine learning model are improved, and the system is improved by reducing data transmission paths, saving transmission time, enhancing data security, and supporting finer-grained operator-level data processing, improving the system's universality.
Smart Images

Figure CN113988307B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an accelerated computing system. Background Art
[0002] Currently, FPGA (Field Programmable Gate Array) clusters can be used to accelerate machine learning models. However, in existing acceleration systems, it is necessary to design FPGA device factory classes and drivers for the FPGA cluster in the server before the computing task can be assigned to the FPGA cluster. In addition, the relevant data to be processed involved in the machine learning model needs to be sent to the FPGA cluster through the onboard FPGA on the server, resulting in slow data transmission efficiency and data security cannot be guaranteed. Among them, the FPGA device factory class and driver for the FPGA cluster are used to implement operations such as starting the FPGA cluster.
[0003] Therefore, how to improve the efficiency and security of the acceleration system of the machine learning model is a problem that technicians in this field need to solve. Summary of the invention
[0004] In view of this, the purpose of this application is to provide an accelerated computing system to improve the efficiency and security of the acceleration system of the machine learning model. The specific scheme is as follows:
[0005] In a first aspect, the present application provides an accelerated computing system, comprising: a server and an FPGA cluster, wherein the server is provided with an onboard accelerator, a TensorFlow architecture end, and a control end of the FPGA cluster; the FPGA cluster is provided with a plurality of cluster accelerators, wherein:
[0006] The TensorFlow architecture end is used to obtain the data to be processed involved in the target operator in any machine learning model;
[0007] The control end is used to obtain the data to be processed from the TensorFlow architecture end, determine a list of cluster accelerators in the FPGA cluster for processing the data to be processed, and transmit the list of cluster accelerators to the TensorFlow architecture end, so that the TensorFlow architecture end transmits the list of cluster accelerators to the onboard accelerator;
[0008] The onboard accelerator is used to start the cluster accelerator in the cluster accelerator list, so that the cluster accelerator in the cluster accelerator list obtains the to-be-processed data from the control end for processing, and returns the processing result obtained by the processing to the onboard accelerator;
[0009] The onboard accelerator is also used to transmit the processing results to the TensorFlow architecture end.
[0010] Preferably, the TensorFlow architecture end is specifically used to: store the data to be processed in its own memory, and open the access rights of its own memory to the control end;
[0011] Accordingly,
[0012] The control end is specifically used to copy the data to be processed from the memory of the TensorFlow architecture end to its own memory.
[0013] Preferably, the cluster accelerators in the cluster accelerator list read the to-be-processed data from the memory of the control end for processing.
[0014] Preferably, the cluster accelerators in the cluster accelerator list read the to-be-processed data from the memory of the control end through the RDMA protocol for processing.
[0015] Preferably, the control end is specifically used to: divide the data to be processed into multiple parts, select a corresponding cluster accelerator for each part of the data to be processed, record accelerator information of the selected cluster accelerator, and obtain the cluster accelerator list.
[0016] Preferably, the control end is specifically used to: take the cluster accelerators in the FPGA cluster that are in an idle state as candidates, and select a corresponding cluster accelerator for each part of the data to be processed from all candidates based on preset conditions.
[0017] Preferably, the preset conditions include: computing power and / or transmission delay of the cluster accelerator to be selected.
[0018] Preferably, the TensorFlow architecture end is further used to: obtain accelerator information of the onboard accelerator;
[0019] Accordingly,
[0020] The control end is further used to obtain accelerator information of the onboard accelerator from the TensorFlow architecture end, and transmit the accelerator information of the onboard accelerator to the cluster accelerators in the cluster accelerator list, so that the cluster accelerators in the cluster accelerator list open access rights to the onboard accelerator.
[0021] Preferably, any cluster accelerator includes:
[0022] A computing sub-kernel, used for processing part of the to-be-processed data acquired by the cluster accelerator;
[0023] The logic control sub-core is used to open an access port of the kernel protocol to the onboard accelerator to open access rights to the onboard accelerator.
[0024] Preferably, the onboard accelerator comprises:
[0025] A data interaction sub-kernel, configured to receive the cluster accelerator list via a bus between the TensorFlow architecture end and receive the processing result via a kernel protocol between the FPGA cluster;
[0026] The computing control sub-kernel is used to start the cluster accelerator in the cluster accelerator list through a kernel protocol with the FPGA cluster.
[0027] It can be seen from the above scheme that the present application provides an accelerated computing system, including: a server and an FPGA cluster, wherein the server is provided with an airborne accelerator, a TensorFlow architecture end and a control end of the FPGA cluster; the FPGA cluster is provided with multiple cluster accelerators, wherein: the TensorFlow architecture end is used to obtain the data to be processed involving the target operator in any machine learning model; the control end is used to obtain the data to be processed from the TensorFlow architecture end, and determine a list of cluster accelerators for processing the data to be processed in the FPGA cluster, and transmit the list of cluster accelerators to the TensorFlow architecture end so that the TensorFlow architecture end transmits the list of cluster accelerators to the airborne accelerator; the airborne accelerator is used to start the cluster accelerator in the cluster accelerator list so that the cluster accelerator in the cluster accelerator list obtains the data to be processed from the control end for processing, and returns the processing result obtained by the processing to the airborne accelerator; the airborne accelerator is also used to transmit the processing result to the TensorFlow architecture end.
[0028] It can be seen that this application does not need to design FPGA device factory classes, drivers, etc. for the FPGA cluster in the server, because the control end of the airborne accelerator and the FPGA cluster can realize the control operations such as starting the FPGA cluster, and the network acceleration function of the airborne accelerator can also improve the control efficiency. At the same time, the relevant data to be processed involved in the machine learning model is not sent to the FPGA cluster through the airborne FPGA on the server, but the control end of the FPGA cluster directly transmits the relevant data to the FPGA cluster, thereby saving transmission time, improving data transmission efficiency, and ensuring data security. Because the FPGA cluster has comprehensive data security protection measures for data, not using the airborne FPGA to forward data can avoid data from being transmitted out of the FPGA cluster, that is: the control end of the FPGA cluster directly transmits the relevant data to the FPGA cluster, so that the data can be transmitted within the FPGA cluster. Furthermore, the FPGA cluster provided in this application supports data processing at the operator level, and can process finer-grained operations in the machine learning model separately, so it has better versatility. Because the operator principles of different machine learning models are essentially the same, the FPGA cluster processes data for finer-grained operators, which is equivalent to the FPGA cluster being able to perform operator-level data processing for various models.
[0029] In summary, the accelerated computing system provided by this application has good data transmission efficiency, security and versatility. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0031] Figure 1 A schematic diagram of an accelerated computing system disclosed in this application;
[0032] Figure 2 A schematic diagram of another accelerated computing system disclosed in this application;
[0033] Figure 3 for Figure 2 Schematic diagram of data processing in the system shown. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0035] To facilitate the introduction of this application, the background technology involved in this application is first introduced as follows.
[0036] Currently, we can use an onboard FPGA accelerator installed on the server to accelerate the processing of machine learning models based on the TensorFlow architecture. However, when the amount of data is too large, such solutions are limited by the number and computing power of onboard FPGA accelerators, resulting in a long computing time. In addition, existing onboard FPGA accelerators do not support operator-level data processing, which limits the universality of the acceleration system.
[0037] Currently, FPGA clusters can be used to accelerate machine learning models. However, in existing acceleration systems, it is necessary to design FPGA device factory classes and drivers for the FPGA cluster in the server before the computing task can be handed over to the FPGA cluster. In addition, the relevant data to be processed involved in the machine learning model needs to be sent to the FPGA cluster through the onboard FPGA on the server, resulting in slow data transmission efficiency and data security. Among them, the FPGA device factory class and driver for the FPGA cluster are used to implement operations such as starting the FPGA cluster.
[0038] To this end, the present application provides an accelerated computing system that can improve the efficiency, security, and versatility of the acceleration system of the machine learning model.
[0039] See also Figure 1 As shown, an embodiment of the present application discloses an accelerated computing system, including: a server and an FPGA cluster, wherein the server is provided with an onboard accelerator, a TensorFlow architecture end, and a control end of the FPGA cluster; and the FPGA cluster is provided with multiple cluster accelerators.
[0040] The TensorFlow architecture end is used to obtain the data to be processed involving the target operator in any machine learning model. The target operator in the machine learning model can be: convolution operation, pooling operation, matrix multiplication, matrix addition, etc. Convolution operations include two-dimensional convolution, etc. The TensorFlow architecture end can be regarded as a software module, which provides the execution logic of the target operator, related data, and the driver of the onboard accelerator.
[0041] TensorFlow is a symbolic mathematics system based on data flow programming, which is widely used in the programming implementation of various machine learning algorithms.
[0042] The control end is used to obtain the data to be processed from the TensorFlow architecture end, determine the list of cluster accelerators used to process the data to be processed in the FPGA cluster, and transmit the list of cluster accelerators to the TensorFlow architecture end, so that the TensorFlow architecture end transmits the list of cluster accelerators to the onboard accelerator. The control end can be regarded as a software application for controlling the FPGA cluster, which provides data such as the control program of the FPGA cluster and the accelerator information of each cluster accelerator in the FPGA cluster. Accelerator information includes IP (Internet Protocol) address, MAC (Media Access Control) address, etc.
[0043] The airborne accelerator is used to start the cluster accelerator in the cluster accelerator list, so that the cluster accelerator in the cluster accelerator list obtains the data to be processed from the control end for processing, and returns the processing result obtained by the processing to the airborne accelerator. The airborne accelerator can be plugged into the server. Before the airborne accelerator starts the cluster accelerator in the cluster accelerator list, the cluster accelerator in the cluster accelerator list already knows the accelerator information of the airborne accelerator, and accordingly opens the access port for the airborne accelerator.
[0044] The onboard accelerator is also used to transmit the processing results to the TensorFlow architecture.
[0045] In this embodiment, multiple cluster accelerators are connected to the network to form an FPGA cluster by means of the optical ports and optical modules of the cluster accelerators, so as to give full play to the computing resources of each cluster accelerator.
[0046] In this embodiment, the control end of the airborne accelerator and the FPGA cluster is used to realize the control operations such as starting the FPGA cluster, and the network acceleration function of the airborne accelerator can also improve the control efficiency. At the same time, the relevant data to be processed involved in the machine learning model allows the control end of the FPGA cluster to directly transmit the relevant data to the FPGA cluster, thereby saving transmission time, improving data transmission efficiency, and ensuring data security. Because the FPGA cluster has comprehensive data security protection measures, not using the airborne FPGA to forward data can avoid data from being transmitted out of the FPGA cluster, that is: the control end of the FPGA cluster directly transmits the relevant data to the FPGA cluster, which can enable data to be transmitted within the FPGA cluster.
[0047] Furthermore, the FPGA cluster in this embodiment can process fine-grained operations such as convolution operations, pooling operations, matrix multiplication, and matrix addition in the machine learning model separately, so that the FPGA cluster supports operator-level data processing, and thus has better versatility. Because the operator principles of different machine learning models are essentially the same, the FPGA cluster processes data for finer-grained operators, which is equivalent to the FPGA cluster being able to perform operator-level data processing for various models.
[0048] It can be seen that the accelerated computing system provided by this embodiment has good data transmission efficiency, security and versatility.
[0049] Based on the above embodiments, it should be noted that, in a specific implementation, the TensorFlow architecture end is specifically used to: store the data to be processed in its own memory, and open the access rights of its own memory to the control end; accordingly, the control end is specifically used to: copy the data to be processed from the memory of the TensorFlow architecture end to its own memory.
[0050] The cluster accelerators in the cluster accelerator list read the data to be processed from the memory of the control end for processing.
[0051] Among them, the cluster accelerators in the cluster accelerator list read the data to be processed from the memory of the control end through the RDMA (Remote Direct Memory Access) protocol for processing. RDMA (Remote Direct Data Access Protocol, RDMA protocol can solve the delay of server-side data processing in network transmission.
[0052] It can be seen that the TensorFlow architecture side has its own memory, and the control side also has its own memory. The control side can access the memory of the TensorFlow architecture side to obtain the data to be processed provided by the TensorFlow architecture side, the accelerator information of the onboard accelerator (such as IP address, MAC address), and other data. The control side communicates directly with the FPGA cluster through the RDMA protocol, and each cluster accelerator in the FPGA cluster can directly access the memory of the control side to obtain the data to be processed. Among them, the cluster accelerator and the onboard accelerator are both FPGA boards.
[0053] Based on the above embodiment, it should be noted that, in a specific implementation, the control end is specifically used to: divide the data to be processed into multiple parts, select a corresponding cluster accelerator for each part of the data to be processed, record the accelerator information of the selected cluster accelerator, and obtain a cluster accelerator list. Each cluster accelerator in the cluster accelerator list can process each part of the data to be processed in parallel, thereby improving data processing efficiency.
[0054] The control end is specifically used to: select the cluster accelerators in the FPGA cluster that are in an idle state as candidates, and select the corresponding cluster accelerator for each part of the data to be processed from all candidates based on preset conditions. The preset conditions include: the computing power and / or transmission delay of the cluster accelerator as the candidate. For example, you can select the first few idle cluster accelerators with large computing power, or you can select the first few idle cluster accelerators with small transmission delay, or you can select based on the two factors of computing power and transmission delay.
[0055] Based on the above embodiment, it should be noted that the TensorFlow architecture end is also used to obtain the accelerator information of the onboard accelerator; accordingly, the control end is also used to obtain the accelerator information of the onboard accelerator from the TensorFlow architecture end, and transmit the accelerator information of the onboard accelerator to the cluster accelerator in the cluster accelerator list, so that the cluster accelerator in the cluster accelerator list opens access rights to the onboard accelerator. Among them, each cluster accelerator in the FPGA cluster communicates directly with the onboard accelerator through the kernel protocol, without having to forward through the server port, thereby improving efficiency. Kernel protocols such as IKL (Inter-Kernel Links, high-speed transmission protocol between kernels) and the like.
[0056] It should be noted that based on the concept of FPGA virtualization, the core corresponding to the logic resources of the FPGA can be divided into multiple sub-cores, so that the logic resources of the FPGA accelerator core can be fully utilized. Therefore, the core of a cluster accelerator can be divided into a computing sub-core and a logic control sub-core, and the core of an airborne accelerator can be divided into a data interaction sub-core and a computing control sub-core, and different functions can be assigned to different sub-cores.
[0057] In a specific implementation, any cluster accelerator includes: a computing sub-core for processing part of the to-be-processed data acquired by the cluster accelerator; and a logic control sub-core for opening an access port of the kernel protocol to the onboard accelerator to open access rights to the onboard accelerator.
[0058] In a specific implementation, the onboard accelerator includes: a data interaction sub-kernel, used to receive a cluster accelerator list through a bus between the TensorFlow architecture end, and receive processing results through a kernel protocol between the FPGA cluster; a computing control sub-kernel, used to start the cluster accelerator in the cluster accelerator list through a kernel protocol between the FPGA cluster.
[0059] join Figure 2 As shown, this embodiment takes the two-dimensional convolution operator as an example to introduce the present application as follows. Figure 2In the computing system shown, the TensorFlow architecture end, the control end of the FPGA cluster (i.e., the FPGA cluster host end), the onboard FPGA accelerator, and the FPGA cluster are organically combined.
[0060] Among them, the TensorFlow architecture can assign the computing task of two-dimensional convolution to each cluster accelerator in the FPGA cluster, so that these cluster accelerators can complete the two-dimensional convolution in parallel.
[0061] It should be noted that the TensorFlow architecture provided in this embodiment supports three additional functions:
[0062] 1. The memory read and write permissions for the data required to perform two-dimensional convolution are opened to the control end of the FPGA cluster, so that the cluster accelerator can read the input data directly from the control end of the FPGA cluster through the RDMA network protocol, thereby avoiding secondary forwarding through the onboard FPGA accelerator and saving time.
[0063] 2. Be able to obtain information such as the IP addresses of each cluster accelerator selected by the control end of the FPGA cluster for completing two-dimensional convolution;
[0064] 3. The IP addresses and other information of each cluster accelerator used to complete the two-dimensional convolution can be transmitted to the airborne FPGA accelerator through the driver of the airborne FPGA accelerator, so as to start the calculation through the airborne FPGA accelerator and collect the calculation results, making full use of the network acceleration function of the airborne FPGA accelerator.
[0065] These three functions ensure the information interaction between the TensorFlow architecture and the control end of the FPGA cluster, the onboard FPGA accelerator, and the cluster accelerator.
[0066] On the other hand, based on the concept of FPGA virtualization, this embodiment implements the data interaction sub-kernel and the computing control sub-kernel based on the optical module in the airborne FPGA accelerator. The two-dimensional convolution computing sub-kernel and the logic control sub-kernel are implemented in each cluster accelerator used to complete the two-dimensional convolution.
[0067] For the onboard FPGA accelerator, the data interaction sub-kernel is responsible for collecting the list of selected cluster accelerators sent by the TensorFlow architecture end, collecting the calculation results fed back by each cluster accelerator in the cluster accelerator list, and feeding back the calculation results to the TensorFlow architecture end. The calculation control sub-kernel is responsible for remotely starting each cluster accelerator in the list according to the cluster accelerator list sent by the TensorFlow architecture end, so that these cluster accelerators can perform two-dimensional convolution calculations.
[0068] For any cluster accelerator used to complete the two-dimensional convolution, the two-dimensional convolution calculation sub-core is responsible for executing the relevant logic of the two-dimensional convolution calculation. The logic control sub-core is responsible for configuring the information of the onboard FPGA accelerator to open its own access port to the onboard FPGA accelerator so as to interact with the onboard FPGA accelerator. Among them, any cluster accelerator obtains the information of the onboard FPGA accelerator from the control end of the FPGA cluster.
[0069] based on Figure 2 The computing system shown in the figure, the corresponding data processing flow can be found in Figure 3 . Figure 3 The process shown specifically includes:
[0070] 1. The TensorFlow architecture first opens the memory copy permission so that the control end of the FPGA cluster can copy the input data and onboard accelerator information of the two-dimensional convolution calculation.
[0071] 2. The control end of the FPGA cluster copies the input data and onboard accelerator information to its own backup memory and evaluates the size of the input data.
[0072] 3. The control end of the FPGA cluster selects cluster accelerators that can perform two-dimensional convolution calculations based on factors such as whether the cluster accelerators in the FPGA cluster are idle, the computing power, and the transmission delay.
[0073] 4. After the screening is completed, the input data is divided. The specific method is to evenly divide the input data at the address of the backup memory, and transmit the memory address of each part of the data to each selected cluster accelerator through the RDMA protocol, so that each cluster accelerator can read the data it needs to process based on the memory address. At the same time, the airborne accelerator information is sent to each selected cluster accelerator, so that these cluster accelerators can open the computing startup permission to the airborne FPGA accelerator.
[0074] Among them, the control end of the FPGA cluster transmits input data to each selected cluster accelerator through the RDMA protocol, which can save time and make full use of the security protection measures of the existing FPGA cluster to prevent security issues such as memory leakage.
[0075] 5. After each cluster accelerator completes recording the memory address and opens access rights to the onboard FPGA accelerator, the control end of the FPGA cluster informs the TensorFlow architecture end of the IP address and other information of the selected cluster accelerator, and the TensorFlow architecture end transmits this information to the onboard FPGA accelerator.
[0076] 6. The onboard FPGA accelerator starts the cluster accelerators through the optical module based on the IP addresses and other information obtained to perform the calculation of the two-dimensional convolution operator.
[0077] 7. Each selected cluster accelerator performs the calculation of the two-dimensional convolution operator. In this process, the data to be processed is obtained from the memory of the control end through the RDMA protocol. After the calculation is completed, the corresponding calculation results are fed back to the onboard FPGA accelerator.
[0078] 8. The onboard FPGA accelerator counts the calculation results of all the selected cluster accelerators and feeds them back to the TensorFlow architecture end;
[0079] 9. The TensorFlow architecture combines the above calculation results to obtain the calculation results of the two-dimensional convolution operator, outputs them and prints them.
[0080] It can be seen that this embodiment enables the TensorFlow architecture end, the onboard FPGA accelerator, the FPGA cluster, and the control end of the FPGA cluster to cooperate with each other to ensure the data transmission rate and security. While utilizing the network acceleration capability of the onboard FPGA accelerator, the parallel computing advantages of the FPGA cluster are fully utilized, and there is no need to implement the device factory class and corresponding scheduling management functions for the FPGA cluster on the TensorFlow architecture end.
[0081] The "first", "second", "third", "fourth", etc. (if any) referred to in this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods or devices.
[0082] It should be noted that the descriptions involving "first", "second", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0083] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0084] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of readable storage medium known in the art.
[0085] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An accelerated computing system, It is characterized in that include: A server and an FPGA cluster, wherein the server is provided with an onboard accelerator, a TensorFlow architecture end, and a control end of the FPGA cluster; The FPGA cluster is provided with a plurality of cluster accelerators, wherein: The TensorFlow architecture end is used to obtain the data to be processed involved in the target operator in any machine learning model; The control end is used to obtain the data to be processed from the TensorFlow architecture end, determine a list of cluster accelerators in the FPGA cluster for processing the data to be processed, and transmit the list of cluster accelerators to the TensorFlow architecture end, so that the TensorFlow architecture end transmits the list of cluster accelerators to the onboard accelerator; The onboard accelerator is used to start the cluster accelerator in the cluster accelerator list, so that the cluster accelerator in the cluster accelerator list obtains the to-be-processed data from the control end for processing, and returns the processing result obtained by the processing to the onboard accelerator; The onboard accelerator is also used to transmit the processing result to the TensorFlow architecture end; Wherein, any cluster accelerator includes: A computing sub-kernel, used for processing part of the to-be-processed data acquired by the cluster accelerator; A logic control sub-core, used to open an access port of a kernel protocol to the airborne accelerator, so as to open access rights to the airborne accelerator; Wherein, the airborne accelerator comprises: A data interaction sub-kernel, configured to receive the cluster accelerator list via a bus between the TensorFlow architecture end and receive the processing result via a kernel protocol between the FPGA cluster; The computing control sub-kernel is used to start the cluster accelerator in the cluster accelerator list through a kernel protocol with the FPGA cluster.
2. The accelerated computing system according to claim 1, It is characterized in that The TensorFlow architecture end is specifically used to: store the data to be processed in its own memory, and open the access rights of its own memory to the control end; Accordingly, The control end is specifically used to copy the data to be processed from the memory of the TensorFlow architecture end to its own memory.
3. The accelerated computing system according to claim 2, It is characterized in that The cluster accelerators in the cluster accelerator list read the data to be processed from the memory of the control end for processing.
4. The accelerated computing system according to claim 3, It is characterized in that The cluster accelerators in the cluster accelerator list read the data to be processed from the memory of the control end through the RDMA protocol for processing.
5. The accelerated computing system according to claim 1, It is characterized in that The control end is specifically used to: divide the data to be processed into multiple parts, select a corresponding cluster accelerator for each part of the data to be processed, record accelerator information of the selected cluster accelerator, and obtain the cluster accelerator list.
6. The accelerated computing system according to claim 5, It is characterized in that The control end is specifically used to: take the cluster accelerators in the FPGA cluster that are in an idle state as candidates, and select a corresponding cluster accelerator for each part of the to-be-processed data from all the candidates based on preset conditions.
7. The accelerated computing system according to claim 6, It is characterized in that The preset conditions include: the computing power and / or transmission delay of the cluster accelerator to be selected.
8. The accelerated computing system according to claim 1, It is characterized in that The TensorFlow architecture end is also used to: obtain accelerator information of the onboard accelerator; Accordingly, The control end is further used to obtain accelerator information of the onboard accelerator from the TensorFlow architecture end, and transmit the accelerator information of the onboard accelerator to the cluster accelerators in the cluster accelerator list, so that the cluster accelerators in the cluster accelerator list open access rights to the onboard accelerator.
Citation Information
Patent Citations
Heterogeneous unmanned aerial vehicle cluster target tracking system and method based on biosocial force
CN109613931A
Task acceleration processing method, device, equipment and readable storage medium
CN110399234A