A data processing method, apparatus and medium
By configuring the FPGA accelerator card, automatic transmission and automatic calculation are realized, and the problem of high distributed computing delay of multiple FPGA accelerator cards is solved, which improves the computing efficiency.
Patent Information
- Application Number
- CN202111425760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-26
AI Technical Summary
In the FPGA cloud platform, when multiple FPGA accelerator cards perform distributed computing, the calculation delay is large and the calculation efficiency is low, mainly because the data transmission between multiple FPGA accelerator cards and the switching between calculation steps is completed by the host software.
By configuring each FPGA accelerator card participating in distributed computing, the automatic transmission of intermediate result data, the automatic calculation of the accelerator card corresponding to the intermediate calculation steps and the automatic return of the final result data, avoiding the host software participating in the distributed computing process.
The calculation delay during distributed computing of multiple FPGA accelerator cards is reduced, thereby improving computing efficiency.
Smart Images

Figure CN114138481B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of FPGA cloud platforms, and particularly to a data processing method, apparatus, and medium. Background Art
[0002] With the continuous enhancement of the processing capabilities of FPGAs (i.e., Field Programmable Gate Arrays), more and more data centers are starting to use FPGAs for acceleration to improve computing power and flexibility. To manage these increasingly numerous and diverse FPGA acceleration cards, FPGA cloud platforms have emerged to address the current difficulties in deploying, maintaining, and managing FPGA acceleration cards.
[0003] Currently, under the management of the cloud platform, due to the limited logical resources of a single FPGA acceleration card, when a complex computing task cannot be achieved by one FPGA acceleration card, the complex computing task needs to be divided into multiple computing steps, and each step is assigned to an FPGA acceleration card for computing. After multiple FPGA acceleration cards complete the computing in sequence, the final result is returned to the host. Among them, the data transmission between multiple FPGA acceleration cards and the switching between computing steps are all completed by the software running on the host. In this way, the distributed computing of multiple cards will have a large delay compared to single-card computing, and the computing efficiency is low. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a data processing method, apparatus, and medium, which can reduce the computing delay during distributed computing by multiple FPGA acceleration cards, thereby improving the computing efficiency. The specific solutions are as follows:
[0005] In a first aspect, this application discloses a data processing method, including:
[0006] When the first target FPGA acceleration card receives a computing start command sent by the target host connected to it, it computes the data to be processed to obtain intermediate result data;
[0007] The first target FPGA acceleration card sends the intermediate result data and the computing type information of the next computing step to the next FPGA acceleration card according to its own configuration information, so that the next FPGA acceleration card computes the intermediate result data to obtain new intermediate result data, and sends the new intermediate result data and the computing type information of the next computing step to the next FPGA acceleration card according to its own configuration information until the last participating second target FPGA acceleration card completes the computing to obtain the final result data;
[0008] Return the final result data to the first target FPGA acceleration card through the second target FPGA acceleration card;
[0009] Send the final result data to the target host through the first target FPGA acceleration card to complete the distributed computing for the data to be processed.
[0010] Optionally, before the first target FPGA acceleration card obtains the calculation start command sent by the target host connected to itself, it further includes:
[0011] Obtain the configuration information of all FPGA acceleration cards participating in the calculation through the target host, and configure the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card;
[0012] Communicate with other hosts through the target host, and send the configuration information corresponding to each of the other hosts to the other hosts respectively, so that the other hosts can configure the corresponding configuration information to the FPGA acceleration cards connected to themselves;
[0013] Among them, the configuration information of the non-second target FPGA acceleration cards in all the FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next FPGA acceleration card participating in the calculation, and the calculation type information of the next step of calculation. And the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the memory stored in the next FPGA acceleration card participating in the calculation; the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory and the physical address of the memory stored in the target host.
[0014] Optionally, the configuring the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card includes:
[0015] Configure the configuration information corresponding to the first target FPGA acceleration card to the internal register of the first target FPGA acceleration card;
[0016] The other hosts configuring the corresponding configuration information to the FPGA acceleration cards connected to themselves includes:
[0017] The other hosts configure the corresponding configuration information to the internal registers of the FPGA acceleration cards connected to themselves.
[0018] Optionally, calculating the data to be processed to obtain intermediate result data includes:
[0019] Invoke the kernel of the first target FPGA acceleration card to calculate the data to be processed, obtaining intermediate result data, so that the kernel writes the intermediate result data into the memory of the first target FPGA acceleration card.
[0020] Optionally, it further includes:
[0021] When the kernel writes data to the memory, detect whether the current write address is within the physical address range of the intermediate result data stored in its own memory according to the preset mapping relationship;
[0022] If so, trigger the step of sending the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card by the first target FPGA acceleration card according to its own configuration information.
[0023] Optionally, the step of sending the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card by the first target FPGA acceleration card, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, includes:
[0024] Convert the intermediate result data into data packets by the first target FPGA acceleration card, and add the calculation type information of the next calculation to the last data packet according to its own configuration information;
[0025] Send the data packets to the next FPGA acceleration card, so that when the next FPGA acceleration card receives the last data packet, generate a kernel call command according to the calculation type information in the last data packet, and use the kernel call command to call its own kernel to perform corresponding calculations on the intermediate result data to obtain new intermediate result data.
[0026] Optionally, the step of the second target FPGA acceleration card returning the final result data to the first target FPGA acceleration card includes:
[0027] Detect the interrupt signal sent to the PCIE by the second target FPGA acceleration card after the kernel calculation is completed;
[0028] When the interrupt signal is detected, send the final result data to the first target FPGA acceleration card.
[0029] Second aspect, the present application discloses a data processing device, which is applied to an FPGA cloud platform and includes multiple FPGA acceleration cards participating in distributed computing, and a host respectively connected to the multiple FPGA acceleration cards. Among the multiple FPGA acceleration cards, there are a first target FPGA acceleration card and a second target FPGA acceleration card. Among them,
[0030] The first target FPGA acceleration card is configured to, when receiving a calculation start command sent by a target host connected to itself, calculate the data to be processed to obtain intermediate result data; according to its own configuration information, send the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, and according to its own configuration information, sends the new intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card until the last participating second target FPGA acceleration card completes the calculation to obtain the final result data;
[0031] The second target FPGA acceleration card is configured to return the final result data to the first target FPGA acceleration card;
[0032] The first target FPGA acceleration card is configured to send the final result data to the target host to complete the distributed calculation for the data to be processed.
[0033] Optionally, the target host is further configured to obtain the configuration information of all FPGA acceleration cards participating in the calculation, and configure the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card; communicate with other hosts, and respectively send the configuration information corresponding to the other hosts to the other hosts, so that the other hosts configure the corresponding configuration information to the FPGA acceleration cards connected to themselves;
[0034] Among them, the configuration information of non-second target FPGA acceleration cards among all the FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next participating FPGA acceleration card, and the calculation type information of the next calculation. And the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the next participating FPGA acceleration card's memory; the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory and the physical address of the final result data stored in the target host's memory.
[0035] In a third aspect, an embodiment of the present application discloses a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the foregoing data processing method.
[0036] It can be seen that in the present application, when the first target FPGA acceleration card receives a calculation start command sent by the target host connected to itself, it calculates the data to be processed to obtain intermediate result data, and then the first target FPGA acceleration card sends the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card according to its own configuration information, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, and sends the new intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card according to its own configuration information until the last participating second target FPGA acceleration card completes the calculation to obtain the final result data. Then, the second target FPGA acceleration card returns the final result data to the first target FPGA acceleration card, and finally the first target FPGA acceleration card sends the final result data to the target host to complete the distributed calculation of the data to be processed. That is, in the present application, by configuring each FPGA acceleration card participating in the distributed calculation, automatic transmission of intermediate result data, automatic calculation of the acceleration card corresponding to the intermediate calculation step, and automatic return of the final result data are realized, avoiding the participation of host software in the distributed calculation process, being able to reduce the calculation latency when multiple FPGA acceleration cards perform distributed calculation, and thus improving the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a data processing method provided by the present application;
[0039] Figure 2 It is a schematic structural diagram of a specific FGPA cloud platform distributed calculation host and acceleration card provided by the present application;
[0040] Figure 3 It is a schematic structural diagram of a static area of an FPGA acceleration card provided by the present application;
[0041] Figure 4 It is a schematic structural diagram of a specific FPGA acceleration card provided by the present application;
[0042] Figure 5 Schematic diagram of a specific FPGA acceleration card structure provided for this application;
[0043] Figure 6 Implementation architecture diagram of a specific data processing solution provided for this application;
[0044] Figure 7 Schematic diagram of the structure of a data processing device provided for this application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0046] Currently, under the management of the cloud platform, due to the limited logical resources of a single FPGA acceleration card, when a complex computing task cannot be implemented by one FPGA acceleration card, the complex computing task needs to be divided into multiple computing steps, and each step is assigned to an FPGA acceleration card for computing. After multiple FPGA acceleration cards complete the computing in sequence, the final result is returned to the host. Among them, the data transmission between multiple FPGA acceleration cards and the switching between computing steps are all completed by the software running on the host. In this way, the distributed computing of multiple cards will have a large delay compared with single-card computing, and the computing efficiency is low. For this reason, the embodiments of the present application provide a data processing solution, which can reduce the computing delay when multiple FPGA acceleration cards perform distributed computing, thereby improving the computing efficiency.
[0047] See Figure 1 As shown, the embodiments of the present application disclose a data processing method, including:
[0048] Step S11: When the first target FPGA acceleration card obtains a computing start command sent by the target host connected to itself, it computes the data to be processed to obtain intermediate result data.
[0049] In a specific embodiment, before the first target FPGA acceleration card receives the calculation start command sent by the target host connected to it, the following steps are further included: obtaining, through the target host, the configuration information of all FPGA acceleration cards participating in the calculation, and configuring the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card; communicating with other hosts through the target host, and respectively sending the configuration information corresponding to each of the other hosts to the other hosts, so that the other hosts can configure the corresponding configuration information to the FPGA acceleration cards connected to themselves.
[0050] Among them, the configuration information of non-second target FPGA acceleration cards among all the FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next FPGA acceleration card participating in the calculation, and the calculation type information of the next step of calculation. And the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the memory storage of the next FPGA acceleration card participating in the calculation; the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory and the physical address of the memory storage in the target host.
[0051] Further, in a specific embodiment, the embodiments of the present application can configure the configuration information corresponding to the first target FPGA acceleration card to the internal register of the first target FPGA acceleration card; the other hosts configure the corresponding configuration information to the internal registers of the FPGA acceleration cards connected to themselves. Among them, the internal register is an internal register in the BSP (i.e., Board Support Package, board-level support package).
[0052] That is to say, before the calculation starts, the embodiments of the present application can configure each FPGA acceleration card participating in the distributed calculation through the target host connected to the first FPGA acceleration card participating in the distributed calculation. In a specific embodiment, the target host configures the configuration information corresponding to the first target FPGA acceleration card to the internal register of the first target FPGA acceleration card through the PCI-E (i.e., peripheral component interconnect express, a high-speed serial computer expansion bus standard) bus, communicates with other hosts through the network, and respectively sends the configuration information corresponding to each of the other hosts to the other hosts, so that the other hosts can configure the corresponding configuration information to the FPGA acceleration cards connected to themselves through the PCI-E bus. After the configuration is completed, the target host sends a start calculation command to the first target FPGA acceleration card.
[0053] Moreover, in the embodiments of the present application, the kernel of the first target FPGA acceleration card is called to calculate the data to be processed to obtain intermediate result data, so that the kernel writes the intermediate result data into the memory of the first target FPGA acceleration card. Similarly, the next FPGA acceleration card also calls its own kernel to calculate the data to be processed to obtain the corresponding intermediate result data.
[0054] Step S12: The first target FPGA acceleration card sends the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card according to its own configuration information, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, and sends the new intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card according to its own configuration information until the last participating second target FPGA acceleration card completes the calculation to obtain the final result data.
[0055] In a specific implementation manner, when the kernel of the present application embodiment writes data into the memory, it detects whether the current write address is within the physical address range of the intermediate result data stored in its own memory according to the preset mapping relationship; if so, it triggers the step of sending the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card by the first target FPGA acceleration card according to its own configuration information.
[0056] The first target FPGA acceleration card converts the intermediate result data into data packets, and adds the calculation type information of the next calculation to the last data packet of the intermediate result data according to its own configuration information; the data packets are sent to the next FPGA acceleration card, so that when the next FPGA acceleration card receives the last data packet, it generates a kernel call command according to the calculation type information in the last data packet, and uses the kernel call command to call its own kernel to perform corresponding calculations on the intermediate result data to obtain new intermediate result data.
[0057] In a specific embodiment, the kernel writes the intermediate result data into the memory of the first target FPGA acceleration card. The BSP in the first target FPGA acceleration card issues an RDMA (Remote Direct Memory Access) command to the MAC (Media Access Control) module of this card. The MAC module converts the intermediate result data in the local memory of the acceleration card into an RDMA data packet according to the configuration information and transmits it to the memory of the next acceleration card. In the last data packet for sending the intermediate result data, the data packet header carries the calculation type information for the next calculation. After receiving the last data packet of the intermediate result data, the next acceleration card calls the kernel for corresponding calculations according to the calculation type information. While the kernel of the next acceleration card calculates the intermediate result data, it automatically issues an RDMA write command to transmit it to the next acceleration card, and so on until the last acceleration card for the calculation. After the kernel calculation of the last acceleration card is completed, the calculation result is fed back to the target host memory according to the configuration information in the BSP.
[0058] Step S13: Return the final result data to the first target FPGA acceleration card through the second target FPGA acceleration card.
[0059] In a specific embodiment, the embodiment of the present application detects the interrupt signal sent to the PCIE by the kernel after the calculation is completed through the second target FPGA acceleration card; when the interrupt signal is detected, the final result data is sent to the first target FPGA acceleration card.
[0060] Step S14: Send the final result data to the target host through the first target FPGA acceleration card to complete the distributed calculation for the data to be processed.
[0061] That is, the second target FPGA acceleration card sends the final result data to the target host according to the configuration information, that is, the network address information of the first target FPGA acceleration card, as well as the physical address range of the final result data stored in its own memory and the physical address of the memory storage in the target host.
[0062] See Figure 2 shown Figure 2 It is a schematic structural diagram of a specific FGPA cloud platform distributed computing host and acceleration card provided by the embodiment of the present application. Under the management of the cloud platform, complex computing tasks are assigned to one or several FPGAs in the FPGA resource pool for acceleration. The acceleration cards in the resource pool are connected to the server through PCI-E, and the acceleration cards transmit data through Ethernet. Figure 2Taking three acceleration cards and three hosts as an example, it includes Host 1, FPGA Acceleration Card 1, Host 2, FPGA Acceleration Card 2, Host 3, and FPGA Acceleration Card 3. The inside of the FPGA acceleration card adopts a general architecture that supports OpenCL programming, and is divided into two parts: a static area (BSP) and a computing unit (kernel). See Figure 3 as shown Figure 3 It is a schematic diagram of the static area structure of an FPGA acceleration card provided by an embodiment of the present application. The static area includes modules such as a PCI-E module connected to the host CPU unit, a network data processing module (MAC) connected to the network, and a memory controller (DDR_controller). The host calls the kernel through the PCI-E to start the calculation and obtains the calculation completion information. The host can send and receive information with other hosts on the network through the PCI-E and the MAC module, or send an RDMA write command to the MAC through the PCI-E. The MAC module converts the local acceleration card memory data into an RDMA data packet and transmits it to the memory of other acceleration cards on the Ethernet. The kernel is a computing unit developed by the user and can be written in OpenCL (i.e., Open Computing Language, an open computing language) or developed in a traditional RTL (i.e., register transfer language) language. The kernel can read and write the FPGA acceleration card memory through the memory controller in the BSP.
[0063] It should be noted that in the prior art, complex calculation tasks are divided into two or more calculation steps, and each step is assigned to a single FPGA acceleration card for calculation. After multiple FPGA acceleration cards complete the calculation in sequence, the final result is returned to the host. Taking two calculation steps as an example, the first host sends an instruction through the PCI-E to make the first FPGA acceleration card start the calculation. After the kernel calculation is completed, an interrupt signal is sent to the first host through the PCI-E. After the first host obtains the calculation completion information of the first FPGA acceleration card, it sends an RDMA write command to the MAC through the PCI-E to transfer the intermediate result data in the memory of the first acceleration card to the memory of the second acceleration card. After the first host confirms that the data transfer is completed, it notifies the second host to perform the next calculation. The second host sends an instruction through the PCI-E to make the second acceleration card start the calculation. After the kernel calculation is completed, an interrupt signal is sent to the second host through the PCI-E, and the second host sends a message to notify the first host that the calculation is over. From the above distributed calculation process, it can be seen that both the data transfer between multiple cards and the switching between calculation steps are completed by the software running on the host, resulting in a large delay. The solution proposed in the present application can significantly reduce the delay of distributed calculation on the FPGA cloud platform without changing the computing unit (kernel).
[0064] See Figure 4As shown in the figure, the embodiment of the present application provides a schematic diagram of a specific FPGA acceleration card structure. The embodiment of the present application is implemented through the memory detection module and the command merging module in the BSP.
[0065] The memory detection module is located between the kernel and the memory controller and can transparently transmit the kernel's read and write memory operations. It internally contains a memory mapping table that records the mapping relationship between the physical address of the intermediate result data stored in the memory of this card and the physical address of the next acceleration card, as well as the information of the next calculation type and the network address information of the next acceleration card. When the kernel writes data to the acceleration card memory, the memory detection module compares the write address with the register setting of the physical address of the intermediate result data stored in the memory of this card. If the data write address is within the range of the physical address of the intermediate result data stored in the memory of this card, it is determined that the data written by the kernel is intermediate result data; the physical address of the next acceleration card memory and the acceleration card network address information are obtained by querying the memory mapping table. The memory detection module issues an RDMA write command to the MAC module, and the MAC module reads the intermediate result data from the memory of this card and forms an RDMA network data packet to be sent to the next acceleration card. When the memory detection module detects the last data of the intermediate result data written by the kernel, it issues an RDMA write command with the next calculation type to the MAC. The last intermediate result data packet sent by the MAC has the next calculation type information in the packet header.
[0066] The command merging module is located between the PCI-E bus and the kernel, and the PCI-E bus operations can be transparently transmitted to the kernel through the command merging module. The command merging module can parse the RDMA data packet received by the MAC module to obtain the information on whether the last packet of the intermediate result data has arrived and the next calculation type. When the last packet of the intermediate result data arrives, it converts the calculation type information contained therein into a PCI-E bus write register command for calling the kernel to start calculation and sends it to the kernel to make the kernel start calculation. The command merging module will detect the interrupt signal sent to the PCI-E after the kernel calculation is completed. When the command merging module belongs to the last acceleration card in the calculation process and is set with the physical address of the target host memory to store the calculation result and the network address information of the first target FPGA acceleration card, it converts the interrupt signal of the kernel calculation completion into an RDMA write command and sends it to the MAC module, and the MAC module sends the calculation result to the memory of the first host through the network.
[0067] In this way, without changing the design of the computing unit of the FPGA acceleration card, multi-step distributed computing is made independent of the scheduling of the host software, realizing the functions of automatically transmitting intermediate result data, automatically performing the next calculation, and automatically returning the result. Without increasing the development workload, the FPGA cloud platform can perform complex large-scale calculations distributively without significantly increasing the calculation latency.
[0068] Taking two-step distributed computing as an example, the data processing solution provided by this application is described below:
[0069] See Figure 5 as shown in Figure 5 FIG. is a schematic diagram of a specific FPGA acceleration card structure provided by an embodiment of this application. The FPGA acceleration card used is the Inspur f10a acceleration card. The FPGA of this acceleration card is an Intel Arria 10 device. There are two 10G Ethernet optical ports connected to the FPGA, and two 4GB SDRAMs as memories, which can be connected to the CPU of the server through PCI-E.
[0070] See Figure 6 as shown in Figure 6This is a specific implementation architecture diagram of the data processing solution provided by the embodiments of this application. Two steps of the calculation are respectively completed by two FPGA acceleration cards connected through a network. The two FPGA acceleration cards are respectively connected to the host through PCI-E. First, the first host sets the BSP register of the first FPGA acceleration card through PCI-E to determine the physical address range of the intermediate result data generated by the first step in the memory of this acceleration card, the network address of the second host, and the physical address range of the intermediate result data in the memory of the second host, as well as the calculation type information of the second step. The first host transmits the configuration information to the second host through the network. The second host configures the BSP register of the second FPGA acceleration card through PCI-E to determine the network address of the first FPGA acceleration card and the physical address of the final result data stored in the memory of this card and the first host. The first host calls the kernel of the first FPGA acceleration card through PCI-E to start the calculation. The kernel writes the calculation result into the memory of this card. The memory detection module in the BSP detects the operation of the kernel writing the memory of this card, and determines that the write address is within the set physical address range of the intermediate result data. The physical address of the intermediate result data in the second FPGA acceleration card is obtained by looking up the table, and an RDMA write command is sent to the MAC module. According to the RDMA write command, the MAC module of this card forms an RDMA network data packet with the intermediate result data in the memory of this card and sends it to the MAC module of the second FPGA acceleration card. The MAC module of the second FPGA acceleration card writes the intermediate result data in the RDMA data packet into the corresponding physical address in the memory of the second FPGA acceleration card. When the memory detection module in the BSP of the first FPGA acceleration card detects the last data of the intermediate result data written by the kernel, an RDMA write command with the next calculation type information is sent to the MAC module, and the MAC module sends the last intermediate result data packet with the next calculation type information. When the last packet of the intermediate result data arrives at the MAC of the second FPGA acceleration card, the command merging module detects that the last packet of the intermediate result arrives and obtains the next calculation type information, and converts this information into a PCI-E bus write register command and sends it to the kernel. The kernel of the second acceleration card starts to calculate. After the calculation is completed, the kernel issues an interrupt signal. The command merging module converts the kernel calculation completion interrupt signal into an RDMA write command and sends it to the MAC module. The MAC module converts the final result data into an RDMA data packet and sends it to the MAC module of the first FPGA acceleration card. The MAC module of the first FPGA acceleration card sends the final result data to the memory of the first host through PCI-E. The first host software polls the calculation result buffer area of the memory of the first host to obtain the final result data, and the distributed calculation is completed.
[0071] It can be seen that in the embodiments of the present application, by configuring each FPGA acceleration card participating in distributed computing, automatic transmission of intermediate result data, automatic calculation of the acceleration card corresponding to the intermediate calculation step, and automatic return of the final result data are achieved. The participation of the host software in the distributed computing process is avoided, and the calculation latency during distributed computing by multiple FPGA acceleration cards can be reduced, thereby improving the calculation efficiency.
[0072] See Figure 7 As shown, the embodiments of the present application provide a data processing device, which is applied to an FPGA cloud platform and includes multiple FPGA acceleration cards participating in distributed computing, and a host respectively connected to the multiple FPGA acceleration cards. Among the multiple FPGA acceleration cards, there are a first target FPGA acceleration card 11 and a second target FPGA acceleration card 12. Among them,
[0073] The first target FPGA acceleration card 11 is configured to, when receiving a calculation start command sent by a target host connected to itself, calculate the data to be processed to obtain intermediate result data; according to its own configuration information, send the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, and according to its own configuration information, sends the new intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card until the last participating second target FPGA acceleration card 12 completes the calculation to obtain the final result data;
[0074] The second target FPGA acceleration card 12 is configured to return the final result data to the first target FPGA acceleration card 11;
[0075] The first target FPGA acceleration card 11 is configured to send the final result data to the target host to complete the distributed calculation for the data to be processed.
[0076] It can be seen that in the embodiments of the present application, by configuring each FPGA acceleration card participating in distributed computing, automatic transmission of intermediate result data, automatic calculation of the acceleration card corresponding to the intermediate calculation step, and automatic return of the final result data are achieved. The participation of the host software in the distributed computing process is avoided, and the calculation latency during distributed computing by multiple FPGA acceleration cards can be reduced, thereby improving the calculation efficiency.
[0077] In a specific embodiment, the target host is further configured to obtain the configuration information of all FPGA acceleration cards participating in the calculation, and configure the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card; communicate with other hosts, and respectively send the configuration information corresponding to the other hosts to the other hosts, so that the other hosts configure the corresponding configuration information to the FPGA acceleration cards connected to themselves.
[0078] Among them, the configuration information of the non-second target FPGA acceleration cards in all the FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next FPGA acceleration card participating in the calculation, and the calculation type information of the next step of calculation. And the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the memory stored in the next FPGA acceleration card participating in the calculation; the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory and the physical address of the memory stored in the target host.
[0079] And, in a specific embodiment, the target host configures the configuration information corresponding to the first target FPGA acceleration card to the internal register of the first target FPGA acceleration card; the other hosts configure the corresponding configuration information to the internal register of the FPGA acceleration cards connected to themselves.
[0080] The first target FPGA acceleration card calls its own kernel to calculate the data to be processed, and obtains intermediate result data, so that the kernel writes the intermediate result data into the memory of the first target FPGA acceleration card.
[0081] Further, when the kernel writes data to the memory, the first target FPGA acceleration card detects whether the current write address is within the physical address range of the intermediate result data stored in its own memory according to the preset mapping relationship; if so, it triggers the step of sending the intermediate result data and the calculation type information of the next step of calculation to the next FPGA acceleration card by the first target FPGA acceleration card according to its own configuration information.
[0082] Moreover, the first target FPGA acceleration card converts the intermediate result data into data packets, and adds calculation type information for the next calculation in the last data packet of the intermediate result data according to its own configuration information; and sends the data packets to the next FPGA acceleration card, so that when the next FPGA acceleration card receives the last data packet, it generates a kernel call command according to the calculation type information in the last data packet, and uses the kernel call command to call its own kernel to perform corresponding calculations on the intermediate result data to obtain new intermediate result data.
[0083] The second target FPGA acceleration card detects the interrupt signal sent to the PCIE after the kernel calculation is completed; when the interrupt signal is detected, it sends the final result data to the first target FPGA acceleration card.
[0084] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, it implements the data processing method disclosed in the foregoing embodiment.
[0085] For the specific process of the above data processing method, reference may be made to the corresponding content disclosed in the foregoing embodiment, and details will not be repeated herein.
[0086] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments may be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference may be made to the description of the method part for related parts.
[0087] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0088] The above has introduced in detail a data processing method, device and medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A data processing method, characterized in that, Applied to the FPGA cloud platform, including: When the first target FPGA acceleration card receives a calculation start command sent by the target host connected to itself, it calculates the data to be processed to obtain intermediate result data; The first target FPGA acceleration card sends the intermediate result data and the calculation type information for the next calculation to the next FPGA acceleration card according to its own configuration information, so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, and sends the new intermediate result data and the calculation type information for the next calculation to the next FPGA acceleration card according to its own configuration information until the last participating second target FPGA acceleration card completes the calculation to obtain the final result data; among them, the configuration information of all non-second target FPGA acceleration cards among the participating FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next participating FPGA acceleration card, and the calculation type information for the next calculation, and the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the memory storage of the next participating FPGA acceleration card; The second target FPGA acceleration card returns the final result data to the first target FPGA acceleration card; The first target FPGA acceleration card sends the final result data to the target host to complete the distributed calculation for the data to be processed.
2. The data processing method according to claim 1, wherein Before the first target FPGA acceleration card receives a calculation start command sent by the target host connected to itself, it further includes: The target host obtains the configuration information of all FPGA acceleration cards participating in the calculation and configures the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card; The target host communicates with other hosts and sends the configuration information corresponding to each of the other hosts to the other hosts respectively, so that the other hosts configure the corresponding configuration information to the FPGA acceleration cards connected to themselves; Among them, the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory and the physical address of the memory storage in the target host.
3. The data processing method according to claim 2, wherein The configuring the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card includes: Configuring the configuration information corresponding to the first target FPGA acceleration card to the internal register of the first target FPGA acceleration card; The other hosts configuring the corresponding configuration information to the FPGA acceleration cards connected to themselves includes: The other hosts configure the corresponding configuration information to the internal registers of the FPGA acceleration cards connected to themselves.
4. The data processing method according to claim 2, wherein The calculating the data to be processed to obtain intermediate result data includes: Invoke the kernel of the first target FPGA acceleration card to calculate the data to be processed, obtaining intermediate result data, so that the kernel writes the intermediate result data into the memory of the first target FPGA acceleration card.
5. The data processing method according to claim 4, wherein It further includes: When the kernel writes data to the memory, detect whether the current write address is within the physical address range of the intermediate result data stored in its own memory according to the preset address mapping relationship; If so, trigger the step of sending the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card by the first target FPGA acceleration card according to its own configuration information.
6. The data processing method according to any one of claims 1 to 5, characterized in that, Sending the intermediate result data and the calculation type information of the next calculation to the next FPGA acceleration card by the first target FPGA acceleration card so that the next FPGA acceleration card calculates the intermediate result data to obtain new intermediate result data, includes: Convert the intermediate result data into data packets by the first target FPGA acceleration card, and add the calculation type information of the next calculation to the last data packet of the intermediate result data according to its own configuration information; Send the data packet to the next FPGA acceleration card, so that when the next FPGA acceleration card receives the last data packet, generate a kernel call command according to the calculation type information in the last data packet, and use the kernel call command to call its own kernel to perform corresponding calculations on the intermediate result data to obtain new intermediate result data.
7. The data processing method according to claim 6, wherein Sending the final result data back to the first target FPGA acceleration card by the second target FPGA acceleration card, includes: Detect the interrupt signal sent to the PCIE by the second target FPGA acceleration card after the kernel calculation is completed; When the interrupt signal is detected, send the final result data to the first target FPGA acceleration card.
8. A data processing device, characterized in that, Applied to the FPGA cloud platform, including multiple FPGA acceleration cards participating in distributed computing, and a host respectively connected to the multiple FPGA acceleration cards. Among the multiple FPGA acceleration cards, there are a first target FPGA acceleration card and a second target FPGA acceleration card, where The first target FPGA acceleration card is configured to, when receiving a calculation start command sent by a target host connected to itself, perform calculations on the data to be processed to obtain intermediate result data; and send the intermediate result data and the calculation type information for the next calculation to the next FPGA acceleration card according to its own configuration information, so that the next FPGA acceleration card performs calculations on the intermediate result data to obtain new intermediate result data, and sends the new intermediate result data and the calculation type information for the next calculation to the next FPGA acceleration card according to its own configuration information, until the last participating second target FPGA acceleration card completes the calculation to obtain the final result data; wherein, the configuration information of all non-second target FPGA acceleration cards among the participating FPGA acceleration cards includes a preset address mapping relationship, the network address information of the next participating FPGA acceleration card, and the calculation type information for the next calculation, and the preset address mapping relationship is the mapping relationship between the physical address range of the intermediate result data stored in its own memory and the physical address range of the memory of the next participating FPGA acceleration card; The second target FPGA acceleration card is configured to return the final result data to the first target FPGA acceleration card; The first target FPGA acceleration card is configured to send the final result data to the target host to complete the distributed calculation for the data to be processed.
9. The data processing device according to claim 8, wherein the target host is further configured to obtain the configuration information of all participating FPGA acceleration cards, and configure the configuration information corresponding to the first target FPGA acceleration card to the first target FPGA acceleration card; communicate with other hosts, and send the configuration information corresponding to each of the other hosts to the other hosts respectively, so that the other hosts configure the corresponding configuration information to the FPGA acceleration cards connected to themselves; wherein, the configuration information of the second target FPGA acceleration card includes the network address information of the first target FPGA acceleration card, the physical address range of the final result data stored in its own memory, and the physical address of the memory in the target host.
10. A computer-readable storage medium, characterized in that, A computer program storage medium for storing a computer program, which when executed by a processor, implements the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
CPU+FPGA-based heterogeneous computing system and acceleration method thereof
CN108776649A
Distributed computing implementation method and system based on FPGA
CN109976912A
FPGA network for stream computing and stream computing system and method
CN110069441A
Data processing method and device, distributed data flow programming framework and related components
CN111324558A
Task deployment method and device based on multi-board FPGA heterogeneous system
CN111736966A