A heterogeneous acceleration system, method, computing device and storage medium

By accumulating interrupt subrequests on heterogeneous accelerator cards and sending target interrupt requests when the threshold is reached, the problem of frequent interrupt requests in heterogeneous accelerator cards resulting in degradation of CPU performance is solved, and the overall performance of computing devices is improved.

CN119376952BActive Publication Date: 2025-05-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411910564.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Heterogeneous acceleration cards send a large number of interrupt requests to the CPU in a short time, causing the CPU to frequently respond to interrupt requests, reducing the overall performance of the computing device.

Method used

By accumulating interrupt sub-requests on a heterogeneous acceleration card, when the accumulation reaches the preset threshold, a target interrupt request is generated and sent to the host side to avoid frequent interrupt requests.

Benefits of technology

It effectively avoids heterogeneous acceleration cards sending a large number of interrupt requests to the host side in a short time, reducing the interrupt response frequency of the host side, thereby improving the overall performance of the computing device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119376952B_ABST
    Figure CN119376952B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and discloses a heterogeneous acceleration system, method, computing device and storage medium, the system comprising: a host end is used to obtain any acceleration task of a computing device, split the acceleration task into multiple subtasks, and unload the multiple subtasks to a heterogeneous acceleration card; the heterogeneous acceleration card is used to receive and execute multiple subtasks, and when any subtask is completed, a corresponding interrupt subrequest is generated, and when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, a target interrupt request is generated, and the target interrupt request is sent to the host end; the host end is used to receive the target interrupt request, and in response to the target interrupt request, reads the execution result of the subtask from the heterogeneous acceleration card. It is avoided that the heterogeneous acceleration card sends a large number of interrupt subrequests to the host end in a short period of time, and then the host end is avoided from frequently responding to interrupt requests, thereby improving the overall performance of the computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a heterogeneous acceleration system, method, computing device and storage medium. Background Art

[0002] With the development of technologies such as big data and artificial intelligence, users have higher and higher requirements for the computing power of computing devices such as servers. However, the central processing unit (CPU) of computing devices has a bottleneck in improving the computing power. Therefore, how to use heterogeneous acceleration cards to improve the computing power of computing devices has become a key research topic.

[0003] In the related art, a heterogeneous acceleration card is usually added to the computing device, and the CPU offloads the computing task to the heterogeneous acceleration card. When the heterogeneous acceleration card completes the computing task based on the acceleration core, it sends an interrupt request to the CPU to notify the CPU to read the task execution result from the heterogeneous acceleration card.

[0004] However, with the increase of acceleration cores of heterogeneous acceleration cards, the parallel computing capabilities of heterogeneous acceleration cards are continuously improved. When the CPU frequently issues computing tasks to give full play to the parallel computing capabilities of heterogeneous acceleration cards, the heterogeneous acceleration card will send a large number of interrupt requests to the CPU in a short period of time. The overall performance of the computing device will be reduced due to the CPU's frequent response to interrupt requests. Summary of the invention

[0005] The present application provides a heterogeneous acceleration system, method, computing device and storage medium to solve the defects in the related art that a heterogeneous acceleration card sends a large number of interrupt requests to a CPU in a short period of time, and the computing device will suffer from a reduction in overall performance due to the CPU frequently responding to the interrupt requests.

[0006] The first aspect of the present application provides a heterogeneous acceleration system, including: a host end and a heterogeneous acceleration card;

[0007] The host end is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and offload the multiple subtasks to the heterogeneous acceleration card;

[0008] The heterogeneous accelerator card is used to receive and execute the multiple subtasks, generate a corresponding interrupt subrequest when any of the subtasks is executed, generate a target interrupt request when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, and send the target interrupt request to the host end;

[0009] The host end is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card.

[0010] In an optional implementation manner, the host end is used to:

[0011] Obtaining the number of acceleration cores of the heterogeneous acceleration card;

[0012] According to the number of acceleration cores of the heterogeneous acceleration card, the acceleration task is split into a plurality of subtasks, so that the subtask loads among the acceleration cores of the heterogeneous acceleration card are balanced.

[0013] In an optional implementation, the heterogeneous accelerator card includes:

[0014] A shared memory, used for caching the multiple subtasks offloaded by the host side;

[0015] The host end writes multiple subtasks and subtask address information corresponding to the acceleration task into the shared memory in a one-time write-to-memory manner, and the subtask address information represents the storage location of the subtask in the shared memory.

[0016] In an optional implementation, the heterogeneous accelerator card includes:

[0017] An instruction parsing module, used for receiving an acceleration request instruction sent by the host end, and sending a memory access request instruction to a memory access module in response to the acceleration request instruction;

[0018] The memory access module is used to receive the memory access request instruction, and in response to the memory access request instruction, read the corresponding target subtask address information from the shared memory, and cache the target subtask address information to the target cache;

[0019] Among them, after the host side caches the multiple subtasks to the shared memory, the host side sends an acceleration request instruction to the host side instruction parsing module, and the memory access request instruction at least includes the storage location of the target subtask address information in the shared memory. The instruction parsing module obtains the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction.

[0020] In an optional implementation, the heterogeneous accelerator card includes:

[0021] an acceleration control module, configured to receive the start-up calculation instruction issued by the instruction parsing module, and in response to the start-up calculation instruction, sequentially obtain the address information of the subtasks to be executed of each acceleration core from the target buffer, and issue the address information of the subtasks to be executed to the corresponding acceleration core, so that the acceleration core reads and executes the subtasks to be executed in the shared memory according to the address information of the subtasks to be executed;

[0022] When the instruction parsing module receives the acceleration request instruction sent by the host, it sends a start calculation instruction to the acceleration core control module.

[0023] In an optional implementation, the heterogeneous acceleration card includes: a plurality of acceleration cores;

[0024] The acceleration core is used to read the corresponding subtask to be executed from the shared memory according to the source address represented by the address information of the subtask to be executed, and execute the subtask to be executed to obtain the corresponding subtask execution result; according to the destination address represented by the address information of the subtask to be executed, write the subtask execution result into the shared memory.

[0025] In an optional implementation manner, the host end is used to:

[0026] In response to the target interrupt request, the subtask execution result is read from the shared memory according to the subtask address information corresponding to each of the subtasks.

[0027] In an optional implementation, the heterogeneous accelerator card includes:

[0028] An interrupt counting module, used for collecting interrupt sub-requests generated by each acceleration core on the heterogeneous acceleration card, and counting the collected interrupt sub-requests to determine the total amount of interrupt sub-requests currently obtained;

[0029] When the acceleration core completes the execution of any subtask to be executed, it generates a corresponding interrupt subrequest.

[0030] In an optional implementation manner, the interrupt counting module is further used to:

[0031] Feedback the total amount of currently obtained interrupt sub-requests to the acceleration control module;

[0032] The acceleration control module is used for generating a target interrupt request when the total amount of the interrupt sub-requests reaches a preset threshold, and sending the target interrupt request to the host end;

[0033] The preset threshold is equal to the total number of subtasks of the acceleration task.

[0034] In an optional implementation, upon receiving the start calculation instruction, the acceleration control module enters a cache read state to obtain configuration information of each acceleration core from a target cache;

[0035] The acceleration control module enters the acceleration core configuration state when obtaining the configuration information, so as to write the configuration information into the configuration register of the corresponding acceleration core, so that the acceleration core executes the subtask to be executed according to the configuration information; wherein the configuration information of the acceleration core includes the address information of the subtask to be executed;

[0036] When determining that the total amount of the interrupt sub-requests reaches a preset threshold, the acceleration control module enters an interrupt reporting state to generate a target interrupt request and sends the target interrupt request to the host end;

[0037] After sending the target interrupt request to the host end, the acceleration control module enters a waiting state to wait for the host end to unload to a subtask again.

[0038] A second aspect of the present application provides a heterogeneous acceleration method, which is applied to a host side, and the method includes:

[0039] Get any acceleration task of the computing device;

[0040] Splitting the acceleration task into a plurality of subtasks, and offloading the plurality of subtasks to a heterogeneous acceleration card, so as to execute the plurality of subtasks based on the heterogeneous acceleration card; wherein the heterogeneous acceleration card generates a corresponding interrupt subrequest when completing the execution of any of the subtasks, and generates a target interrupt request when the total amount of the currently obtained interrupt subrequests reaches a preset threshold;

[0041] In response to the target interrupt request, the execution result of the subtask is read from the heterogeneous acceleration card.

[0042] A third aspect of the present application provides a heterogeneous acceleration method, which is applied to a heterogeneous acceleration card. The method includes:

[0043] Receiving multiple subtasks unloaded by the host side; wherein, when the host side obtains any acceleration task of the computing device, the acceleration task is split into multiple subtasks;

[0044] Executing the plurality of subtasks, and generating a corresponding interrupt sub-request when any of the subtasks is completed;

[0045] When the total amount of interrupt sub-requests currently obtained reaches a preset threshold, a target interrupt request is generated and sent to the host end, so that the host end responds to the target interrupt request and reads the execution result of the sub-task from the heterogeneous acceleration card.

[0046] A fourth aspect of the present application provides a computing device, comprising: a heterogeneous acceleration system as described in the first aspect and various possible designs of the first aspect.

[0047] The fifth aspect of the present application provides a computer-readable storage medium, which stores computer execution instructions. When a processor executes the computer execution instructions, it implements the method described in the second aspect and various possible designs of the second aspect or the method described in the third aspect and various possible designs of the third aspect.

[0048] The sixth aspect of the present application provides a computer program product, including computer instructions, which are used to enable a computer to execute the method described in the second aspect and various possible designs of the second aspect or the method described in the third aspect and various possible designs of the third aspect.

[0049] The technical solution of this application has the following advantages:

[0050] The present application provides a heterogeneous acceleration system, method, computing device and storage medium, the system includes: a host end and a heterogeneous acceleration card; the host end is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and unload the multiple subtasks to the heterogeneous acceleration card; the heterogeneous acceleration card is used to receive and execute multiple subtasks, and when any subtask is completed, a corresponding interrupt subrequest is generated, and when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, a target interrupt request is generated, and the target interrupt request is sent to the host end; the host end is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card. The system provided by the above scheme sends the target interrupt request to the host end through the heterogeneous acceleration card when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, instead of sending it to the host end every time an interrupt subrequest is obtained, so as to avoid the heterogeneous acceleration card sending a large number of interrupt subrequests to the host end in a short time, thereby avoiding the host end from frequently responding to interrupt requests, thereby improving the overall performance of the computing device. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the following is a brief introduction to the drawings required for use in the embodiments or the related technical descriptions. Obviously, the drawings described below are some embodiments of the present application, and a person skilled in the art can also obtain other drawings based on these drawings.

[0052] Figure 1 A schematic diagram of the interaction process of the heterogeneous acceleration system provided in an embodiment of the present application;

[0053] Figure 2 A schematic diagram of the structure of a heterogeneous acceleration card provided in an embodiment of the present application;

[0054] Figure 3A schematic diagram of the structure of another heterogeneous acceleration card provided in an embodiment of the present application;

[0055] Figure 4 A schematic diagram of the structure of the instruction parsing module provided in the embodiment of the present application;

[0056] Figure 5 A schematic diagram of the coding structure of the address information of the subtask to be executed provided in an embodiment of the present application;

[0057] Figure 6 A schematic diagram of the structure of a memory access module provided in an embodiment of the present application;

[0058] Figure 7 A schematic diagram of the structure of an interrupt counting module provided in an embodiment of the present application;

[0059] Figure 8 A schematic diagram of the structure of an acceleration control module provided in an embodiment of the present application;

[0060] Fig. 9 A state transition diagram of an acceleration control module provided in an embodiment of the present application;

[0061] Fig.10 A schematic diagram of a heterogeneous acceleration method provided in an embodiment of the present application;

[0062] Fig.11 A schematic diagram of a process flow of another heterogeneous acceleration method provided in an embodiment of the present application;

[0063] Fig.12 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.

[0064] The above drawings have shown clear embodiments of the present application, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0066] In addition, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. In the description of the following embodiments, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0067] With the widespread popularity and publicity of AI chips and other specialized chips in recent years, more and more people have gradually realized that using field programmable gate arrays (FPGAs) or specialized chips to implement certain traditional software algorithms can bring several orders of magnitude acceleration effects. At present, hardware acceleration cards based on FPGAs or application-specific integrated circuits (ASICs) have been widely used in cloud platforms, network data centers, etc.

[0068] FPGA-based heterogeneous acceleration cards are a commonly used heterogeneous device with powerful parallel computing capabilities and flexibility, and are widely used. Currently, the more common acceleration operators in FPGAs include compression, decompression, regular expression matching, encryption and decryption, and database filtering. These acceleration operators can be designed as independent IP cores (acceleration cores) to complete independent computing acceleration tasks. Usually in a heterogeneous acceleration system, the CPU will handle other tasks after unloading related tasks. When the accelerator (acceleration card) completes the calculation, it notifies the CPU through an interrupt to take the calculation results, or continue to issue other acceleration tasks. However, with the increase in the types and number of acceleration cores, in the case of high-density IP interrupt generation, the CPU will need to frequently process interrupt programs, constantly switch states, and communicate with the FPGA in handshakes, which will reduce the system's work efficiency and fail to achieve the acceleration effect.

[0069] In view of the above problems, the embodiment of the present application provides a heterogeneous acceleration system, method, computing device and storage medium, the system comprising: a host end and a heterogeneous acceleration card; the host end is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and unload the multiple subtasks to the heterogeneous acceleration card; the heterogeneous acceleration card is used to receive and execute multiple subtasks, and when any subtask is completed, a corresponding interrupt subrequest is generated, and when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, a target interrupt request is generated, and the target interrupt request is sent to the host end; the host end is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card. The system provided by the above scheme sends the target interrupt request to the host end through the heterogeneous acceleration card when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, instead of sending it to the host end every time an interrupt subrequest is obtained, so as to avoid the heterogeneous acceleration card sending a large number of interrupt subrequests to the host end in a short time, thereby avoiding the host end from frequently responding to interrupt requests, thereby improving the comprehensive performance of the computing device.

[0070] The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0071] The embodiment of the present application provides a heterogeneous acceleration system for implementing heterogeneous acceleration of a computing device while preventing the host from reducing the overall performance of the computing device due to frequent responses to interrupt requests sent by a heterogeneous acceleration card.

[0072] like Figure 1 As shown, it is a schematic diagram of the interaction process of the heterogeneous acceleration system provided in an embodiment of the present application. The heterogeneous acceleration system includes: a host end and a heterogeneous acceleration card.

[0073] Among them, the host side is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and offload the multiple subtasks to the heterogeneous acceleration card; the heterogeneous acceleration card is used to receive and execute multiple subtasks, and when any subtask is completed, a corresponding interrupt subrequest is generated. When the total amount of interrupt subrequests currently obtained reaches a preset threshold, a target interrupt request is generated and sent to the host side; the host side is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card.

[0074] Among them, the preset threshold can be set according to the actual scenario.

[0075] It should be noted that the computing device can specifically refer to the server, and the acceleration tasks include server compression, decompression, regular expression matching, encryption and decryption, database filtering and other computing tasks that can be offloaded to the heterogeneous acceleration card. The host side can specifically refer to the CPU deployed on the service, etc., and the heterogeneous acceleration card can specifically refer to FPGA deployed on the service.

[0076] Specifically, in order to give full play to the computing resources of the heterogeneous acceleration card and to improve the efficiency of heterogeneous acceleration, the host side first splits the computing task into multiple subtasks after obtaining the acceleration task, and unloads the subtasks to the heterogeneous acceleration card. The heterogeneous acceleration card performs parallel processing of subtasks based on the multiple acceleration cores deployed by it. Each time the acceleration core completes a subtask, it generates a corresponding interrupt subrequest. When the total amount of interrupt subrequests accumulated by the heterogeneous acceleration card reaches a preset threshold, a target interrupt request is sent to the host side to remind the host side to read the obtained subtask execution results from the heterogeneous acceleration card. When the host side obtains the execution results of all subtasks of a single computing task, the subtask execution results are spliced ​​to obtain the target execution result of the computing task. The heterogeneous acceleration system provided in the embodiment of the present application can effectively reduce the problem of low system working efficiency caused by frequent interruptions of the heterogeneous acceleration card to the CPU, reduce the interaction frequency between the host side and the acceleration IP core (heterogeneous acceleration card), and improve the working efficiency of the computing device system.

[0077] On the basis of the above embodiments, since the heterogeneous acceleration card includes multiple acceleration cores, in order to fully utilize the acceleration core computing resources of the heterogeneous acceleration card while achieving load balancing between the acceleration cores, as an implementable method, in one embodiment, the host side is used to obtain the number of acceleration cores of the heterogeneous acceleration card; according to the number of acceleration cores of the heterogeneous acceleration card, the acceleration task is split into multiple subtasks to balance the subtask load among the acceleration cores of the heterogeneous acceleration card.

[0078] It should be noted that the host side divides a time-consuming acceleration task into multiple subtasks with small data volumes. The calculation of each subtask is performed by an acceleration core, and multiple subtasks are distributed in multiple acceleration cores in a time-sharing manner, making full use of the parallelism of the acceleration cores of the heterogeneous acceleration card FPGA. In this way, the bandwidth of the data path can be fully utilized to improve the acceleration effect of the computing device.

[0079] Specifically, the host side can divide the acceleration tasks according to the acceleration tasks and the support capabilities of the heterogeneous acceleration card. Taking the acceleration task as a compression task as an example, if the heterogeneous acceleration card has a 16-way compression acceleration IP core (acceleration core), then the acceleration task can be split into subtasks that are multiples of 16, such as 16 subtasks or 32 subtasks, etc., to achieve load balancing of the acceleration cores of the heterogeneous acceleration card.

[0080] Based on the above embodiment, as an implementable method, Figure 2 , which is a schematic diagram of the structure of a heterogeneous acceleration card provided in an embodiment of the present application. In one embodiment, the heterogeneous acceleration card includes: a shared memory.

[0081] Shared memory is used to cache multiple subtasks offloaded on the host side.

[0082] The host end writes multiple subtasks and subtask address information corresponding to the acceleration task into the shared memory in a one-time write-to-memory manner, and the subtask address information represents the storage location of the subtask in the shared memory.

[0083] It should be noted that the host side can actually send subtask data to the heterogeneous acceleration card FPGA by writing registers, but this is too inefficient and will waste too many host resources and time when the computing task is very cumbersome. Therefore, the subtask data (subtask and subtask address information) is transferred by writing to the memory once. After the data is written, the heterogeneous acceleration card FPGA is informed through registers. After the data is written, the parameters can be parsed.

[0084] Specifically, Figure 2 The heterogeneous accelerator card shown is based on FPGA. FPGA generally uses advanced development tools such as HLS to develop and design the IP core (acceleration core) of the functional unit. The interface of the IP core can be designed as AXI-Lite or AXI4 type, which can be easily integrated into the overall design of FPGA. The encryption IP core implemented by HLS and the conventional connection method are as follows Figure 2 As shown. The heterogeneous acceleration card includes a communication module, an acceleration core and a shared memory. An Axi_lite interface (bus interface) is provided between the communication module and the acceleration core. The acceleration core is also provided with a clock input interface (clk_in) for receiving a clock signal and a reset interface (reset_n) for receiving a reset signal. The communication module controls and configures the acceleration core through the Axi_lite interface. When the communication module determines that the subtask unloaded from the host side has been cached to the shared memory, it controls the acceleration core AXI4_MM interface to access the shared memory to read the subtask from the shared memory. After the acceleration core completes the execution, it sends an interrupt subrequest to the communication module through the irq_ip interface. The AXI_Lite interface is also called Advanced eXtensibleInterface Lite, which is a bus interface protocol commonly used in hardware designs such as system-on-chip (SoC) and FPGA.

[0085] Specifically, in one embodiment, if Figure 3FIG. 1 is a schematic diagram of the structure of another heterogeneous acceleration card provided in an embodiment of the present application, wherein the heterogeneous acceleration card includes:

[0086] An instruction parsing module, used for receiving an acceleration request instruction sent by the host end, and sending a memory access request instruction to the memory access module in response to the acceleration request instruction;

[0087] The memory access module is used to receive a memory access request instruction, and in response to the memory access request instruction, read the corresponding target subtask address information (parameter information) from the shared memory, and cache the target subtask address information to the target cache.

[0088] Among them, after the host side caches multiple subtasks into the shared memory, it sends an acceleration request instruction to the host side instruction parsing module. The memory access request instruction at least includes the storage location of the target subtask address information in the shared memory. The instruction parsing module obtains the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction.

[0089] It should be noted that in Figure 3 In the heterogeneous accelerator card shown, the heterogeneous accelerator card communicates with the host through the direct access memory interface module (XDMA), PCIe_RX represents PCIe reception, that is, the heterogeneous accelerator card receives data transmitted from the host through the PCIe bus, PCIe_TX represents PCIe transmission, that is, the heterogeneous accelerator card sends data to the host through the PCIe bus, and the first crossbar switch module (crossbar_axi4) realizes the function of multiple initiator Masters sharing one memory when multiple Axi4 interface modules need to access the shared memory through the first crossbar switch module; the second crossbar switch module (crossbar_axi_lite) realizes the function of one Master controlling multiple responder Slave modules, and the interface uses the Axi_lite protocol; the XDMA module (direct access memory interface module) mainly realizes the analysis of the PCIe protocol, realizes the transmission of host memory data to FPGA memory data through DMA, and realizes the control of subsequent modules through the Axi_lite interface (by writing registers).

[0090] Specifically, after receiving the acceleration request instruction indicating that the data transmission of the host terminal task is completed, the instruction parsing module determines the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction. For the instruction parsing module, the target subtask is the subtask that the host side unloads to the heterogeneous acceleration card. The instruction parsing module generates a memory access request instruction based on the determined storage location of the target subtask address information in the shared memory, and sends the memory access request instruction to the memory access module. The memory access module responds to the memory access request instruction, reads the corresponding target subtask address information from the specified location of the shared memory (the storage location of the target subtask address information in the shared memory), and caches the target subtask address information to the target buffer (RAM) for subsequent modules to read.

[0091] Accordingly, in one embodiment, a heterogeneous acceleration card includes: an acceleration control module, which is used to receive a startup computing instruction issued by an instruction parsing module, and in response to the startup computing instruction, obtain the address information of the subtask to be executed of each acceleration core from the target cache in sequence, and send the address information of the subtask to be executed to the corresponding acceleration core, so that based on the acceleration core, the subtask to be executed is read and executed in the shared memory according to the address information of the subtask to be executed.

[0092] When the instruction parsing module receives the acceleration request instruction sent by the host, it sends a start calculation instruction to the acceleration core control module.

[0093] Specifically, after receiving the notification to start the calculation (start calculation instruction), the acceleration control module sequentially obtains the address information of the subtask to be executed of each acceleration core from the target cache, and sends the address information of the subtask to be executed to the corresponding acceleration core to start the acceleration core, that is, to control the acceleration core to start the execution of the subtask.

[0094] Among them, Figure 4 As shown, it is a structural diagram of the instruction parsing module provided in the embodiment of the present application. The instruction parsing module includes an Axi-lite response end protocol implementation module and a CPU command parsing module. The Axi-lite response end protocol implementation module is used to receive the acceleration request instruction sent by the host side. The instruction parsing module is connected to the CPU side (host side) through the Axi-lite interface output from the direct access memory interface module xdma-ip. Therefore, it is necessary to implement the parsing function of the Axi-lite protocol based on the CPU command parsing module to obtain the acceleration request instruction information of the host side. After receiving the acceleration request instruction, a start_read_ram start signal (memory access request instruction) and a start calculation instruction are sent to the subsequent module (memory access module).

[0095] Specifically, in one embodiment, a heterogeneous acceleration card includes: multiple acceleration cores; the acceleration core is used to read the corresponding subtask to be executed from the shared memory according to the source address represented by the address information of the subtask to be executed, and execute the subtask to be executed to obtain the corresponding subtask execution result; according to the destination address represented by the address information of the subtask to be executed, the subtask execution result is written into the shared memory.

[0096] For example, Figure 5 As shown, it is a schematic diagram of the encoding structure of the address information of the subtask to be executed provided in the embodiment of the present application. The encoding of the address information of the subtask to be executed is written into the shared memory by the host side in a one-time write-into-memory manner. With the feature of the internal module bit width of the heterogeneous acceleration card of 512 bits (64 bytes), 20 bytes are used for each calculation. In order to reduce the number of data transmissions, the address information of every two calculations is configured together. The heterogeneous acceleration card stipulates that it supports up to 8192 calculations. The data size of each calculation is not fixed, and the number of calculations is not fixed. Among them, the source address represents the storage address (starting address) of the subtask in the shared memory, the destination address represents the storage address of the execution result of the subtask, the subtask size is represented as size, and Cnt_num represents that the current acceleration task is split into Cnt_num subtasks.

[0097] Among them, Figure 6 As shown, it is a structural diagram of the memory access module provided by the embodiment of the present application. The memory access module includes an Axi4 protocol implementation module, a control module and a dual-port RAM access module. The memory access module implements two functions: reading parameter information (target subtask address information) from the hbm / ddr memory (shared memory) based on the Axi4 protocol implementation module and storing the valid parameter information in the dual-port RAM (target buffer). After receiving the start_read_ram start signal (memory access request instruction), the control module sends a control signal to start the Axi4_master interface protocol conversion function, reads parameter information from the specified address of the hbm / ddr memory, and stores the valid parameter information in the RAM according to the encoding rules of the information. According to the content of the previous section, the valid information of each calculation instruction is 20 bytes, that is, the valid information does not include the reserved bytes. Through the dual-port RAM access module, the valid information of each calculation is sequentially stored in the dual-port RAM. The write control port of the RAM is controlled by the memory access module, and the read control port is controlled by the acceleration control module to read out the parameter information.

[0098] Accordingly, in one embodiment, the host side responds to the target interrupt request and reads the subtask execution result from the shared memory according to the subtask address information corresponding to each subtask.

[0099] Specifically, after receiving the interrupt request, the host side reads the subtask execution result from the shared memory according to the destination address represented by the subtask address information.

[0100] Based on the above embodiment, as an implementable manner, in one embodiment, the heterogeneous accelerator card includes:

[0101] The interrupt counting module is used to collect the interrupt sub-requests generated by each acceleration core on the heterogeneous acceleration card, and count the collected interrupt sub-requests to determine the total amount of interrupt sub-requests currently obtained.

[0102] When the acceleration core completes the execution of any subtask to be executed, it generates a corresponding interrupt subrequest.

[0103] Specifically, Figure 7 As shown, it is a structural schematic diagram of the interrupt counting module provided in an embodiment of the present application. The interrupt counting module includes a basic counting unit and an upper counting unit of each acceleration core, so as to count the counting results of the interrupt sub-requests generated by each acceleration core based on the basic counting unit. The upper counting unit obtains the total number of interrupt sub-requests currently obtained by counting the counting results of each basic counting unit.

[0104] Specifically, in one embodiment, the interrupt counting module is also used to feed back the total amount of interrupt sub-requests currently obtained to the acceleration control module; the acceleration control module is used to generate a target interrupt request when the total amount of interrupt sub-requests reaches a preset threshold, and send the target interrupt request to the host end.

[0105] Among them, in order to enable the host side to read the subtask execution results of the entire acceleration task at one time, the preset threshold value can be equal to the total number of subtask splits of the acceleration task (Cnt_num). The setting of the interrupt counting module and the acceleration control module is equivalent to designing a "shell" for the acceleration core, making the function of the IP core (acceleration core) with a fixed interface more flexible. It can not only reduce the impact of frequent interrupts on CPU performance, but also reduce the information interaction between the CPU and external devices, and sink part of the management and scheduling work of the CPU to the FPGA implementation, which releases the computing power of the CPU, reduces the interaction delay, and improves the work efficiency of the FPGA. Most of the IP cores of the same interface type generated by advanced tools such as Hls, as long as the interface types of all IP cores are consistent, can use the above-mentioned interrupt counting module and acceleration control module as the "shell" to improve the flexibility of the IP core. In actual applications, the form and meaning of the encoding format information of the host side can be modified according to the actual application scenario. The system provided in the embodiment of the present application can be widely used and promoted.

[0106] Specifically, Figure 8As shown, it is a structural diagram of the acceleration control module provided in an embodiment of the present application. The acceleration control module includes an Axi-lite initiator protocol implementation module and a control module. The acceleration control module is based on the Axi-lite initiator protocol implementation module, and sends the address information of the subtask to be executed to the corresponding acceleration core. Based on the control module, the generation and sending of the target interrupt request are controlled.

[0107] Specifically, in one embodiment, when the acceleration control module receives the start calculation instruction, it enters the read cache state to obtain the configuration information of each acceleration core from the target cache; when the acceleration control module obtains the configuration information, it enters the acceleration core configuration state to write the configuration information into the configuration register of the corresponding acceleration core, so that the acceleration core executes the subtask to be executed according to the configuration information; wherein, the configuration information of the acceleration core includes the address information of the subtask to be executed; when the acceleration control module determines that the total amount of interrupt sub-requests reaches a preset threshold, it enters the interrupt reporting state to generate a target interrupt request and send the target interrupt request to the host end; after sending the target interrupt request to the host end, the acceleration control module enters the waiting state to wait for the host end to unload to the subtask again.

[0108] The configuration information of the acceleration core includes the address information of the subtask to be executed, the interrupt enable configuration information, and the computing configuration, etc. The computing configuration represents the computing requirements of the subtask, such as the selection of the encryption algorithm, etc.

[0109] Among them, Fig. 9 As shown, it is a state transition diagram of the acceleration control module provided in an embodiment of the present application. After receiving the start signal (start calculation instruction), the state machine enters the state of reading RAM (reading cache state), obtains parameter information (configuration information of each acceleration core), and then writes the parameter information into the idle IP core according to the Axi-lite interface protocol and starts it to start calculation; if all calculations are not completed, the Read-RAM (reading cache state)-->Write_IP (acceleration core configuration state) is executed in a loop until all subtask calculations are completed, and enters the state of initiating interruption (interrupt reporting state), and after completion, it ends and returns to the Idel waiting state to wait for the next calculation to start.

[0110] Specifically, in the waiting state, the start signal (start calculation instruction) initiated by the instruction parsing module is waited for. After each start signal is initiated, there will be multiple calculation operations, which are determined by the value of the parameter cal_mun sent by the CPU. The heterogeneous acceleration card generally supports a maximum of 8192 times; in the read cache state, the acceleration control module will initiate a read operation on the RAM (target cache). The working parameters of the IP core can be configured once per clock cycle, and the parameter information read back is saved in the register. If there is a configurable IP core (idle acceleration core), it enters the next state, otherwise it waits in this state; in the acceleration core configuration state, the acceleration control module initiates a write register operation to the corresponding IP core to write the configuration information into the corresponding acceleration core. Configure registers, including the starting address (source address) of the data to be calculated (subtask to be executed), the destination address of the execution result, the size of the calculated data (subtask size), interrupt enable configuration, calculation configuration and other parameters. Initiate a write operation according to the Axi-lite-master interface protocol (Axi-lite initiator protocol). When all calculations are completed, it will enter the interrupt reporting state, otherwise it will return to the acceleration core configuration state to continue the next round of calculations; in the interrupt reporting state, the acceleration control module will initiate an irq_top interrupt signal (target interrupt request) to the host side to notify the host side that the acceleration task calculation is completed, and then the state returns to the starting Idle state (waiting state) to wait for the next calculation.

[0111] On the basis of the above embodiments, the host side's selection of acceleration tasks often directly affects the utilization rate of the heterogeneous acceleration system. As an implementable method, in one embodiment, when the host side obtains the computing tasks, it first analyzes the computational complexity and resource demand characteristics of each computing task, and selects computing-intensive tasks as acceleration tasks. For computing-intensive tasks, such as encryption and decryption of large-scale data, these tasks usually require a large amount of computing resources and take a long time to calculate. Since the heterogeneous acceleration card has a powerful parallel computing capability, it can significantly improve the computing efficiency when processing such tasks. Therefore, the host side prioritizes such tasks as acceleration tasks and unloads them to the heterogeneous acceleration card, thereby improving the computing efficiency of the tasks while improving the utilization rate of the heterogeneous acceleration system.

[0112] Specifically, in one embodiment, when there is no computationally intensive task suitable as an acceleration task on the host side, the host side predicts the execution time of each computing task on the host side and the execution time on the heterogeneous acceleration card based on the current resource situation of the host side and the resource situation of the heterogeneous acceleration card. If it is determined that the execution time of any computing task on the heterogeneous acceleration card is less than the execution time on the host side, the computing task is offloaded to the heterogeneous acceleration card for execution as an acceleration task, so as to further improve the computing efficiency of the computing device for the computing tasks.

[0113] The heterogeneous acceleration system provided by the embodiment of the present application includes: a host end and a heterogeneous acceleration card; the host end is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and unload the multiple subtasks to the heterogeneous acceleration card; the heterogeneous acceleration card is used to receive and execute multiple subtasks, and when any subtask is completed, a corresponding interrupt subrequest is generated, and when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, a target interrupt request is generated, and the target interrupt request is sent to the host end; the host end is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card. The system provided by the above scheme sends the target interrupt request to the host end through the heterogeneous acceleration card when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, instead of sending it to the host end every time an interrupt subrequest is obtained, so as to avoid the heterogeneous acceleration card from sending a large number of interrupt subrequests to the host end in a short period of time, thereby avoiding the host end from frequently responding to interrupt requests, thereby improving the overall performance of the computing device. In addition, the host can issue multiple subtasks at a time without issuing a startup instruction for each subtask separately, which reduces the number of communications between the host and the heterogeneous accelerator card and further improves the acceleration efficiency of the heterogeneous accelerator card.

[0114] The embodiment of the present application provides a heterogeneous acceleration method for realizing heterogeneous acceleration of a computing device while preventing the host from frequently responding to interrupt requests sent by a heterogeneous acceleration card, thereby reducing the overall performance of the computing device. The execution subject of the embodiment of the present application is the host in the heterogeneous acceleration system provided by the above embodiment.

[0115] like Fig.10 FIG. 1 is a flow chart of a heterogeneous acceleration method provided in an embodiment of the present application, the method comprising:

[0116] Step 1001, obtaining any acceleration task of a computing device;

[0117] Step 1002: split the acceleration task into multiple subtasks, and offload the multiple subtasks to the heterogeneous acceleration card, so as to execute the multiple subtasks based on the heterogeneous acceleration card; wherein, when the heterogeneous acceleration card completes the execution of any subtask, it generates a corresponding interrupt subrequest, and generates a target interrupt request when the total amount of the currently obtained interrupt subrequests reaches a preset threshold;

[0118] Step 1003: In response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card.

[0119] Regarding the heterogeneous acceleration method in this embodiment, its specific implementation has been described in detail in the embodiments of the system and will not be elaborated here.

[0120] The heterogeneous acceleration method provided in the embodiment of the present application is applied to the host side in the heterogeneous acceleration system provided in the above embodiment. Its implementation method and principle are the same and will not be repeated here.

[0121] The embodiment of the present application provides a heterogeneous acceleration method for realizing heterogeneous acceleration of a computing device while preventing the host from frequently responding to interrupt requests sent by a heterogeneous acceleration card, thereby reducing the overall performance of the computing device. The execution subject of the embodiment of the present application is the heterogeneous acceleration card in the heterogeneous acceleration system provided in the above embodiment.

[0122] like Fig.11 FIG. 1 is a flow chart of another heterogeneous acceleration method provided in an embodiment of the present application, the method comprising:

[0123] Step 1101, receiving multiple subtasks unloaded by the host side; wherein, when the host side obtains any acceleration task of the computing device, the acceleration task is split into multiple subtasks;

[0124] Step 1102, executing multiple subtasks, and generating a corresponding interrupt subrequest when any subtask is completed;

[0125] Step 1103, when the total amount of interrupt sub-requests currently obtained reaches a preset threshold, a target interrupt request is generated and sent to the host side, so that the host side responds to the target interrupt request and reads the execution result of the sub-task from the heterogeneous acceleration card.

[0126] Regarding the heterogeneous acceleration method in this embodiment, its specific implementation has been described in detail in the embodiments of the system and will not be elaborated here.

[0127] The heterogeneous acceleration method provided in the embodiment of the present application is applied to the heterogeneous acceleration card in the heterogeneous acceleration system provided in the above embodiment. Its implementation method and principle are the same and will not be repeated here.

[0128] An embodiment of the present application provides a computing device for deploying the heterogeneous acceleration system provided by the above embodiment.

[0129] like Fig.12 FIG. 1 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. The computing device includes: a heterogeneous acceleration system provided in the above embodiment.

[0130] The computing device provided in the embodiment of the present application is used to deploy the heterogeneous acceleration system provided in the above embodiment. The implementation method and principle of the heterogeneous acceleration are the same and will not be repeated here.

[0131] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the heterogeneous acceleration method provided in any of the above embodiments is implemented.

[0132] The storage medium containing computer executable instructions provided in the embodiments of the present application can be used to store computer executable instructions of the heterogeneous acceleration method provided in the aforementioned embodiments. The implementation method and principle are the same and will not be repeated here.

[0133] An embodiment of the present application provides a computer program product, including computer instructions, which are used to enable a computer to execute the heterogeneous acceleration method provided in the aforementioned embodiment.

[0134] The embodiments of the present application provide a computer program product that can be used to execute computer instructions of the heterogeneous acceleration method provided in the aforementioned embodiments. The implementation method and principle are the same and will not be repeated here.

[0135] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, modules or units, which can be electrical, mechanical or other forms.

[0136] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0137] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0138] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0139] A part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0140] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. The specific working process of the method described above can refer to the corresponding process in the aforementioned system embodiment, and will not be repeated here.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A heterogeneous acceleration system, characterized in that: include: Host side and heterogeneous accelerator cards; The host end is used to obtain any acceleration task of the computing device, split the acceleration task into multiple subtasks, and offload the multiple subtasks to the heterogeneous acceleration card; The acceleration task is a computationally intensive task selected by the host according to the computational complexity and resource requirement characteristics of each computing task obtained; The heterogeneous accelerator card is used to receive and execute the multiple subtasks, generate a corresponding interrupt subrequest when any of the subtasks is executed, generate a target interrupt request when the total amount of the currently obtained interrupt subrequests reaches a preset threshold, and send the target interrupt request to the host end; The host end is used to receive the target interrupt request, and in response to the target interrupt request, read the execution result of the subtask from the heterogeneous acceleration card; The heterogeneous accelerator card includes: An instruction parsing module is used to receive an acceleration request instruction sent by the host end, and send a memory access request instruction to the memory access module in response to the acceleration request instruction; A memory access module, used for receiving a memory access request instruction, reading corresponding target subtask address information from a shared memory in response to the memory access request instruction, and caching the target subtask address information into a target buffer; Among them, the host side writes multiple subtasks and subtask address information corresponding to the acceleration task into the shared memory by a one-time write-to-memory method, and the subtask address information represents the storage location of the subtask in the shared memory; after the host side caches multiple subtasks into the shared memory, the host side sends an acceleration request instruction to the host side instruction parsing module, and the memory access request instruction at least includes the storage location of the target subtask address information in the shared memory. The instruction parsing module obtains the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction.

2. The system according to claim 1, characterized in that The host side is used for: Obtaining the number of acceleration cores of the heterogeneous acceleration card; According to the number of acceleration cores of the heterogeneous acceleration card, the acceleration task is split into a plurality of subtasks, so that the subtask loads among the acceleration cores of the heterogeneous acceleration card are balanced.

3. The system according to claim 1, characterized in that The heterogeneous accelerator card includes: an acceleration control module, configured to receive the start-up calculation instruction issued by the instruction parsing module, and in response to the start-up calculation instruction, sequentially obtain the address information of the subtasks to be executed of each acceleration core from the target buffer, and issue the address information of the subtasks to be executed to the corresponding acceleration core, so that the acceleration core reads and executes the subtasks to be executed in the shared memory according to the address information of the subtasks to be executed; When the instruction parsing module receives the acceleration request instruction sent by the host, it sends a start calculation instruction to the acceleration core control module.

4. The system according to claim 3, characterized in that The heterogeneous acceleration card includes: a plurality of acceleration cores; The acceleration core is used to read the corresponding subtask to be executed from the shared memory according to the source address represented by the address information of the subtask to be executed, and execute the subtask to be executed to obtain the corresponding subtask execution result; according to the destination address represented by the address information of the subtask to be executed, write the subtask execution result into the shared memory.

5. The system according to claim 4, characterized in that The host side is used for: In response to the target interrupt request, the subtask execution result is read from the shared memory according to the subtask address information corresponding to each of the subtasks.

6. The system according to claim 1, characterized in that The heterogeneous accelerator card includes: An interrupt counting module, used for collecting interrupt sub-requests generated by each acceleration core on the heterogeneous acceleration card, and counting the collected interrupt sub-requests to determine the total amount of interrupt sub-requests currently obtained; When the acceleration core completes the execution of any subtask to be executed, it generates a corresponding interrupt subrequest.

7. The system according to claim 6, characterized in that The interrupt counting module is further used for: Feedback the total amount of currently obtained interrupt sub-requests to the acceleration control module; The acceleration control module is used for generating a target interrupt request when the total amount of the interrupt sub-requests reaches a preset threshold, and sending the target interrupt request to the host end; The preset threshold is equal to the total number of subtasks of the acceleration task.

8. The system according to claim 7, characterized in that When receiving the start calculation instruction, the acceleration control module enters the cache read state to obtain the configuration information of each acceleration core from the target cache; The acceleration control module enters the acceleration core configuration state when obtaining the configuration information, so as to write the configuration information into the configuration register of the corresponding acceleration core, so that the acceleration core executes the subtask to be executed according to the configuration information; wherein the configuration information of the acceleration core includes the address information of the subtask to be executed; When determining that the total amount of the interrupt sub-requests reaches a preset threshold, the acceleration control module enters an interrupt reporting state to generate a target interrupt request and sends the target interrupt request to the host end; After sending the target interrupt request to the host end, the acceleration control module enters a waiting state to wait for the host end to unload to a subtask again.

9. A heterogeneous acceleration method, characterized in that: Applied to the host side, the method includes: Obtain any acceleration task of the computing device; the acceleration task is a computing-intensive task selected by the host end according to the computing complexity and resource requirement characteristics of each computing task obtained; Splitting the acceleration task into a plurality of subtasks, and offloading the plurality of subtasks to a heterogeneous acceleration card, so as to execute the plurality of subtasks based on the heterogeneous acceleration card; wherein the heterogeneous acceleration card generates a corresponding interrupt subrequest when completing the execution of any of the subtasks, and generates a target interrupt request when the total amount of the currently obtained interrupt subrequests reaches a preset threshold; In response to the target interrupt request, reading the execution result of the subtask from the heterogeneous accelerator card; Wherein, the heterogeneous acceleration card includes: An instruction parsing module is used to receive an acceleration request instruction sent by the host end, and send a memory access request instruction to the memory access module in response to the acceleration request instruction; A memory access module, used for receiving a memory access request instruction, reading corresponding target subtask address information from a shared memory in response to the memory access request instruction, and caching the target subtask address information into a target buffer; Among them, the host side writes multiple subtasks and subtask address information corresponding to the acceleration task into the shared memory by a one-time write-to-memory method, and the subtask address information represents the storage location of the subtask in the shared memory; after the host side caches multiple subtasks into the shared memory, the host side sends an acceleration request instruction to the host side instruction parsing module, and the memory access request instruction at least includes the storage location of the target subtask address information in the shared memory. The instruction parsing module obtains the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction.

10. A heterogeneous acceleration method, characterized in that: Applied to a heterogeneous accelerator card, the method includes: Receiving multiple subtasks unloaded by the host side; wherein, when the host side obtains any acceleration task of the computing device, the acceleration task is split into multiple subtasks; the acceleration task is a computing-intensive task selected by the host side according to the computing complexity and resource requirement characteristics of each computing task obtained; Executing the plurality of subtasks, and generating a corresponding interrupt sub-request when any of the subtasks is completed; When the total amount of the currently obtained interrupt sub-requests reaches a preset threshold, a target interrupt request is generated and sent to the host end, so that the host end reads the execution result of the sub-task from the heterogeneous accelerator card in response to the target interrupt request; The heterogeneous accelerator card includes: An instruction parsing module is used to receive an acceleration request instruction sent by the host end, and send a memory access request instruction to the memory access module in response to the acceleration request instruction; A memory access module, used for receiving a memory access request instruction, reading corresponding target subtask address information from a shared memory in response to the memory access request instruction, and caching the target subtask address information into a target buffer; Among them, the host side writes multiple subtasks and subtask address information corresponding to the acceleration task into the shared memory by a one-time write-to-memory method, and the subtask address information represents the storage location of the subtask in the shared memory; after the host side caches multiple subtasks into the shared memory, the host side sends an acceleration request instruction to the host side instruction parsing module, and the memory access request instruction at least includes the storage location of the target subtask address information in the shared memory. The instruction parsing module obtains the storage location of the target subtask address information in the shared memory by parsing the acceleration request instruction.

11. A computing device, characterized in that: include: A heterogeneous acceleration system as claimed in any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to claim 9 or 10 is implemented.

13. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method according to claim 9 or 10.

Citation Information

Patent Citations

  • DMA communication system and method for RDMA communication equipment

    CN113742267A

  • Data near-storage calculation method and device and storage medium

    CN116627892A