Heterogeneous calculation acceleration method based on fpga
By building a CPU+FPGA heterogeneous computing system and utilizing the parallel processing and dynamic reconfigurability of FPGA, the problem that FPGA cannot accelerate computing after image acquisition is solved, and efficient parallel processing and pipeline processing of computing tasks are achieved, which improves computing performance and throughput, especially improving detection speed in SAR image detection.
Patent Information
- Application Number
- CN202510789314.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, FPGA cannot perform calculation acceleration after image acquisition, and its performance and power consumption are relatively high, which cannot meet the needs of large-scale computing tasks.
By adopting the FPGA-based heterogeneous computing acceleration method, a CPU+FPGA heterogeneous computing system is constructed through the steps of task decomposition and coordination, accelerator construction, data flow optimization, task parallel processing, dynamic reconstruction and software and hardware coordination. The parallel processing capability and dynamic reconfiguration characteristics of FPGA are utilized to optimize data flow and resource allocation to achieve efficient computing.
It greatly improves the computing throughput and algorithm running speed, reduces software-level overhead, and is suitable for executing highly repetitive computing tasks and matrix operations, especially in the detection of train wheel tread defects in SAR images, significantly shortening the detection time.
Smart Images

Figure CN120708035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computing technology, and in particular to an FPGA-based heterogeneous computing acceleration method. Background Art
[0002] Heterogeneous computing is a special form of parallel and distributed computing. It uses a single independent computer that can support both SIMD and MIMD methods, or a group of independent computers interconnected by a high-speed network to complete computing tasks. The advantage of heterogeneous computing is that it can combine the characteristics of different computing units. CPUs are good at handling general computing tasks, GPUs perform well in parallel processing and floating-point operations, and FPGAs provide customizable hardware acceleration capabilities. They are suitable for executing specific algorithms and protocols and can be applied in a wide range of scenarios, including the comprehensive size detection of high-speed rail wheelsets and the use of logical operations such as corrosion and dilation commonly used in image morphological processing to locate bridge positions.
[0003] In Chinese patent 202410241050.2, image acquisition can only be performed through FPGA, and calculation acceleration cannot be performed. Since FPGA has a high performance-to-power ratio and limited acceleration capability, it cannot meet the computing requirements of some large-scale computing tasks. Therefore, the present invention proposes a heterogeneous computing acceleration method based on FPGA to solve the problems existing in the prior art. Summary of the Invention
[0004] In response to the above problems, the purpose of the present invention is to propose a heterogeneous computing acceleration method based on FPGA, which can realize highly parallel data processing and is suitable for executing highly repetitive computing tasks, matrix operations and signal processing. Through pipeline technology, FPGA can process multiple data within one clock cycle, greatly improving throughput. The direct hardware implementation of FPGA reduces the overhead at the software level, thereby reducing calculation.
[0005] To achieve the purpose of the present invention, the present invention is implemented through the following technical solutions: a method for accelerating heterogeneous computing based on FPGA, comprising the following steps:
[0006] Step 1: Task decomposition and coordination. A high-level language is used to describe the train wheel tread defect detection algorithm based on SAR images using corrosion and dilation operations. Multiple FPGA acceleration units are used to collaborate with the CPU main control unit to complete the same computing task. Multiple FPGA acceleration units collaborate with the CPU main control unit to form a CPU+FPGA heterogeneous computing system, which decomposes the computing task into multiple subtasks that can be processed in parallel.
[0007] Step 2: Build an accelerator. After parallel processing of the subtasks, the accelerator is built using the digital signal processor, lookup table, registers and on-chip storage resources inside the FPGA.
[0008] Step 3: Data flow optimization. After the accelerator is built, the acceleration unit inside the FPGA deploys different algorithm modules to optimize the flow of data inside the FPGA, reduce data transmission delay and overhead, and achieve efficient memory access mode.
[0009] Step 4: Task parallel processing. During the algorithm module processing, multiple identical algorithm modules are deployed inside the same FPGA acceleration unit, and different FPGA acceleration units deploy different algorithm modules. The upper-level FPGA acceleration unit processes the data in parallel.
[0010] Step 5: Dynamic Reconfiguration: After the data is processed in parallel, the dynamic reconfiguration characteristics of FPGA are used to dynamically configure hardware resources according to different computing requirements. The FPGA acceleration unit is divided into a static area and a dynamically reconfigurable area.
[0011] Step 6: Software-hardware collaboration. During the switching of computing tasks, a software-hardware collaborative design method is used to migrate image processing, video encoding, and signal processing tasks suitable for execution on the FPGA to the FPGA, while retaining operating system management, file system operations, and database query tasks on the CPU.
[0012] Step 7: Platform management. After completing all computing tasks, manage all resources, manage resource allocation and task scheduling on the heterogeneous computing platform, achieve load balancing, and ensure that each computing unit works efficiently.
[0013] The further improvement lies in that: in the step one, when decomposing multiple groups of tasks, multiple groups of subtasks are mapped to different parts of the FPGA according to the resources and performance characteristics of the FPGA. At the same time, the CPU+FPGA heterogeneous computing system is programmed using the OpenCL programming model. The FPGA acceleration unit is interconnected and communicated with the CPU main control unit through the PCIe-DMA bus, and multiple FPGA acceleration units are interconnected and communicated using the SRIO bus.
[0014] A further improvement is that in step 2, when building the accelerator, the CPU main control unit is responsible for logical judgment, management control, and allocation of computing tasks to the FPGA acceleration unit, and the FPGA acceleration unit accelerates the computing tasks, and the FPGA acceleration unit is internally divided into a static area and a dynamically reconfigurable area.
[0015] A further improvement is that: in the step three, the acceleration unit inside the FPGA deploys different algorithm modules, and there is a logical cascade relationship between the algorithm modules. The CPU main control unit sends the data to be processed to the first algorithm module for processing. The first group of algorithm modules inputs the processing results to the second group of algorithm modules for further processing. The second group of algorithm modules then transmits them to the third group of algorithm modules, and so on. When each algorithm module completes the processing, it notifies the upper-level algorithm module in an interrupt manner to receive the new processing data to fully utilize the parallel processing capability of the FPGA.
[0016] A further improvement is that in step 4, the acceleration unit processes multiple sets of data and transmits the processing results to the next-level FPGA acceleration unit, utilizing the parallel processing capability of the FPGA to execute multiple computing tasks simultaneously.
[0017] A further improvement is that in step five, PCIe-DMA communication, SRIO communication and DDR control are performed through the static area of the FPGA acceleration unit, and the kernel function issued by the CPU main control unit is executed through the dynamic reconfigurable area of the FPGA acceleration unit to accelerate the computing task and realize fast switching between different computing tasks.
[0018] Further improvements are: in step six, in the process of software and hardware collaborative conversion, the exchange of data and control signals is realized through efficient software and hardware interfaces, and in the FPGA heterogeneous computing nodes, dynamic allocation scheduling technology is used to achieve efficient collaboration between the CPU and FPGA heterogeneous computing chips inside each FPGA heterogeneous computing node.
[0019] A further improvement is that in step seven, when each node in the heterogeneous computing platform performs a verification operation, the corresponding FPGA heterogeneous computing is performed through the FPGA heterogeneous computing chip in each node, the consensus algorithm embedded in the FPGA heterogeneous computing node is determined, and the FPGA heterogeneous computing is used to implement the consensus operation in the platform.
[0020] A further improvement is that the FPGA heterogeneous computing chip is used to provide a programming interface for the CPU and provide its own job scheduling and online reconstruction.
[0021] A further improvement is that the FPGA heterogeneous computing node integrates multiple groups of high-speed bus protocols according to the heterogeneous protocol interconnection fusion standard, so that the FPGA heterogeneous computing node supports multiple heterogeneous interconnection protocols and improves the compatibility of the system.
[0022] The beneficial effects of the present invention are as follows: the present invention can realize highly parallel data processing, is suitable for executing highly repetitive computing tasks, matrix operations and signal processing, the FPGA processes multiple data within one clock cycle, greatly improving throughput, the direct hardware implementation of the FPGA reduces the software-level overhead, thereby reducing calculations, and running the above-generated executable file on the FPGA end to implement the SAR image train wheel tread defect detection algorithm for corrosion and expansion operations to obtain processing results, and running the generated executable file on the FPGA end greatly reduces the time required for running the SAR image train wheel tread defect detection algorithm, and significantly improves the running speed of the algorithm, and by deploying FPGA heterogeneous computing nodes, a system of multi-processor collaboration and parallel computing is constructed, and the FPGA heterogeneous computing chip can be responsible for computationally intensive, highly parallel computing tasks, and through the powerful computing power and sufficient flexibility provided by the FPGA, efficient collaboration between the CPU and the FPGA heterogeneous computing chip inside each FPGA heterogeneous computing node is achieved, and a heterogeneous architecture with multiple acceleration modes of parallel processing and pipeline processing is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 is a flow chart of the steps of the present invention;
[0025] Figure 2 This is a decomposition flow chart of step 1 of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] In the description of the present invention, it should be noted that, unless otherwise specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0028] In file 202410241050.2, the flexibility and high performance of FPGA are utilized to receive and process serial data from the camera, and convert the serial data into MIPIDPHY protocol format so that it is directly compatible with the ARM-based processing platform. The accuracy and reliability of data processing are improved by integrating an efficient frame control module, a data cache module and a MIPID-PHY transmission controller. However, since FPGA has a high performance-to-power ratio, highly parallel data processing can be achieved in this application, which is suitable for executing highly repetitive computing tasks, matrix operations and signal processing. FPGA can process multiple data within one clock cycle, and the FPGA acceleration unit is configured into parallel acceleration mode to realize a computing method that combines parallel acceleration processing and pipeline acceleration processing of computing tasks, which can greatly improve the task processing throughput, shorten the task execution time, and thus significantly improve the computing performance of the computer.
[0029] Example 1
[0030] according to Figure 1 As shown, this embodiment provides an FPGA-based heterogeneous computing acceleration method, including the following steps:
[0031] Step 1: Task decomposition and coordination. A high-level language is used to describe the train wheel tread defect detection algorithm based on SAR images using corrosion and dilation operations. Multiple FPGA acceleration units are used to collaborate with the CPU main control unit to complete the same computing task. Multiple FPGA acceleration units collaborate with the CPU main control unit to form a CPU+FPGA heterogeneous computing system, which decomposes the computing task into multiple subtasks that can be processed in parallel.
[0032] Step 2: Build an accelerator. After parallel processing of the subtasks, the accelerator is built using the digital signal processor, lookup table, registers and on-chip storage resources inside the FPGA.
[0033] Step 3: Data flow optimization. After the accelerator is built, the acceleration unit inside the FPGA deploys different algorithm modules to optimize the flow of data inside the FPGA, reduce data transmission delay and overhead, and achieve efficient memory access mode.
[0034] Step 4: Task parallel processing. During the algorithm module processing, multiple identical algorithm modules are deployed inside the same FPGA acceleration unit, and different FPGA acceleration units deploy different algorithm modules. The upper-level FPGA acceleration unit processes the data in parallel.
[0035] Step 5: Dynamic Reconfiguration: After the data is processed in parallel, the dynamic reconfiguration characteristics of FPGA are used to dynamically configure hardware resources according to different computing requirements. The FPGA acceleration unit is divided into a static area and a dynamically reconfigurable area.
[0036] Step 6: Software-hardware collaboration. During the switching of computing tasks, a software-hardware collaborative design method is used to migrate image processing, video encoding, and signal processing tasks suitable for execution on the FPGA to the FPGA, while retaining operating system management, file system operations, and database query tasks on the CPU.
[0037] Step 7: Platform management. After completing all computing tasks, manage all resources, manage resource allocation and task scheduling on the heterogeneous computing platform, achieve load balancing, and ensure that each computing unit works efficiently.
[0038] In step one, when decomposing multiple groups of tasks, multiple groups of subtasks are mapped to different parts of the FPGA according to the resources and performance characteristics of the FPGA. At the same time, the CPU+FPGA heterogeneous computing system is programmed using the OpenCL programming model. The FPGA acceleration unit communicates with the CPU main control unit through the PCIe-DMA bus, and multiple FPGA acceleration units communicate with each other using the SRIO bus.
[0039] In step 2, when building the accelerator, the CPU main control unit is responsible for logic judgment, management control, and computing task allocation to the FPGA acceleration unit, and the FPGA acceleration unit accelerates the computing tasks. The FPGA acceleration unit is divided into a static area and a dynamically reconfigurable area. PCIe-DMA communication, SRIO communication, and DDR control are performed through the static area of the FPGA acceleration unit, and the kernel function issued by the CPU main control unit is executed through the dynamically reconfigurable area of the FPGA acceleration unit to accelerate the computing tasks.
[0040] In step three, the acceleration unit inside the FPGA deploys different algorithm modules, and there is a logical cascade relationship between the algorithm modules. The CPU main control unit sends the data to be processed to the first algorithm module for processing. The first group of algorithm modules inputs the processing results to the second group of algorithm modules for further processing. The second group of algorithm modules then passes them to the third group of algorithm modules, and so on. When each algorithm module is processed, it notifies the upper-level algorithm module in the form of an interrupt to receive new processing data to fully utilize the parallel processing capabilities of the FPGA. The algorithm modules deployed inside the FPGA acceleration unit are developed in the C++ high-level language.
[0041] In step four, the acceleration unit processes multiple sets of data and transmits the processing results to the next-level FPGA acceleration unit, which uses the parallel processing capability of the FPGA to execute multiple computing tasks simultaneously.
[0042] In step five, PCIe-DMA communication, SRIO communication and DDR control are performed through the static area of the FPGA acceleration unit, and the kernel function issued by the CPU main control unit is executed through the dynamic reconfigurable area of the FPGA acceleration unit to accelerate the computing task and realize fast switching between different computing tasks.
[0043] In step six, during the process of collaborative conversion between software and hardware, the exchange of data and control signals is achieved through efficient software and hardware interfaces. In the FPGA heterogeneous computing nodes, dynamic allocation and scheduling technology is used to achieve efficient collaboration between the CPU and FPGA heterogeneous computing chips inside each FPGA heterogeneous computing node, realizing a heterogeneous architecture that supports multiple acceleration modes of parallel processing and pipeline processing.
[0044] In step seven, when each node in the heterogeneous computing platform performs a verification operation, the corresponding FPGA heterogeneous computing is performed through the FPGA heterogeneous computing chip in each node to determine the consensus algorithm embedded in the FPGA heterogeneous computing node. FPGA heterogeneous computing is used to implement consensus operations in the platform, effectively improve the computing performance of the platform, and solve the performance problems of the platform.
[0045] FPGA heterogeneous computing chip is used to provide a programming interface for the CPU and provide its own job scheduling and online reconstruction.
[0046] FPGA heterogeneous computing nodes integrate multiple sets of high-speed bus protocols according to the heterogeneous protocol interconnection fusion standard, so that FPGA heterogeneous computing nodes support multiple heterogeneous interconnection protocols and improve the compatibility of the system.
[0047] Example 2
[0048] In step one, multiple groups of subtasks are mapped to different parts of the FPGA. At the same time, it can be applied to the scenario of comprehensive size detection of high-speed railway wheelsets. Logical operations such as corrosion and dilation, which are commonly used in SAR image morphological processing, are used to locate the bridge position. The SAR image train wheel tread defect detection algorithm based on corrosion and dilation operations is compiled into an AOCX executable file. Then, the high-resolution SAR image to be processed is sent to the FPGA board, and the compiled AOCX executable file is run on the FPGA board to finally obtain the processing results of the target data.
[0049] When using the FPGA-based heterogeneous computing acceleration method, multiple FPGA acceleration units are used to collaboratively complete the same computing task. For different task types, the FPGA acceleration mode can be configured to a parallel acceleration mode to achieve parallel acceleration processing and pipeline acceleration processing of computing tasks, which can greatly improve the task processing throughput and shorten the task execution time. Multiple identical algorithm module kernels are deployed in all FPGA acceleration units, and the kernels are executed in parallel. The CPU main control unit serves as a management unit and sends different data segments to be processed to different kernels for processing. The kernels process each data segment in parallel and return the processing results to the CPU main control unit. The CPU main control unit integrates the processing results. Different FPGA acceleration units deploy different kernels. The upper-level FPGA acceleration unit processes the data in parallel and transmits the processing results to the lower-level FPGA acceleration unit for further processing, thereby forming an FPGA acceleration unit that improves the use speed.
[0050] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A heterogeneous computing acceleration method based on FPGA, characterized in that: The following steps are included: Step 1: Task decomposition and coordination: A high-level language is used to describe the train wheel tread defect detection algorithm based on SAR images using erosion and dilation operations. Multiple FPGA acceleration units are used to collaborate with the CPU main control unit to complete the same computing task. Step 2: Build an accelerator. After parallel processing of the subtasks, the accelerator is built using the digital signal processor, lookup table, registers and on-chip storage resources inside the FPGA. Step 3: Data flow optimization. After the accelerator is built, the acceleration unit inside the FPGA deploys different algorithm modules to optimize the flow of data inside the FPGA, reduce data transmission delay and overhead, and achieve efficient memory access mode. Step 4: Task parallel processing. During the algorithm module processing, multiple identical algorithm modules are deployed inside the same FPGA acceleration unit, and different FPGA acceleration units deploy different algorithm modules. The upper-level FPGA acceleration unit processes the data in parallel. Step 5: Dynamic Reconfiguration: After the data is processed in parallel, the dynamic reconfiguration characteristics of the FPGA are used to dynamically configure hardware resources according to different computing requirements. The FPGA acceleration unit is divided into a static area and a dynamically reconfigurable area. Step 6: Software-hardware collaboration. During the switching of computing tasks, a software-hardware collaborative design method is used to migrate image processing, video encoding, and signal processing tasks suitable for execution on the FPGA to the FPGA, while retaining operating system management, file system operations, and database query tasks on the CPU. Step 7: Platform management. After completing all computing tasks, manage all resources, manage resource allocation and task scheduling on heterogeneous computing platforms, and achieve load balancing.
2. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In the step 1, when decomposing multiple groups of tasks, multiple groups of subtasks are mapped to different parts of the FPGA according to the resources and performance characteristics of the FPGA. At the same time, the CPU+FPGA heterogeneous computing system is programmed using the OpenCL programming model.
3. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In the step 2, when building the accelerator, the CPU main control unit is responsible for logic judgment, management control, and computing task allocation to the FPGA acceleration unit, and the FPGA acceleration unit accelerates the computing task, and the FPGA acceleration unit is internally divided into a static area and a dynamically reconfigurable area.
4. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In the step three, the acceleration unit inside the FPGA deploys different algorithm modules, and there is a logical cascade relationship between the algorithm modules. The CPU main control unit sends the data to be processed to the first algorithm module for processing. The first group of algorithm modules inputs the processing results to the second group of algorithm modules for further processing. The second group of algorithm modules then transmits them to the third group of algorithm modules, and so on. When each algorithm module completes the processing, it notifies the upper-level algorithm module in an interrupt manner to receive the new processing data to fully utilize the parallel processing capability of the FPGA.
5. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In the step 4, the acceleration unit processes the multiple sets of data and transmits the processing results to the next level FPGA acceleration unit.
6. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In step five, PCIe-DMA communication, SRIO communication and DDR control are performed through the static area of the FPGA acceleration unit, and the kernel function issued by the CPU main control unit is executed through the dynamic reconfigurable area of the FPGA acceleration unit to accelerate the computing task and realize fast switching between different computing tasks.
7. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In step six, during the process of software and hardware collaborative conversion, the exchange of data and control signals is realized through the software and hardware interface, and the dynamic allocation scheduling technology is used in the FPGA heterogeneous computing node to achieve efficient collaboration between the CPU and the FPGA heterogeneous computing chip inside each FPGA heterogeneous computing node.
8. The FPGA-based heterogeneous computing acceleration method according to claim 1, characterized in that: In step seven, when each node in the heterogeneous computing platform performs a test operation, the corresponding FPGA heterogeneous computing is performed through the FPGA heterogeneous computing chip in each node to solve the performance problem of the platform.
9. The FPGA-based heterogeneous computing acceleration method according to claim 7, characterized in that: The FPGA heterogeneous computing chip is used to provide a programming interface for the CPU and provide its own job scheduling and online reconstruction.
10. The FPGA-based heterogeneous computing acceleration method according to claim 7, characterized in that: The FPGA heterogeneous computing node integrates multiple groups of high-speed bus protocols according to the heterogeneous protocol interconnection fusion standard, so that the FPGA heterogeneous computing node supports multiple heterogeneous interconnection protocols and improves the compatibility of the system.
Citation Information
Patent Citations
CPU+FPGA-based heterogeneous computing system and acceleration method thereof
CN108776649A
A bridge detection method based on FPGA heterogeneous computing
CN109472777A
Embedded high-performance heterogeneous computing platform based on FPGA and ARM
CN117271430A
SAR load on-satellite processing evaluation system and method based on CGRA chip
CN117271944A
Heterogeneous computing-based task processing method and software and hardware framework system
US20220350669A1