Real-time target detection device based on FPGA acceleration

CN118823704BActive Publication Date: 2026-09-29BEIJING MECHANICAL EQUIP INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310423425.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-09-29
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

[0005]鉴于上述的分析,本发明实施例旨在提供一种基于FPGA加速的实时目标检测装置,用以解决现有技术中目标检测模块存在检测速度慢以及不易维护的问题

Benefits of technology

[0047]与现有技术相比,本发明至少可实现如下有益效果之一:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823704B_ABST
    Figure CN118823704B_ABST
Patent Text Reader

Abstract

The application relates to a real-time target detection device based on FPGA acceleration and belongs to the technical field of image recognition, and solves the problem of slow detection speed of a target detection module in the prior art. The real-time target detection device comprises an image acquisition module, which is used for acquiring a to-be-detected image, performing pretreatment, and obtaining a to-be-detected feature map in a preset format; a PS end DDR, which is used for storing the to-be-detected feature map in the preset format and preset weights; a weight loading module, which is used for loading the preset weights in the PS end DDR into a PL end DDR during convolution operation; an acceleration module, which is used for determining the size of a convolution layer PE array; a target detection module, which is used for outputting a target detection result to the PS end DDR for storage; and a post-processing module, which is used for determining the position range and category of each target in the to-be-detected image and sending the position range and category to a visual device for display. The speed of the target detection module for detecting an image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a real-time target detection device based on FPGA acceleration. Background Technology

[0002] Automatic driving system refers to a train operation system in which the work performed by the train driver is fully automated and highly centrally controlled. Automatic driving system has functions such as automatic train wake-up and start-up and hibernation, automatic entry and exit from the depot, automatic cleaning, automatic driving, automatic stopping, automatic door opening and closing, and automatic fault recovery. It also has multiple operating modes such as normal operation, degraded operation, and operation interruption.

[0003] The primary function of autonomous driving systems is to accurately identify objects around the vehicle in real time to ensure safe and correct control decisions. Object detection modules based on deep learning and semantic segmentation algorithms have become crucial in the field of autonomous driving, as rapid inference speed is key to ensuring the safety of autonomous driving.

[0004] Currently, the inference speed of target detection modules in existing technologies is relatively slow, and they are also quite troublesome to maintain, which has brought inconvenience to the development of the autonomous driving field. Summary of the Invention

[0005] Based on the above analysis, the present invention aims to provide a real-time target detection device based on FPGA acceleration to solve the problems of slow detection speed and difficulty in maintenance of target detection modules in the prior art.

[0006] This invention provides a real-time target detection device based on FPGA acceleration. The real-time target detection device includes an image acquisition module, a PS-end DDR, a PL-end DDR, a weight loading module, a target detection module, an acceleration module, a post-processing module, and a visualization device.

[0007] The image acquisition module is used to acquire the image to be detected and perform preprocessing to obtain a feature map to be detected in a preset format;

[0008] The PS-side DDR is used to store the feature map to be detected and the preset weights in a preset format; the preset weights are weight data pre-trained based on the target detection module.

[0009] The weight loading module is used to load the preset weights in the PS-side DDR into the PL-side DDR during the convolution operation.

[0010] The acceleration module is used to determine the size of the convolutional layer PE array based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers.

[0011] The target detection module is used to schedule the feature map to be detected stored in DDR at the PS end and the weight data stored in DDR at the PL end during the convolution operation. It uses the PE array of convolutional layers to perform convolution operations on multiple convolutional layers in the target detection module and outputs the target detection results to the DDR at the PS end for storage.

[0012] The post-processing module is used to determine the location range and category of each target in the image to be detected based on the target detection results, and send it to the visualization device for display.

[0013] Based on further improvements to the above-mentioned device, the image acquisition module includes:

[0014] Video interleaved array camera group, used to acquire raw images;

[0015] The deserializer is used to deserialize the original image and output the image to be detected.

[0016] The preprocessing module is used to preprocess the image to be detected to obtain a feature map to be detected in a preset format.

[0017] Based on further improvements to the above-mentioned device, the preprocessing module includes one or more of the following:

[0018] The demosaic module is used to perform demosaicing on the image to be detected;

[0019] The gamma correction module is used to perform gamma correction on the image to be detected.

[0020] The telescopic module is used to perform telescopic actions on the image to be inspected;

[0021] The color space conversion module is used to perform color space conversion on the image to be detected.

[0022] Based on further improvements to the above device, the PS-side DDR includes a first image cache space and a second image cache space; the first image cache space is used to store multiple feature maps to be detected; each feature map to be detected is quantized by 8 bits and then reordered in the order of depth, row, and column, and each reordered feature map to be detected is mapped and stored in the second image cache space.

[0023] Based on further improvements to the above-mentioned device, the real-time target detection device also includes a weight training module;

[0024] The weight training module is used to train the object detection module using the publicly available KITTI dataset and obtain the weight data after training. The weight data after training is quantized into 8 bits, and the quantized weight data is stored as preset weights in the order of depth, row and column.

[0025] The weight loading module is also used to load preset weights from the SD card into the PS-side DDR during the convolution operation.

[0026] Based on a further improvement of the above device, the step of determining the size of the convolutional layer PE array according to the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers includes:

[0027] The parallelism of multiple execution depths and multiple execution columns is determined based on the input feature map size of the multiple convolutional layers included in the object detection module.

[0028] The number of execution convolution kernels is determined based on the number of convolution kernels in the multiple convolutional layers included in the object detection module;

[0029] The size of the convolutional layer PE array is determined based on the number of available DSPs at the PL end, the multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolution kernels; the size of the convolutional layer PE array is C*K*L, where C is the execution depth, K is the number of execution convolution kernels, and L is the parallelism of the execution columns.

[0030] Based on further improvements to the above-mentioned device, the step of determining the size of the convolutional layer PE array according to the number of available DSPs at the PL end, multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolutional kernels includes:

[0031] From multiple execution depths, multiple execution column parallelisms, and multiple execution convolution kernel numbers, select one execution depth, one execution column parallelism, and one execution convolution kernel number, respectively, such that the selected execution depth, execution column parallelism, execution convolution kernel number, and available DSP number satisfy the following conditions:

[0032] C*K*L≤2*N;

[0033] Where C*K*L represents the size of the convolutional layer PE array, N represents the number of available DSPs at the PL end, C represents the execution depth, K represents the number of execution convolutional kernels, and L represents the parallelism of the execution columns.

[0034] Based on a further improvement of the above device, the step of using a convolutional layer PE array to perform convolution operations on multiple convolutional layers in the target detection module includes:

[0035] For each of the plurality of convolutional layers, the input feature map H1*W1*P1 of the current convolutional layer is divided into columns according to the parallelism L of the execution columns and the execution stride S of the convolutional kernel of the current convolutional layer. There are 1 region, and the size of the input feature map in each region is H1*(L*S)*P1;

[0036] Perform the following steps on each region sequentially, from the first column to the last column of the input feature map of the current convolutional layer:

[0037] Based on the row X of the current convolutional kernel and the stride S of the current convolutional kernel, the input feature map H1*(L*S)*P1 of the current region is divided into... Each sub-region contains an input feature map of size X*(L*S)*P1; the kernel size of the current convolutional layer is X*Y*P1, where X is the row, Y is the column, and P1 is the depth.

[0038] Convolution operations are performed on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer.

[0039] Based on a further improvement of the above device, the step of performing convolution operations on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer includes:

[0040] Based on the execution depth C of the convolutional layer PE array, the current sub-region X*(L*S)*P1 is divided into P1 / C spaces, and the input feature map size of each space is X*(L*S)*C.

[0041] Based on the input feature map of the current convolutional layer, each space is processed sequentially from the depth of the first layer to the depth of the last layer as follows:

[0042] The number of kernel executions M / K is determined based on the number of kernels M in the current convolutional layer and the number of kernels K in the PE array of the convolutional layer.

[0043] Each time, the K convolutional kernels of the current convolutional layer and their corresponding preset weights are scheduled to perform convolution operations on the input feature map X*(L*S)*C included in the current space. The kernels are executed M / K times until all M convolutional kernels of the current convolutional layer have been executed.

[0044] Based on further improvements to the above-mentioned device, the step of scheduling the K convolutional kernels of the current convolutional layer and their corresponding preset weights to perform convolution operations on the input feature map X*(L*S)*C included in the current space each time includes:

[0045] Perform X*Y cycles of convolution operations on the corresponding input feature map in the current space in the direction from the first row to the Xth row and from the first column to the Yth column of the convolution kernel of the current convolutional layer.

[0046] Within each cycle, based on the execution stride S of the convolutional kernel, the weight data of the preset weights of one row, one column, and depth C from K convolutional kernels are selected, multiplied and then added to the corresponding current space including the input feature map data of one row, L columns, and depth C. Each convolutional kernel obtains L intermediate results.

[0047] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0048] 1. By using the acceleration module, the appropriate size of the convolutional layer PE array is determined by utilizing the number of available DSPs on the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers. Then, the convolutional layer PE array is used to perform convolution operations on the multiple convolutional layers in the target detection module. This fully utilizes the number of DSPs on the PL end, speeds up the detection speed of the target detection module, and enables the target detection module to identify the image to be detected more quickly.

[0049] 2. By organically combining the image acquisition module, weight loading module, acceleration module, target detection module, and post-processing module, real-time detection of the target's category and location range is achieved. At the same time, when the target detection module changes, the acceleration module can promptly determine the new convolutional layer PE array, making the maintenance of the target detection module more convenient, its portability stronger, and facilitating the iterative upgrade of the target detection algorithm.

[0050] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0051] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0052] Figure 1 A schematic diagram of a real-time target detection device based on FPGA acceleration provided for an embodiment of the present invention;

[0053] Figure 2 A schematic diagram of a target detection module structure based on the YOLO V3 tiny algorithm is provided for an embodiment of the present invention;

[0054] Figure 3 This is a schematic diagram of convolution operation of a PE array of convolutional layers, provided as an embodiment of the present invention. Detailed Implementation

[0055] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0056] A specific embodiment of the present invention discloses a real-time target detection device based on FPGA acceleration, such as... Figure 1 As shown, the real-time target detection device includes an image acquisition module, a PS-end DDR, a PL-end DDR, a weight loading module, a target detection module, an acceleration module, a post-processing module, and a visualization device.

[0057] The image acquisition module is used to acquire the image to be detected and perform preprocessing to obtain a feature map to be detected in a preset format;

[0058] The PS-side DDR is used to store the feature map to be detected and the preset weights in a preset format; the preset weights are weight data pre-trained based on the target detection module.

[0059] The weight loading module is used to load the preset weights in the PS-side DDR into the PL-side DDR during the convolution operation.

[0060] The acceleration module is used to determine the size of the convolutional layer PE array based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers.

[0061] The target detection module is used to schedule the feature map to be detected stored in DDR at the PS end and the weight data stored in DDR at the PL end during the convolution operation. It uses the PE array of convolutional layers to perform convolution operations on multiple convolutional layers in the target detection module and outputs the target detection results to the DDR at the PS end for storage.

[0062] The post-processing module is used to determine the location range and category of each target in the image to be detected based on the target detection results, and send it to the visualization device for display.

[0063] Preferably, the image acquisition module includes:

[0064] Video interleaved array camera group, used to acquire raw images;

[0065] The deserializer is used to deserialize the original image and output the image to be detected.

[0066] The preprocessing module is used to preprocess the image to be detected to obtain a feature map to be detected in a preset format.

[0067] Specifically, the video interlaced array camera group is a multi-camera array structure with a spherical reference surface. In this multi-camera array structure, cameras are paired at a preset angle, so that the images captured by the video interlaced array camera group form an interlaced visual field based on the spherical array. It is worth noting that by processing the image information captured by the camera array group with cameras paired at a preset angle based on a circular reference surface, rapid and accurate perception of information around the vehicle is achieved, improving the target recognition rate and enhancing the system's versatility.

[0068] Specifically, the video interleaved array camera group can be implemented using imaging devices such as cameras, camcorders, scanners, or other devices with photo-taking capabilities (mobile phones, tablets, etc.). The original images acquired by different types of devices are deserialized using a deserializer to obtain the image to be detected.

[0069] Preferably, the preprocessing module includes one or more of the following:

[0070] The demosaic module is used to perform demosaicing on the image to be detected;

[0071] The gamma correction module is used to perform gamma correction on the image to be detected.

[0072] The telescopic module is used to perform telescopic actions on the image to be inspected;

[0073] The color space conversion module is used to perform color space conversion on the image to be detected.

[0074] Specifically, since the images to be detected acquired through different devices may not meet the input format requirements of the target detection module, preprocessing is required. Preprocessing methods can include demosaic, gamma correction, or video scaling. After preprocessing, the images to be detected will meet the input format requirements of the target detection module.

[0075] It is worth noting that by preprocessing the image to be detected using one or more of the following modules—de-mosaic, gamma correction, scaling, and color space conversion—the effectiveness of the image and the accuracy of the target detection results can be further improved.

[0076] To facilitate human visual recognition of images, Gamma correction is required. Gamma correction is a non-linear operation performed on the gray values ​​of the input image, so that the gray values ​​of the output image are exponentially related to the gray values ​​of the input image.

[0077] Specifically, the scaling module and color space conversion module can be implemented using Xilinx Video Processing SubSystem IP. Xilinx Video Processing SubSystem IP can perform various video image processing functions on the image to be detected, such as deinterlacing, video scaling (up and down scaling), color space conversion, frame rate conversion, etc., and finally obtain RGB image data with a preset format of 416*416*3.

[0078] It is worth noting that, such as Figure 2 As shown, a target detection module based on the YOLO V3 tiny algorithm is presented. This module includes 23 layers, where `layer` represents the layer level, `filters` represents the number of convolutional kernels, `size / strd` represents the kernel size / stride, `input` represents the input feature map, and `output` represents the output feature map. The input image must meet the preset format of 416*416*3, i.e., 416 rows, 416 columns, and 3 depths of feature map data.

[0079] The preset format can be the input format requirement of the target detection module. The feature map corresponding to the preprocessed image to be detected can be saved to the Double Data Rate (DDR) synchronous dynamic random access memory of the processing system (PS).

[0080] Preferably, the PS-side DDR includes a first image cache space and a second image cache space; the first image cache space is used to store multiple feature maps to be detected; each feature map to be detected is quantized by 8 bits and then reordered in the order of depth, row, and column, and each reordered feature map to be detected is mapped and stored in the second image cache space.

[0081] For example, after preprocessing, the image to be detected becomes 416*416*3 RGB format data, with each data byte being 8 bits. This data forms the AXI-Stream video stream data and is written to DDR. In DDR, each address stores one byte of data. Addresses 0x01000000, 0x01500000, and 0x02000000 are the starting addresses of the DDR addresses used for three-frame buffering, and 0x01000000, 0x01500000, and 0x02000000 are the first image buffer space. An image contains 416*416*3 = 519168 bytes of data. Storing an image into address 0x01000000 means storing the first byte of the image into 0x01000000, the second byte into 0x01000001, and so on, until all 519168 bytes are stored.

[0082] Three-frame buffering means that the image is stored in three different DDR addresses. In this embodiment of the invention, the first captured image frame is stored in DDR with the starting address 0x01000000, the second frame is stored in 0x01500000, and the third frame is stored in 0x02000000. Three-frame buffering is a reasonable design because single-frame buffering has a drawback: when the data source is continuously input, the frame buffer may store the result of two or more frames of image data superimposed, thus requiring three-frame buffering.

[0083] The acquired image data is saved to the DDR in the order of depth->width->height. That is, the first row and first column of RGB data is saved first, then the first row and second column of RGB data are saved, and so on, until the first row and 416th column are saved. Then the second row and first column are saved, and so on. In this embodiment of the invention, when the convolutional layer PE array performs convolution operations, it needs to retrieve data row by row; that is, the image needs to be reordered to the order of depth->height->width before being stored in the DDR.

[0084] The three captured images are reordered and placed into the DDR of the PS end with the starting address as 0x02500000, 0x03000000, and 0x03500000. 0x02500000, 0x03000000, and 0x03500000 are the second image buffer space.

[0085] Preferably, the real-time target detection device further includes a weight training module;

[0086] The weight training module is used to train the object detection module using the publicly available KITTI dataset and obtain the weight data after training. The weight data after training is quantized into 8 bits, and the quantized weight data is stored as preset weights in the order of depth, row and column.

[0087] The weight loading module is also used to load preset weights from the SD card into the PS-side DDR during the convolution operation.

[0088] Specifically, the pre-trained weight data of the object detection module is pre-stored in the DDR of the PS terminal. After obtaining the feature map to be detected in the preset format, the preset weights are loaded into the Programmable Logic (PL) terminal.

[0089] Understandably, the YOLO V3 tiny algorithm is trained using the publicly available KITTI dataset. The weight and bias data generated during training are separated into two binary files, allowing for the fusion of CONV and BN layers on the network weights. The fused new convolutional layer is then quantized using 8-bit fixed-point quantization. The quantized weight and bias data are rearranged according to memory access order and saved as a bin file to an SD card. The Vitis development software reads the weight data into address 0x06000000 in the DDR memory of the Processor System (PS). All weight data in the PS DDR is then loaded into the Programmable Logic (PL) DDR memory (starting address 0x00000000), and remapped to the AXI bus address (AXI address signal -0x06000000) when reading the weight data.

[0090] Specifically, the logic resources on the PL side can be used to complete the inference process of the target detection module. This requires determining the number of available DSPs on the PL side. To increase the inference speed of the target detection module, the number of available DSPs on the PL side should be utilized as much as possible.

[0091] Specifically, in the acceleration module, the size of the convolutional layer PE array is determined based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers.

[0092] Depend on Figure 2 It can be seen that the object detection module includes 13 convolutional layers (conv), namely layers 0, 2, 4, 6, 8, 10, 12, 13, 14, 15, 18, 21 and 22. The number of convolutional kernels in each convolutional layer is 16, 32, 64...255 respectively. The input feature map formats of each convolutional layer are 416*416*3, 208*208*16, 104*104*32...26*26*256 respectively.

[0093] The size of the convolutional layer PE array is determined based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the object detection module, and the input feature maps of the multiple convolutional layers. The size of the convolutional layer PE array is C*K*L, where C is the execution depth, K is the number of execution convolutional kernels, and L is the parallelism of the execution columns.

[0094] Preferably, determining the size of the convolutional layer PE array based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers includes:

[0095] The parallelism of multiple execution depths and multiple execution columns is determined based on the input feature map size of the multiple convolutional layers included in the object detection module.

[0096] The number of execution convolution kernels is determined based on the number of convolution kernels in the multiple convolutional layers included in the object detection module;

[0097] The size of the convolutional layer PE array is determined based on the number of available DSPs at the PL end, the multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolution kernels; the size of the convolutional layer PE array is C*K*L, where C is the execution depth, K is the number of execution convolution kernels, and L is the parallelism of the execution columns.

[0098] It is understandable that, such as Figure 2 As shown, layers 2, 4, 6 and 8 of the target detection module are selected for convolution operations on the PE array of convolutional layers, and the number of available DSPs at the PL end is 1000.

[0099] The number of convolutional kernels in the 2-layer, 4-layer, 6-layer, and 8-layer object detection module are 32, 64, 128, and 256, respectively. To facilitate the calculation of convolutional kernels, multiple execution convolutional kernels of 32, 16, 8, 4, and 2 can be calculated at the same time. The number of each execution convolutional kernel can be the common divisor of the smallest number of convolutional kernels in the selected convolutional layer.

[0100] Based on the 208 columns in the 208*208*16 input feature map of layer 2, the 104 columns in the 104*104*32 input feature map of layer 4, the 52 columns in the 52*52*64 input feature map of layer 6, and the 26 columns in the 26*26*128 input feature map of layer 8, the parallelism of multiple execution columns of 13 columns and 26 columns can be selected. The parallelism of the selected multiple execution columns is the common divisor of the columns in the input feature map with the smallest number of columns.

[0101] Based on the depth of 16 in the 2-layer input feature map 208*208*16, the depth of 32 in the 4-layer input feature map 104*104*32, the depth of 64 in the 6-layer input feature map 52*52*64, and the depth of 128 in the 8-layer input feature map 26*26*128, multiple execution depths such as 32, 16, and 8 can be selected. The selected multiple execution depths are the common divisors of the depths in the feature maps with the smallest depth in each input feature map.

[0102] Preferably, determining the size of the convolutional layer PE array based on the number of available DSPs at the PL end, multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolutional kernels includes:

[0103] From multiple execution depths, multiple execution column parallelisms, and multiple execution convolution kernel numbers, select one execution depth, one execution column parallelism, and one execution convolution kernel number, respectively, such that the selected execution depth, execution column parallelism, execution convolution kernel number, and available DSP number satisfy the following conditions:

[0104] C*K*L≤2*N;

[0105] Where C*K*L represents the size of the convolutional layer PE array, N represents the number of available DSPs at the PL end, C represents the execution depth, K represents the number of execution convolutional kernels, and L represents the parallelism of the execution columns.

[0106] It is understandable that, since convolution operations utilize the logic resources of the PL side to complete the convolution operation, the selected execution depth, the parallelism of the execution columns, the number of execution convolution kernels, and the number of available DSPs must satisfy the following condition: C*K*L≤2*N.

[0107] Specifically, in order to better utilize the number of available DSPs at the PL end, one execution depth, one parallelism of one execution column, and one number of execution convolution kernels can be selected from multiple execution depths, multiple execution column parallelisms, and multiple execution convolution kernel numbers. Under the premise of satisfying the condition C*K*L≤2*N, the combination that makes C*K*L the maximum value can be selected. At this time, the computing resources of the available DSPs can be fully utilized, the convolution operation of the target detection module can be completed more quickly, and the detection speed of the target detection module can be significantly improved.

[0108] Compared with existing technologies, determining the size of the convolutional layer PE array based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature map can improve the applicability of the target detection module on different hardware.

[0109] Specifically, the target detection module is used to schedule the feature map to be detected stored in DDR at the PS end and the weight data stored in DDR at the PL end during the convolution operation, and to perform convolution operation on multiple convolutional layers in the target detection module using the convolutional layer PE array, and output the target detection result to the DDR at the PS end for storage.

[0110] Specifically, the convolutional layer PE array is a configuration method for performing convolution operations in the object detection module. When this convolutional layer PE array is executed, each execution requires selecting L columns of input feature maps with a depth of C and selecting K convolution kernels to participate in the convolution operation.

[0111] The PE array of convolutional layers is used to perform convolution operations on multiple convolutional layers in the object detection module. For example, one can select... Figure 2Layers 2, 4, and 6 of the 12 convolutional layers in the module participate in convolution operations. Other layers in the object detection module can be processed using different methods. After the object detection module completes inference, the completed object detection results are obtained.

[0112] Understandably, when performing inference on the target detection module, it is necessary to schedule the feature map to be detected and the corresponding weight data with preset weights.

[0113] Specifically, the post-processing module is used to determine the location range and category of each target in the image to be detected based on the target detection results, and send it to the visualization device for display.

[0114] Specifically, the object detection module outputs the location and category of objects in an image. Based on the detection results, it outlines the objects on the image with bounding boxes and displays the object category information near the bounding boxes. After this processing, the image is sent to a visualization device, where the outlined objects and their categories are displayed for user viewing. It is understandable that during autonomous driving, the vehicle's onboard system can determine the next action based on the object detection results.

[0115] During implementation, a target detection module is pre-selected, and the weight data trained by this module is stored in the DDR of the PS terminal. The convolutional layers in the target detection module that need to be operated on by the convolutional layer PE array provided according to this embodiment of the invention are determined. During actual target detection, the target detection module performs inference in real time based on the received image to be detected, and obtains the target detection result.

[0116] Compared with existing technologies, the FPGA-accelerated real-time target detection device provided in this embodiment of the invention can perform convolution operations on the convolutional layers of the target detection module in the prior art based on the convolutional layer PE array, making full use of the number of available DSPs on the PL end and improving the detection speed of the target detection module. Meanwhile, the convolutional layer PE array determined by the target detection module participates in multiple convolutional layer operations of the target detection module, making the target detection module easier to maintain, more portable, and facilitating iterative upgrades of the target detection algorithm.

[0117] Furthermore, the step of using a convolutional layer PE array to perform convolution operations on multiple convolutional layers in the target detection module includes:

[0118] For each of the plurality of convolutional layers, the input feature map H1*W1*P1 of the current convolutional layer is divided into columns according to the parallelism L of the execution columns and the execution stride S of the convolutional kernel of the current convolutional layer. There are 1 region, and the size of the input feature map in each region is H1*(L*S)*P1;

[0119] Perform the following steps on each region sequentially, from the first column to the last column of the input feature map of the current convolutional layer:

[0120] Based on the row X of the current convolutional kernel and the stride S of the current convolutional kernel, the input feature map H1*(L*S)*P1 of the current region is divided into... Each sub-region contains an input feature map of size X*(L*S)*P1; the kernel size of the current convolutional layer is X*Y*P1, where X is the row, Y is the column, and P1 is the depth.

[0121] Convolution operations are performed on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer.

[0122] Specifically, such as Figure 3 As shown, the convolutional layer PE array C*K*L is 13*16*4, the input feature map of the current convolutional layer is 416*416*8, the output feature map is 416*416*32, the number of convolutional kernels of the current convolutional layer is 32, the kernel size is 3*3*8, and the execution stride of the kernel of the current convolutional layer is 1.

[0123] Based on the parallelism of the execution column (13) and the execution stride of the current convolutional layer kernel (1), the input feature map of the current convolutional layer (416*416*8) is divided into 32 regions, each region containing an input feature map of size 416*13*8.

[0124] For each of the 32 regions, based on the 3 rows of the current convolution kernel and the execution stride of the current convolutional layer's kernel of 1, the input feature map of the current region, which is 416*13*8, is divided into 416 sub-regions, each of which includes an input feature map of size 3*13*8.

[0125] Convolution operations are performed on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer.

[0126] Preferably, the step of performing convolution operations on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer includes:

[0127] Based on the execution depth C of the convolutional layer PE array, the current sub-region X*(L*S)*P1 is divided into P1 / C spaces, and the input feature map size of each space is X*(L*S)*C.

[0128] Based on the input feature map of the current convolutional layer, each space is processed sequentially from the depth of the first layer to the depth of the last layer as follows:

[0129] The number of kernel executions M / K is determined based on the number of kernels M in the current convolutional layer and the number of kernels K in the PE array of the convolutional layer.

[0130] Each time, the K convolutional kernels of the current convolutional layer and their corresponding preset weights are scheduled to perform convolution operations on the input feature map X*(L*S)*C included in the current space. The kernels are executed M / K times until all M convolutional kernels of the current convolutional layer have been executed.

[0131] Specifically, within each sub-region, the input feature map size is 3*13*8. Based on the execution depth of the convolutional layer PE array of 4, the current sub-region 3*13*8 is divided into 2 spaces, each space containing an input feature map size of 3*13*4. These 2 spaces are processed sequentially according to the direction from the first layer depth to the last layer depth of the current convolutional layer's input feature map.

[0132] Since the current convolutional layer has 32 convolutional kernels and the PE array of the convolutional layer has 16 execution kernels, the number of kernel executions is 2. Each time, only 16 convolutional kernels of the current convolutional layer and their corresponding preset weights can be scheduled to perform convolution operations on the current space 3*13*4. Execution is performed 2 times, which can complete the execution of all 32 convolutional kernels of the current convolutional layer.

[0133] Preferably, the step of scheduling the K convolutional kernels of the current convolutional layer and their corresponding preset weights to perform convolution operations on the input feature map X*(L*S)*C in the current space each time includes:

[0134] Perform X*Y cycles of convolution operations on the corresponding input feature map in the current space in the direction from the first row to the Xth row and from the first column to the Yth column of the convolution kernel of the current convolutional layer.

[0135] Within each cycle, based on the execution stride of the convolutional kernel, the weight data of the preset weights of one row, one column, and depth C from the K convolutional kernels are selected, multiplied by the corresponding current space including the input feature map data of one row, L columns, and depth C, and then summed. Each convolutional kernel obtains L intermediate results.

[0136] Specifically, such as Figure 3 As shown, the corresponding input feature map in the current space is subjected to 9 cycles of convolution operations in the direction of the first row to the third row and the first column to the third column of the convolution kernel of the current convolutional layer.

[0137] Within each cycle, the weight data of the preset weights of one row, one column, and 4 depth from the 16 convolutional kernels are multiplied and summed with the feature map data of one row, 13 columns, and 4 depth in the current space of 3*13*4. Each convolutional kernel obtains 13 intermediate results.

[0138] For example:

[0139] In the first cycle, the weight data of the first row, first column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0140] In the second cycle, the weight data of the second row, first column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the second row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0141] In the third cycle, the weight data of the third row, first column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the third row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0142] In the fourth cycle, the weight data of the first row, second column, and depth 4 of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and depth 4 of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0143] In the 5th cycle, the weight data of the second row, second column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0144] In the 6th cycle, the weight data of the third row, second column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0145] In the 7th cycle, the weight data of the first row, third column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0146] In the 8th cycle, the weight data of the second row, third column, and 4 depth of the current 16 convolutional kernels are multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0147] In the 9th cycle, the weight data of the third row, third column, and 4 depth of the current 16 convolutional kernels are selected and multiplied with the feature map data of the first row, 13 columns, and 4 depth of the current space. This process requires 13×16×4 operations. Then, the corresponding values ​​of each convolutional kernel are added together to obtain 13 intermediate results of the 16 convolutional kernels.

[0148] It is worth noting that the 13 columns in the first 3 periods are columns 1-13, the 13 columns in the fourth to sixth periods are columns 2-14, and the 13 columns in the seventh to ninth periods are columns 3-15.

[0149] Compared with existing technologies, the real-time target detection device based on FPGA acceleration provided in this invention, through an acceleration module, utilizes the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers to determine the appropriate size of the convolutional layer PE array. Then, the convolutional layer PE array is used to perform convolution operations on the multiple convolutional layers in the target detection module, making full use of the number of DSPs at the PL end, accelerating the detection speed of the target detection module, and enabling the target detection module to identify the image to be detected more quickly. Through the organic combination of the image acquisition module, weight loading module, acceleration module, target detection module, and post-processing module, real-time detection of the target category and location range is achieved. At the same time, when the target detection module changes, the acceleration module can also determine the new convolutional layer PE array in a timely manner, making the maintenance of the target detection module more convenient, the portability stronger, and facilitating the iterative upgrade of the target detection algorithm.

[0150] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0151] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A real-time target detection device based on FPGA acceleration, characterized in that, The real-time target detection device includes an image acquisition module, a PS-end DDR, a PL-end DDR, a weight loading module, a target detection module, an acceleration module, a post-processing module, and a visualization device; The image acquisition module is used to acquire the image to be detected and perform preprocessing to obtain a feature map to be detected in a preset format; PS-side DDR is used to store the feature map to be detected and the preset weights in a preset format; The preset weights are based on the weight data pre-trained by the object detection module; The weight loading module is used to load the preset weights in the PS-side DDR into the PL-side DDR during the convolution operation. The acceleration module is used to determine the size of the convolutional layer PE array based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers. The target detection module is used to schedule the feature map to be detected stored in DDR at the PS end and the weight data stored in DDR at the PL end during the convolution operation. It uses the PE array of convolutional layers to perform convolution operations on multiple convolutional layers in the target detection module and outputs the target detection results to the DDR at the PS end for storage. The parallelism of multiple execution depths and multiple execution columns is determined based on the input feature map sizes of the multiple convolutional layers included in the object detection module; the size of the convolutional layer PE array is C. K L and C represent the execution depth, K represents the number of execution convolution kernels, and L represents the parallelism of the execution columns; The post-processing module is used to determine the location range and category of each target in the image to be detected based on the target detection results, and send it to the visualization device for display; The image acquisition module includes: Video interleaved array camera group, used to acquire raw images; The deserializer is used to deserialize the original image and output the image to be detected. The method of using a PE array of convolutional layers to perform convolution operations on multiple convolutional layers in the target detection module includes: For each of the plurality of convolutional layers, the input feature map H1 of the current convolutional layer is processed according to the parallelism L of the execution column and the execution stride S of the convolutional kernel of the current convolutional layer. W1 P1 is divided into columns. W1 / (L) S) There are 10 regions, each containing an input feature map of size H1. (L) S) P1, H1 are the height, and W1 is the width; Perform the following steps on each region sequentially, from the first column to the last column of the input feature map of the current convolutional layer: Based on the row X of the current convolutional kernel and the stride S of the current convolutional kernel, the input feature map H1 of the current region is... (L) S) P1 is divided into H1 / S There are X sub-regions, each containing an input feature map of size X. (L) S) P1; The kernel size of the current convolutional layer is X. Y P1, where X is the row, Y is the column, and P1 is the depth; Convolution operations are performed on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer.

2. The real-time target detection device according to claim 1, characterized in that, The image acquisition module also includes: The preprocessing module is used to preprocess the image to be detected to obtain a feature map to be detected in a preset format.

3. The real-time target detection device according to claim 2, characterized in that, The preprocessing module includes one or more of the following: The demosaic module is used to perform demosaicing on the image to be detected; The gamma correction module is used to perform gamma correction on the image to be detected. The telescopic module is used to perform telescopic actions on the image to be inspected; The color space conversion module is used to perform color space conversion on the image to be detected.

4. The real-time target detection device according to claim 1, characterized in that, The PS-side DDR includes a first image cache space and a second image cache space; the first image cache space is used to store multiple feature maps to be detected; each feature map to be detected is quantized by 8 bits and then reordered in the order of depth, row, and column, and each reordered feature map to be detected is mapped and stored in the second image cache space.

5. The real-time target detection device according to claim 1, characterized in that, The real-time target detection device also includes a weight training module; The weight training module is used to train the object detection module using the publicly available KITTI dataset and obtain the weight data after training. After training, the weight data is quantized to 8 bits. The quantized weight data is then stored as preset weights in the order of depth, row, and column. The weight loading module is also used to load preset weights from the SD card into the PS-side DDR during the convolution operation.

6. The real-time target detection device according to any one of claims 1-5, characterized in that, The determination of the convolutional layer PE array size based on the number of available DSPs at the PL end, the number of convolutional kernels in the multiple convolutional layers included in the target detection module, and the input feature maps of the multiple convolutional layers includes: The number of execution convolution kernels is determined based on the number of convolution kernels in the multiple convolutional layers included in the object detection module; The size of the convolutional layer PE array is determined based on the number of available DSPs at the PL end, the multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolution kernels.

7. The real-time target detection device according to claim 6, characterized in that, The determination of the convolutional layer PE array size based on the number of available DSPs at the PL end, multiple execution depths, the parallelism of multiple execution columns, and the number of multiple execution convolutional kernels includes: From multiple execution depths, multiple execution column parallelisms, and multiple execution convolution kernel numbers, select one execution depth, one execution column parallelism, and one execution convolution kernel number, respectively, such that the selected execution depth, execution column parallelism, execution convolution kernel number, and available DSP number satisfy the following conditions: C K L≤2 N; Among them, C K L represents the size of the convolutional layer PE array, N represents the number of available DSPs at the PL end, C represents the execution depth, K represents the number of execution convolution kernels, and L represents the parallelism of the execution columns.

8. The real-time target detection device according to claim 1, characterized in that, The step of performing convolution operations on each sub-region sequentially from the first row to the last row of the input feature map of the current convolutional layer includes: The current sub-region X is determined based on the execution depth C of the convolutional layer PE array. (L) S) P1 is divided into P1 / C spaces, and each space contains input feature maps of size X. (L S) C; Based on the input feature map of the current convolutional layer, each space is processed sequentially from the depth of the first layer to the depth of the last layer as follows: The number of kernel executions M / K is determined based on the number of kernels M in the current convolutional layer and the number of kernels K in the PE array of the convolutional layer. Each time, the K convolutional kernels of the current convolutional layer and their corresponding preset weights are used to adjust the input feature map X in the current space. (L S) C performs convolution operations, executing the kernel M / K times, until all M convolution kernels of the current convolutional layer have been executed.

9. The real-time target detection device according to claim 8, characterized in that, Each time, the weight data of the K convolutional kernels and their corresponding preset weights of the current convolutional layer are used to adjust the input feature map X in the current space. (L S) C performs convolution operations, including: The input feature map in the current space is processed sequentially along the direction from the first row to the Xth row and the direction from the first column to the Yth column of the current convolutional layer kernel. Y cycles of convolution operation; Within each cycle, based on the execution stride S of the convolutional kernel, the weight data of the preset weights of one row, one column, and depth C from K convolutional kernels are selected, multiplied and then added to the corresponding current space including the input feature map data of one row, L columns, and depth C. Each convolutional kernel obtains L intermediate results.

Citation Information

Patent Citations

  • Winograd YOLOv2 target detection model method based on FPGA acceleration

    CN111459877A

  • MobileNet-SSD target detection device and method based on FPGA acceleration

    CN113051216A