Low-coupling Transform target detection method and system based on FPGA acceleration

By constructing a loosely coupled optimized Transformer object detection network model and a multi-framework general-purpose FPGA accelerator, the problem of limited hardware resources for the Transformer object detection framework on edge devices is solved, achieving efficient hardware acceleration and loosely coupled object detection.

CN120911516AActive Publication Date: 2025-11-07SHANGHAI JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511019120.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing Transformer object detection frameworks face hardware resource constraints when deployed on edge devices, especially in their inability to effectively accelerate matrix multiplication, nonlinear layers, and convolution operators, making it difficult to achieve full forward inference acceleration.

Method used

We employ a loosely coupled optimization method based on FPGA to construct a Transformer object detection network model. Combined with the Winograd fast algorithm, we design a multi-framework general-purpose FPGA accelerator that supports convolutional neural networks and Transformer codecs. We process different types of computationally intensive operators through a unified data flow and perform high-precision approximation and data flow optimization for special functions.

Benefits of technology

It achieves a high-efficiency acceleration of the Transformer object detection algorithm, reduces computational requirements, improves inference efficiency and detection accuracy, and is suitable for the simplified model requirements in the field of edge computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911516A_ABST
    Figure CN120911516A_ABST
Patent Text Reader

Abstract

The invention provides a low-coupling Transform target detection method and system based on FPGA (Field Programmable Gate Array) acceleration. The method comprises the following steps: constructing a Transform target detection network model based on a low-coupling optimization method; and constructing a multi-framework universal FPGA accelerator based on a Winograd fast algorithm, wherein the multi-framework universal FPGA accelerator is used for simultaneously supporting a convolutional neural network used for feature extraction and a Transform codec based on a multi-head attention mechanism in the Transform target detection network model. According to the method, an efficient memory access transformation scheme is provided for the special type of operators, so that the accelerator obtains higher convolution operation performance while having high calculation efficiency for all the operators; a high-precision approximation scheme and corresponding data flow optimization are provided for various special functions contained in an attention mechanism, a difference fitting method is provided for GELU, approximation precision is improved, attention scaling operation and Softmax operator are fused, and scaling operation overhead during operation is saved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, in particular, to a low-coupling Transformer target detection method and system based on FPGA acceleration. In particular, it relates to a method for low-coupling optimization of a Transformer target detection algorithm. BACKGROUND

[0002] In recent years, deep learning technology has developed rapidly. Neural network models based on the Transformer framework have shown excellent performance in many fields, especially in the core task of target detection in the field of computer vision. By introducing the Transformer framework that can model the global dependence between features, the accuracy of target detection has been significantly improved, and the algorithm design has ushered in a new direction of development. However, as the demand for model accuracy continues to increase, the complexity, computation and size of the target detection algorithm model have all increased significantly, which poses a serious challenge to its deployment on edge devices.

[0003] Currently, the forward inference process of neural networks is usually run on hardware platforms such as graphic processing units (GPUs), central processing units (CPUs), application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Among them, GPUs have a large number of parallel computing units and high-efficiency floating-point matrix operation capabilities, with high throughput and high memory bandwidth advantages, but their power consumption is high; CPUs have strong versatility, but the number of computing units is limited, the control logic is complex, and the efficiency of neural network inference is low; ASICs can achieve performance improvement or power reduction through specialized circuit optimization, but the development cost is high, the flexibility is poor, and the cycle is long, which limits its wide application; in contrast, FPGAs have obvious advantages in power consumption, parallelism and flexibility, and the development cycle is shorter, so they have become the preferred hardware platform for implementing lightweight Transformer / CNN model inference acceleration.

[0004] The mainstream Transformer object detection framework, such as the DETR series, adopts an architecture combining a CNN backbone network and a Transformer encoder-decoder. This architecture not only improves performance but also complicates the coupling relationship between layers, making it difficult to deploy the algorithm model on resource-constrained edge devices. In addition, some existing Transformer accelerator architectures only optimize operators such as matrix multiplication, while moving non-linear operators, normalization operators, and other operations to other platforms for calculation. Such an architecture is difficult to achieve complete hardware acceleration. Therefore, the design of a hardware accelerator for a Transformer object detection framework needs to optimize general matrix multiplication (GEMM) operators while considering the hardware acceleration needs of non-linear operators, normalization operators, and convolution operators in feature extraction networks to achieve complete forward inference acceleration capabilities.

[0005] Patent document CN111459877A discloses a Winograd YOLOv2 object detection model method based on FPGA acceleration. A PYNQ board card is used, and the main control chip of the PYNQ board card includes a processing system end PS and a programmable logic end PL. The PS end caches the YOLO model and the feature map data of the image to be detected. The PL end caches the parameters of the YOLO model and the image to be detected in the on-chip RAM, deploys a YOLO accelerator with a Winograd algorithm, completes the model acceleration operation, forms a data path of the hardware accelerator, and realizes the object detection of the image to be detected. The acceleration circuit operation result can also be read out and image preprocessing and display can be performed.

[0006] However, patent document CN111459877A cannot be applied to Transformer object detection. Transformer object detection contains a large number of matrix multiplication and matrix transpose operations, as well as complex non-linear layers such as the Softmax layer and the LayerNorm layer. The FPGA accelerator in the patent document can only support convolution operations and cannot support the above-mentioned Transformer-related operations (matrix multiplication and non-linear layers). That is, the operator type supported by the patent document is relatively simple, and the convolution operator in a model such as YOLOv2, which has a small number of operator types, can be supported, but for a Transformer object detection model containing a large number of complex operators, most of the calculations cannot be supported.

[0007] In summary, from the perspective of algorithm and hardware co-design, through algorithm simplification and FPGA-based hardware accelerator design, a target detection algorithm model with high efficiency and easy-to-use deployment and its hardware implementation scheme are researched, which has important theoretical value and application prospect. SUMMARY

[0008] In view of the defects in the prior art, the purpose of the present application is to provide a low-coupling Transformer target detection method and system based on FPGA acceleration.

[0009] According to the low-coupling Transformer target detection method based on FPGA acceleration provided by the present application, the following steps are included:

[0010] Step S1: Construct a low-coupling optimization method-based Transformer target detection network model;

[0011] Step S2: Construct a multi-frame general FPGA accelerator based on the Winograd fast algorithm, which is used to simultaneously support the convolutional neural network for feature extraction in the Transformer target detection network model and the Transformer encoder-decoder based on the multi-head attention mechanism.

[0012] Preferably, the target detection network model includes a first target detection model and a second target detection model;

[0013] The first target detection model includes a feature fusion module for low-coupling optimization and a low-coupling encoder structure; the feature fusion module adopts multi-gradient feature splicing, and the convolution structure introduces a reparameterization operation, and the convolution structure is a single 3x3 convolution during inference; the low-coupling encoder structure includes a trainable DECONV4 deconvolution operator, the DECONV4 deconvolution operator is inserted into a feature recovery path, and the aggregation path is two, obtaining an optimized feature fusion module in the low-coupling encoder;

[0014] The second target detection model is a simplified model with a straight-through structure without a horizontal path, which only uses a deconvolution layer as an up-sampling path, and directly splices the output of the Transformer encoder and the result of the up-sampling path as the input of the decoder.

[0015] Preferably, the FPGA accelerator includes a configurable multi-type computation-intensive operator operation module, an input cache module, a weight cache module, a layer bias cache module, an accumulation module, a data rearrangement module, an overall master control unit, and a special function processing unit SFU for calculating nonlinear operators in the attention mechanism;

[0016] The FPGA accelerator can process the two types of computation-intensive operators, matrix multiplication and multi-type convolution, included in the Transformer target detection network model, perform memory access processing of each operator in the input cache module and the weight cache module, generate a unified data stream to the configurable computing array, maximize the utilization rate of the configurable computing unit, and also need to process different data streams based on different operator types in the output accumulation part.

[0017] Preferably, the generating the unified data stream comprises: converting the left multiplication matrix row block into an input channel of the input feature map, converting the right multiplication matrix column block into a convolution kernel input channel, converting the output matrix column direction traversal into an input feature map pixel point traversal, and converting each column in each row in the output matrix into an output channel direction.

[0018] Preferably, the configurable multi-type computation-intensive operator operation module comprises a plurality of configurable processing units CPEs.

[0019] The configurable processing unit comprises an input conversion module, a weight conversion module, and a point multiplication and output conversion module, which correspond to Winograd domain conversion of input data, Winograd domain conversion of weight data, and Winograd domain conversion of point multiplication results in different convolution types, respectively.

[0020] Preferably, the configurable processing unit forms a configurable processing unit group and a configurable computing array through parallelism of the output channel dimension and the input channel dimension.

[0021] The computing mode of the configurable computing array comprises a CONV2 mode, a GEMM / CONV1 mode, and a CONV3 mode.

[0022] Preferably, the input buffer module, the weight buffer module, and the accumulation buffer module all adopt a ping-pong buffer structure.

[0023] The input buffer involves main memory control, and correct memory operation and generation operation of a computing array configuration signal are completed based on different output pixel point parallelism.

[0024] The accumulation buffer module integrates a residual buffer, a bias buffer, a quantization and dequantization unit, an activation module containing multiple activations, an accumulation logic unit, and an accumulation control unit.

[0025] Preferably, the special function processing unit SFU comprises a ReLU part, a GeLU part, and a SiLU part.

[0026] For the ReLU part, an activation is directly realized through a selector.

[0027] For the GeLU part, a difference fitting method based on piecewise linear fitting is adopted, so that the piecewise error can reach the same error magnitude as direct fitting.

[0028] For the SiLU part, a direct fitting method is adopted for nonlinear calculation. The range of input data is decoded to determine the range to which the data belongs, and the corresponding slope and intercept are found through k-LUT and b-LUT according to the range, and then a multiplication and addition operation is performed to obtain the output.

[0029] Preferably, the special function computing unit further comprises a multi-stage pipelined Softmax computing module and a two-stage LayerNorm computing module.

[0030] In the multi-stage pipelined Softmax computing module, the scaling operation of the attention weight is fused, and the scaling operation is decomposed so that part of it becomes a shift operation and part of it is integrated into a constant operation based on a selector.

[0031] The multi-stage pipelined Softmax data flow traverses the entire sequence based on fine-grained pipelining within a single stage, and distributes multiple cache sequences to the computing modules corresponding to each stage based on coarse-grained pipelining between stages, so that the operations between different pipelining levels overlap with each other.

[0032] According to the low-coupling Transformer target detection system based on FPGA acceleration provided by the application, the system comprises:

[0033] Step M1: constructing a Transformer target detection network model based on a low-coupling optimization method;

[0034] Step M2: constructing a multi-frame general FPGA accelerator based on a Winograd fast algorithm, which is used to simultaneously support a convolutional neural network for feature extraction in the Transformer target detection network model and a Transformer encoder-decoder based on a multi-head attention mechanism.

[0035] Compared with the prior art, the application has the following beneficial effects:

[0036] 1. The FPGA accelerator based on the Winograd fast algorithm is constructed, the data flow of two major types of calculation-intensive operators, i.e., multi-type convolution and matrix multiplication, is unified in the Winograd domain, an efficient padding strategy is proposed to adapt to the same, an efficient memory access transformation scheme is proposed for special type operators, so that the accelerator has high calculation efficiency for all operators and obtains higher convolution operation performance.

[0037] 2. The application proposes a high-precision approximation scheme and corresponding data flow optimization for the multiple special functions contained in the attention mechanism, proposes a difference fitting method for GeLU to improve the approximation accuracy, and fuses the attention scaling operation and the Softmax operator to save the scaling operation overhead at runtime.

[0038] 3. The FPGA hardware accelerator of the application divides the ping-pong cache in a fine-grained manner, can realize effective ping-pong control, stores most of the accumulated data on-chip, greatly reduces the data transfer time, and conforms to the processing mode of special functions.

[0039] 4、The hardware accelerator for the Transformer target detection framework needs to optimize the general matrix multiplication (GEMM) operator while considering the hardware acceleration needs of nonlinear operators, normalization operators, and convolution operators in the feature extraction network, and has stronger inference ability. BRIEF DESCRIPTION OF DRAWINGS

[0040] Other features, objects, and advantages of the application will become more apparent from the following detailed description of non-limiting embodiments, when read in conjunction with the accompanying drawings:

[0041] Figure 1 Two low-coupling network structure diagrams based on the DETR framework;

[0042] Figure 2 The overall structure diagram of the multi-framework FPGA accelerator for the DETR;

[0043] Figure 3 The data flow diagram of matrix multiplication and convolution;

[0044] Figure 4 The configurable processing unit structure diagram based on Winograd;

[0045] Figure 5 The configurable processing unit group and configurable computing array structure diagram;

[0046] Figure 6 The processing diagram for two special operators CONV3_2 and DECONV4;

[0047] Figure 7 The memory access process diagram of different types of operators;

[0048] Figure 8 The column direction padding process example diagram;

[0049] Figure 9 The multi-type activation function calculation module structure diagram;

[0050] Figure 10 The multi-stage pipelined Softmax data flow diagram. DETAILED DESCRIPTION

[0051] The application will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be noted that for those skilled in the art, without departing from the concept of the application, a number of changes and improvements can be made. These all belong to the protection scope of the application.

[0052] The application provides a hardware accelerator implementation method based on an FPGA and oriented to a DETR framework, and is developed from hardware-software co-design. Through low-coupling optimization algorithm design and in combination with a Winograd fast algorithm, the application provides an efficient and simplified Transformer target detection network model and a hardware accelerator, solves the acceleration and deployment problems of various operators in the DETR framework, and is a general hardware accelerator supporting the DETR framework and compatible with multi-framework structures, realizes FPGA acceleration of the low-coupling Transformer target detection algorithm, optimizes the convolution operation performance, reduces the operation overhead, and guarantees high detection precision and real-time detection speed. In the software design part, the application provides a low-coupling optimized Transformer target detection network model, adopts a CSP-Rep Fusion feature fusion module and a low-coupling encoder structure, reduces the calculation requirement and improves the inference efficiency. In addition, the application provides a low-coupling model with a straight-through structure, which further reduces the coupling degree while guaranteeing the detection performance. In the hardware design part, the application provides a multi-framework general FPGA accelerator based on the Winograd fast algorithm, improves the hardware acceleration performance by uniformly processing different types of convolution and matrix multiplication operators. For special functions such as GeLU, Softmax and LayerNorm, the application provides a high-precision approximation method and a multi-stage pipeline calculation module, and optimizes the operation efficiency. The FPGA accelerator adopts a ping-pong cache structure, supports memory access transformation of different operator types, optimizes the padding and memory access operations, and splits the scaling operation in the attention mechanism, thereby improving the hardware utilization and acceleration efficiency. Based on the above, the FPGA acceleration-based low-coupling Transformer target detection method can be flexibly applied to FPGA acceleration scenarios oriented to the DETR framework, and the low-coupling optimization method and structure can be used to meet the simplified model requirement in the edge computing field.

[0053] Embodiment 1

[0054] According to the FPGA acceleration-based low-coupling Transformer target detection method provided by the application, the method comprises the following steps.

[0055] Step S1: constructing a Transformer target detection network model based on a low-coupling optimization method. As shown in the following formula, the method comprises two low-coupling network structures based on the DETR framework, which are a first target detection model and a second target detection model. Figure 1 As shown in the following formula, the method comprises two low-coupling network structures based on the DETR framework, which are a first target detection model and a second target detection model. Figure 1As shown on the left side, the first target detection model includes a feature fusion module for low-coupling optimization and a low-coupling encoder structure. The CSP-Rep Fusion feature fusion module adopts multi-gradient feature splicing, can realize more effective feature aggregation, and evenly distributes the calculation amount of each layer, thereby significantly reducing the calculation requirement. Meanwhile, the convolution structure introduces a reparameterization operation, decouples the inference process from the training process, and the convolution structure is a single 3*3 convolution during inference, thereby reducing the deployment difficulty and improving the algorithm inference speed. The low-coupling encoder structure introduces a trainable DECONV4 deconvolution operator and inserts it into the feature recovery path, and simultaneously reduces the aggregation path from three to two, thereby obtaining an optimized LC-PAN feature fusion module in the low-coupling encoder, and reducing the feature aggregation to twice. The introduced learnable up-sampling operator will be helpful for recovering more accurate feature information and positioning information, and meanwhile, discarding the aggregation path of the shallow layer feature greatly reduces the input sequence length sent into the decoder, thereby further improving the inference efficiency.

[0056] The CSP-Rep Fusion feature fusion module realizes more effective feature aggregation through multi-gradient feature splicing, and evenly distributes the calculation amount, thereby significantly reducing the calculation requirement. The reparameterization operation is introduced in the convolution, decouples the inference from the training process, and is a single 3*3 convolution during inference, thereby improving the speed and reducing the deployment difficulty. The low-coupling encoder structure introduces a trainable DECONV4 deconvolution operator and simplifies the aggregation path to two, and the optimized LC-PAN feature fusion module reduces the feature aggregation times. The introduced up-sampling operator is helpful for accurately recovering the feature information, and discarding the aggregation path of the shallow layer feature further improves the inference efficiency.

[0057] In addition, as Figure 1 As shown on the right side, the present application further includes a second target detection model, which is a simplified model with a straight-through structure without a transverse path, only uses a deconvolution layer as an up-sampling path, obtains more accurate feature recovery effect, and fully considers that the low-resolution feature map after the intra-scale interaction contains more rich semantic information, proposes that on the basis of introducing the up-sampling path, by means of the sequence input characteristics of the Transformer decoder, the output of the encoder and the result of the up-sampling path are directly spliced as the input of the decoder, and a straight-through structure algorithm model with lower coupling degree and capable of guaranteeing the detection performance is obtained.

[0058] Only the deconvolution layer is used as the up-sampling path, and more accurate feature recovery is obtained. Considering that the low-resolution feature map contains rich semantic information, it is proposed that the output of the Transformer encoder and the up-sampling result are directly spliced and input into the decoder, thereby further reducing the coupling degree while guaranteeing the detection performance.

[0059] Step S2: constructing a multi-frame general FPGA accelerator based on Winograd fast algorithm, which is used to support convolutional neural network for feature extraction and Transformer encoder-decoder based on multi-head attention mechanism in the Transformer object detection network model simultaneously. As shown in Figure 2 The accelerator mainly includes configurable multi-type computation-intensive operator operation module, input cache module, weight cache module, layer bias cache module, accumulation module, data rearrangement module, overall master control unit and special function processing unit (Special Function Unit, SFU) for calculating nonlinear operators in attention mechanism, etc. The two types of computation-intensive operators, matrix multiplication and multi-type convolution, contained in the Transformer object detection network model can be processed, the memory access processing of each operator is performed in the input cache module and the weight cache module, and a unified data stream is generated to the configurable computing array, which maximizes the utilization rate of the configurable computing unit. Different data streams also need to be processed based on different operator types in the output accumulation part. In order to realize resource reuse and efficient mapping of different operators, the application unifies the data streams of the two types of operators based on the transformation process shown in Figure 3 The left multiplication matrix row block is converted into the input channel of the input feature map, the right multiplication matrix column block is converted into the input channel of the convolution kernel, the output matrix column direction traversal is converted into the input feature map pixel point traversal, and each row of the output matrix is converted into the output channel direction. Therefore, matrix multiplication can be converted to 1x1 convolution in the data arrangement stage, the matrix block process can be realized by input channel traversal and input feature map block operation, and the data stream unification of multi-type operators is also converted to the data stream unification of multi-type convolution.

[0060] In order to further realize the data stream unification of multi-type convolution, the application provides a configurable processing unit (Configurable Process Element, CPE) based on Winograd fast algorithm as shown in Figure 4 The configurable processing unit includes input conversion module, weight conversion module, point multiplication and output conversion module, which correspond to Winograd domain conversion of input data, Winograd domain conversion of weight data and Winograd domain conversion of point multiplication result in different convolution types respectively. According to the Winograd fast algorithm, the configuration conversion operation of CPE needs to input the transformation matrix B T , the convolution kernel weight transformation matrix G and the output transformation matrix A T, the input data, the weight, the output are converted to Winograd domain respectively. The input conversion module, the weight conversion module, the point multiplication and the output conversion module are realized by adding adders and selectors. In the Winograd domain, the matrix multiplication, the 1*1 convolution, the 2*2 convolution and the 3*3 convolution are unified in data flow, so that the calculation array can flexibly calculate different types of operators based on the operation mode control signal.

[0061] In addition, the application introduces the parallelism of the output channel dimension and the input channel dimension respectively, and provides a configurable processing unit group and a configurable calculation array as shown in Figure 5 . Figure 5 CPE1, CPE2...CPEpout in the figure respectively represent a plurality of configurable processing units, and the configurable units constitute a CPEGroup. The CPEGroup extracts and integrates the input data conversion modules of a plurality of CPEs into one, and the input data conversion result is broadcast shared. A single input value vector d is shared by Pout CPEs, and Pout weight data conversion modules respectively receive Pout groups of weight value vectors, thereby realizing Pout parallelism of the output channel dimension. In the actual working process, a single input value vector d corresponds to 4 input feature data of a certain row in the input channel, Pout groups of weight value vectors correspond to the weight data of the input channel positions corresponding to Pout different convolution kernels, and then the intermediate results of convolution corresponding to Pout channels on the output feature map are obtained. These channels are independent of each other, so the parallel output results in a single CPE Group do not need to be accumulated. The structure of the configurable calculation array includes a plurality of CPEGroups and an accumulator. In the figure, d1-dPin are input data of Pin channels, g11-gPinPout are calculation weights, and r0, r1...rPout respectively represent output data of Pout channels. The output value vector of a single CPE Group corresponds to 4 groups of output intermediate results, and the accumulation needs to adopt a 4-parallel flow addition tree structure, so the total number of flow addition trees in the multi-dimensional parallel calculation array is 4*Pin. The calculation mode of the configurable calculation array includes a CONV2 mode, a GEMM / CONV1 mode and a CONV3 mode.

[0062] For the two special types of operators of down-sampling convolution with a step of 2 and 4*4 deconvolution with a step of 1, the application provides targeted processing as shown in Figure 6 . Taking 3*3 convolution with a step of 2 as an example, as shown in Figure 6On the upside, the odd and even pixel points of the down-sampling convolution are split, and the split accumulation operation is performed by using the CONV2 mode and the GEMM / CONV1 mode in the configurable computing array, and the data flow unification is completed through the memory access transformation. Figure 6 On the downside, the 4x4 deconvolution process is split into 4 times of 2x2 convolution, which avoids all invalid zero value calculation and saves the zero insertion operation, and meanwhile, array multiplexing can be realized to improve the hardware efficiency.

[0063] The input buffer module, the weight buffer module and the accumulation buffer module in the FPGA hardware accelerator of the present application all adopt the ping-pong buffer structure. Figure 7 As shown, the correct memory access operation and the generation operation of the computing array configuration signal can be completed based on different output pixel parallel degrees. In addition, the input buffer makes corresponding memory optimization for padding, and can skip the whole row of zero values in units of rows. Figure 8 The accumulation buffer module integrates the residual buffer, the bias buffer, the quantization and dequantization unit, the activation module containing multiple types of activation, the accumulation logic unit and the accumulation control unit.

[0064] The FPGA hardware accelerator of the present application can calculate multiple types of operators that cannot be calculated by the comparison file by flexibly setting multiple computing modes, and in the data carrying process, the present application can support input feature maps of any size based on configuration information and perform corresponding large-size input feature map blocking, has better adaptability, and the blocking process can also completely adapt to the pipeline design of the present application.

[0065] The low-coupling Transform target detection method based on FPGA acceleration of the present application further includes a Transform special function calculation unit. Figure 9 As shown, the Transform special function calculation unit includes a ReLU part, a GeLU part and a SiLU part. For ReLU, the activation is directly realized by a selector. For GeLU, the difference fitting method based on piecewise linear fitting makes the piecewise error reach the same error magnitude as direct fitting, Figure 9The GeLU part in the formula is approximated by reusing the result of ReLU activation and calculating the difference between ReLU and the piecewise linear fitting result, and the above difference has the property of even function, and the piecewise linear fitting process realizes higher fitting accuracy by fully utilizing this feature. For the SiLU part, a direct fitting method is used for nonlinear calculation, the range decoding is performed on the input data to determine the range to which the data belongs, and the corresponding slope and intercept are looked up through k-LUT and b-LUT according to the range, and then the multiplication and addition operation is completed to obtain the output.

[0066] The Transformer special function calculation unit also includes a multi-stage pipelined Softmax calculation module and a two-stage LayerNorm calculation module. In the multi-stage pipelined Softmax calculation module, the scaling operation of the attention weight is fused, the scaling operation is decomposed, part of which becomes a hardware-friendly shift operation, and part of which is integrated into a constant operation based on a selector, thereby saving the overhead of the scaling operation in the operation process. Specifically, the scaling operation of the attention weight is decomposed into 2 s and 1 0.5 or 1.2 s which can be directly realized by a right shift operation on hardware. The latter calculation process is integrated into the approximation process of the exp index by the transformation process shown in formula (1):

[0067]

[0068] The subsequent operation of the scaling operation and the exp index solving process are completed at the same time. The definition domain of the input x is x i -max(x) corresponds to the range, which can be directly obtained by multiplying the input fixed-point number by the fixed-point constant based on the square root configuration, and the calculation process is shown in formula (2):

[0069]

[0070] Taking the Softmax calculation module as an example, the multi-stage pipelined Softmax data flow is shown in formula (3): Figure 10 The Softmax calculation process includes four stages including data loading, and the multi-stage pipelined Softmax data flow in the application traverses the entire sequence based on fine-grained pipelining within a single stage, and distributes multiple cache sequences to the calculation modules corresponding to each stage based on coarse-grained pipelining between multiple stages, so that the operations between different pipelining levels overlap with each other, thereby improving the hardware utilization rate and acceleration efficiency.

[0071] The application aims to solve the problem of software and hardware cooperation in the prior art that various operators in the DETR framework cannot be effectively accelerated and efficiently deployed, by providing an efficient and easy-to-deploy target detection algorithm model and a general hardware accelerator supporting DETR framework deployment mapping and compatible with multi-framework structure, to realize FPGA acceleration of low-coupling Transform target detection algorithm. Through low-coupling optimization algorithm design, an efficient and simplified network model is provided, which has high detection accuracy and real-time detection speed. Based on the Winograd fast algorithm, a general FPGA accelerator is provided, which improves the convolution operation performance, and an efficient data flow unified strategy is designed to solve the acceleration needs of different operators. At the same time, for the special functions in the Transformer, a high-precision approximation and data flow optimization scheme is provided to reduce the operation overhead.

[0072] Embodiment 2

[0073] The application also provides a low-coupling Transform target detection system based on FPGA acceleration, which can be realized by executing the process steps of the low-coupling Transform target detection method based on FPGA acceleration, that is, the low-coupling Transform target detection method based on FPGA acceleration can be understood by those skilled in the art as the preferred embodiment of the low-coupling Transform target detection system based on FPGA acceleration.

[0074] According to the low-coupling Transform target detection system based on FPGA acceleration provided by the application, the system comprises:

[0075] Module M1: Construct a Transform target detection network model based on a low-coupling optimization method. The target detection network model comprises a first target detection model and a second target detection model. The first target detection model comprises a feature fusion module for low-coupling optimization and a low-coupling encoder structure. The feature fusion module adopts multi-gradient feature splicing, and a convolution structure introduces a reparameterization operation, and the convolution structure is a single 3x3 convolution during inference. The low-coupling encoder structure comprises a trainable DECONV4 deconvolution operator, the DECONV4 deconvolution operator is inserted into a feature recovery path, and the aggregation path is two, to obtain an optimized feature fusion module in the low-coupling encoder. The second target detection model is a simplified model with a straight-through structure without a horizontal path, which only adopts a deconvolution layer as an up-sampling path, and directly splices the output of the Transform encoder and the result of the up-sampling path as the input of the decoder.

[0076] Module M2: a multi-frame general FPGA accelerator based on Winograd fast algorithm is constructed, which is used to support convolutional neural network for feature extraction and Transformer encoder-decoder based on multi-head attention mechanism in the Transformer object detection network model. The FPGA accelerator includes a configurable multi-type computation-intensive operator operation module, an input cache module, a weight cache module, a layer bias cache module, an accumulation module, a data rearrangement module, an overall master control unit, and a special function processing unit SFU for calculating nonlinear operators in the attention mechanism. The input cache module, the weight cache module, and the accumulation cache module all adopt the structure of ping-pong cache. The input cache involves main memory control, and based on different output pixel point parallelism, correct memory operation and generation operation of calculation array configuration signal are completed. The accumulation cache module integrates residual cache, bias cache, quantization and dequantization unit, activation module containing multiple activations, accumulation logic unit, and accumulation control unit. The FPGA accelerator can process two types of computation-intensive operators, matrix multiplication and multi-type convolution, which are contained in the Transformer object detection network model. The memory processing of each operator is performed in the input cache module and the weight cache module to generate a unified data stream to the configurable calculation array, which maximizes the utilization of the configurable calculation unit. Different data streams also need to be processed based on different operator types in the output accumulation part. The generation of the unified data stream includes: converting the left multiplication matrix row block into the input channel of the input feature map, converting the right multiplication matrix column block into the convolution kernel input channel, converting the output matrix column direction traversal into the input feature map pixel point traversal, and converting each column in the output matrix row into the output channel direction. The configurable multi-type computation-intensive operator operation module includes a plurality of configurable processing units CPE. The configurable processing unit includes an input conversion module, a weight conversion module, and a point multiplication and output conversion module, which correspond to Winograd domain conversion of input data, Winograd domain conversion of weight data, and Winograd domain conversion of point multiplication result in different convolution types, respectively. The configurable processing unit forms a configurable processing unit group and a configurable calculation array through the parallelism of the output channel dimension and the input channel dimension. The calculation mode of the configurable calculation array includes CONV2 mode, GEMM / CONV1 mode, and CONV3 mode.

[0077] The special function processing unit SFU includes a ReLU part, a GeLU part and a SiLU part. For the ReLU part, the activation is directly implemented through a selector. For the GeLU part, a difference fitting method based on piecewise linear fitting is used, so that the piecewise error can reach the same error level as direct fitting. For the SiLU part, a direct fitting method is used for nonlinear calculation. The range decoding is performed on the input data to determine the range to which the data belongs, and the corresponding slope and intercept are found through k-LUT and b-LUT according to the range, and then the multiplication and addition operation is performed to obtain the output. The special function calculation unit also includes a multi-stage pipelined Softmax calculation module and a two-stage LayerNorm calculation module. In the multi-stage pipelined Softmax calculation module, the scaling operation of the attention weight is fused, and the scaling operation is decomposed, so that part of it becomes a shift operation, and part of it is integrated into the constant operation based on the selector. In a single stage, the multi-stage pipelined Softmax data flow is based on fine-grained pipelining to traverse the entire sequence, and between multiple stages, the multi-stage pipelined Softmax data flow is based on coarse-grained pipelining to distribute multiple cache sequences to the corresponding calculation modules of each stage, so that the operations between different pipelining stages overlap with each other.

[0078] Those skilled in the art know that, in addition to implementing the system provided by the present application and each device, module and unit thereof in a pure computer readable program code manner, the same functions can also be achieved by logically programming the method steps in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers. Therefore, the system provided by the present application and each device, module and unit thereof can be considered as a hardware component, and the devices, modules and units included therein for achieving various functions can also be considered as structures within the hardware component. The devices, modules and units for achieving various functions can also be considered as both software modules for implementing methods and structures within hardware components.

[0079] The specific embodiments of the present application are described above. It should be understood that the present application is not limited to the specific embodiments described above, and various changes or modifications can be made by those skilled in the art within the scope of the claims, without affecting the essential content of the present application. The embodiments of the present application and the features in the embodiments can be combined with each other in any manner without conflict.

Claims

1. A low-coupling Transform target detection method based on FPGA acceleration, characterized in that, The application relates to a low-coupling optimization method-based Transformer target detection network model and a Winograd fast algorithm-based multi-frame general FPGA accelerator. The target detection network model comprises a first target detection model and a second target detection model. The first target detection model comprises a feature fusion module for low-coupling optimization and a low-coupling encoder structure; the feature fusion module adopts multi-gradient feature splicing, a convolution structure introduces a reparameterization operation, and the convolution structure is a single 3*3 convolution during reasoning; the low-coupling encoder structure comprises a trainable DECONV4 deconvolution operator, the DECONV4 deconvolution operator is inserted into a feature recovery path, there are two aggregation paths, and an optimized feature fusion module in the low-coupling encoder is obtained.

2. The FPGA-accelerated low-coupling Transform target detection method according to claim 1, wherein, The second target detection model is a simplified model without a transverse path and with a straight-through structure, only adopts a deconvolution layer as an up-sampling path, and directly splices the output of a Transformer encoder and the result of the up-sampling path as the input of a decoder. The object of target detection is product surface defects, sea surface ships or moving objects in a monitoring area. The FPGA accelerator comprises a configurable multi-type computation-intensive operator operation module, an input cache module, a weight cache module, a layer bias cache module, an accumulation module, a data rearrangement module, an overall master control unit and a special function processing unit SFU for computing a nonlinear operator in an attention mechanism. The FPGA accelerator can process two types of computation-intensive operators, namely matrix multiplication and multi-type convolution, which are simultaneously contained in the Transformer target detection network model, performs memory processing of each operator in the input cache module and the weight cache module, generates a unified data stream, maximizes the utilization rate of a configurable computing unit, and processes different data streams based on different operator types in the output accumulation part.

3. The FPGA-accelerated low-coupling Transform target detection method of claim 1, wherein, The generation of the unified data stream comprises the following steps: converting left-multiplying matrix rows into input channels of an input feature map, converting right-multiplying matrix columns into convolution kernel input channels, converting an output matrix column direction into input feature image pixel point traversal, and converting each row and column in the output matrix into an output channel direction. The configurable multi-type computation-intensive operator operation module comprises a plurality of configurable processing units CPEs.

4. The FPGA-accelerated low-coupling Transform target detection method of claim 3, wherein, The configurable processing unit comprises an input conversion module, a weight conversion module and a point multiplication and output conversion module, which correspond to Winograd domain conversion of input data, Winograd domain conversion of weight data and Winograd domain conversion of point multiplication results in different convolution types, respectively.

5. The FPGA-accelerated low-coupling Transform target detection method according to claim 3, wherein, The configurable processing unit CPE is configured by the parallelism of the output channel dimension and the input channel dimension to form a configurable processing unit group and a configurable computing array. ​ 6. The FPGA-accelerated low-coupling Transform target detection method according to claim 5, wherein, ​ The computing mode of the configurable computing array includes a CONV2 mode, a GEMM / CONV1 mode and a CONV3 mode.

7. The FPGA-accelerated low-coupling Transform target detection method of claim 3, wherein, The input buffer module, the weight buffer module and the accumulation buffer module all adopt a ping-pong buffer structure. The input buffer involves main memory access control, and correct memory access operation and computing array configuration signal generation operation are completed based on different output pixel point parallelism. The accumulation buffer module integrates a residual buffer, a bias buffer, a quantization and dequantization unit, an activation module containing multiple activations, an accumulation logic unit and an accumulation control unit.

8. The FPGA-accelerated low-coupling Transform target detection method of claim 3, wherein, The special function processing unit SFU includes a ReLU part, a GeLU part and a SiLU part. For the ReLU part, activation is directly realized through a selector. For the GeLU part, based on a difference fitting method of piecewise linear fitting, the piecewise error can be made to reach the same error magnitude as direct fitting. For the SiLU part, a direct fitting method is used for nonlinear calculation, the range decoding of input data is performed to determine the range to which the data belongs, and the corresponding slope and intercept are looked up through k-LUT and b-LUT according to the range, and then the multiplication and addition operation is performed to obtain the output.

9. The FPGA-accelerated low-coupling Transform target detection method of claim 8, wherein, The special function computing unit also includes a multi-stage pipelining Softmax computing module and a two-stage LayerNorm computing module. In the multi-stage pipelining Softmax computing module, scaling operation of attention weight is fused, and the scaling operation is decomposed, so that part of it becomes a shift operation and part of it is integrated into constant operation based on a selector. In a single stage, the multi-stage pipelining Softmax data flow is based on fine-grained pipelining to traverse the entire sequence, and between multiple stages, multiple buffer sequences are distributed to the corresponding computing modules of each stage based on coarse-grained pipelining, so that the operations between different pipelining stages overlap with each other.

10. A low-coupling Transform object detection system based on FPGA acceleration, characterized in that, The method comprises the following steps: Step M1: constructing a Transformer target detection network model based on a low-coupling optimization method; Step M2: constructing a multi-framework general FPGA accelerator based on a Winograd fast algorithm, which is used to simultaneously support a convolutional neural network for feature extraction in the Transformer target detection network model and a Transformer encoder-decoder based on a multi-head attention mechanism.

Citation Information

Patent Citations

  • Winograd YOLOv2 target detection model method based on FPGA acceleration

    CN111459877A

  • Target detection-oriented lightweight network hardware acceleration method and target detection method

    CN119295722A

  • Multi-head attention mechanism operator optimization method and system for domestic heterogeneous processor

    CN119690522A

  • Object detection method and apparatus, device, and storage medium

    WO2024183181A1