High energy efficiency computing system for underwater image recognition CNN based on FPGA
By employing depth-first block-by-block convolutional units and segmented adaptive quantization on FPGAs, the energy efficiency of underwater image recognition CNNs is optimized, solving the problems of high memory overhead and high inference latency in existing architectures, and achieving efficient underwater image recognition.
Patent Information
- Application Number
- CN202511099542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-07
AI Technical Summary
When performing underwater image recognition on FPGAs, existing CNN architectures suffer from low energy efficiency, especially multi-CE architectures which have high memory overhead, while single-CE architectures have insufficient inference latency and throughput.
We employ depth-first block-wise convolutional units and segmented adaptive quantization, combined with weighting and dequantization units. Through depth-first and block-wise processing, we utilize multiple sets of parallel multiplication and accumulation units to process multiple data points, and optimize the bit width format of the calculation results through segmented adaptive quantization and dequantization.
It significantly reduces on-chip memory utilization and related power consumption, improves computing throughput and energy efficiency, adapts to the complexity of underwater environments, and enhances system flexibility and scalability.
Smart Images

Figure CN120598766B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, and in particular to a high-energy-efficiency CNN computing system for underwater image recognition based on FPGA. BACKGROUND
[0002] In recent years, with the rapid development of underwater camera technology, underwater image recognition has made significant progress, especially in the fields of fish detection, classification, tracking and population estimation. The application of deep learning technology in these fields has significantly improved the recognition accuracy, but the inference based on GPU generates a large amount of energy cost and memory demand, which conflicts with the strict energy efficiency and thermal flexibility requirements of underwater devices. Field programmable gate array (FPGA) has a reconfigurable hardware structure, which can greatly reduce the operating costs related to underwater visual processing. However, when executing neural networks on FPGA for underwater image recognition, special optimization techniques are needed to achieve high energy efficiency, which involves the combination of software and hardware optimization techniques.
[0003] In terms of software, quantization can reduce the bit width used to represent weights and activations, which can reduce power consumption and resource utilization in edge inference scenarios while accelerating computation. However, the complexity of underwater environments such as variable lighting, turbidity and noise leads to a significant decrease in the accuracy of underwater image quantization processing.
[0004] In terms of hardware, improving energy efficiency means ensuring computing throughput while reducing power consumption caused by computation or storage. Currently, convolutional neural network (CNN) acceleration architectures on FPGA are usually divided into multi-computing engine (CE) and single CE architectures. Multi-CE architecture utilizes pipelined computation, allowing each processing element (PE) to compute continuously, thereby improving hardware utilization. However, these architectures generate huge on-chip memory overhead when storing weights and intermediate results, resulting in increased power consumption. In contrast, single CE architecture usually adopts a tiled computation strategy, dividing convolutional layers into multiple blocks, which can effectively reduce the demand for on-chip memory, but requires frequent off-chip memory access, resulting in increased inference delay, reduced computing throughput and energy efficiency. SUMMARY
[0005] The present application provides a high-energy-efficiency CNN computing system for underwater image recognition based on FPGA, which solves the technical problem of how to improve the energy efficiency of CNN execution on FPGA for underwater image recognition.
[0006] To solve the above technical problems, the application provides a high-energy-efficiency CNN computing system for underwater image recognition based on FPGA, which is provided with a computing engine responsible for computation, a re-quantization unit for quantization processing, and a de-quantization unit for de-quantization processing; the computing engine is provided with a depth-first block-by-block convolution unit, which adopts a depth-first and block-by-block processing method and is equipped with multiple sets of parallel multiplication and accumulation units to simultaneously process multiple data points and perform convolution operation; the re-quantization unit is used for segmenting and adaptively quantizing the computation results of the computing engine to convert them into a reduced bit width format; and the de-quantization unit is used for de-quantizing the initial activation data after quantization in a manner opposite to the segmenting and adaptive quantization and then inputting the de-quantized data into the computing engine for next round of computation.
[0007] Preferably, the re-quantization unit performs segmenting and adaptive quantization, in particular:
[0008] The original floating point data X is divided into left, middle and right regions, the middle region is designated as a quantization concentrated region, and each region is independently and uniformly quantized; the optimal length, position and number of quantization levels of the quantization concentrated region are determined through a calibration process, and the remaining number of quantization levels is allocated to the left and right regions according to the proportion of the region length.
[0009] Preferably, the calibration process comprises:
[0010] An initial quantization concentrated region is established between the minimum and maximum values of the data in a calibration data set composed of randomly extracted part of the training set;
[0011] An iterative search is started, the effective data range is gradually reduced between the minimum and maximum values of the data, the length and quantization level density of the quantization concentrated region are gradually increased, and the quantization concentrated region is slid within the effective data range to search for the segment position that best matches the current data distribution;
[0012] The quantization concentrated region with the minimum reconstruction error between quantization and de-quantization is taken as the final quantization concentrated region.
[0013] Preferably, for the left region, the quantized data , represents the left floating point boundary of the left region, represents the scaling factor of the left region, is a rounding function, and min represents the minimum value;
[0014] For the middle region, the quantized data , represents the left floating point boundary of the middle region, represents the scaling factor of the middle region, represents the number of quantization levels of the left region;
[0015] for the right region, the quantized data , denotes the left floating-point boundary of the right region, denotes the scaling factor of the right region, denotes the number of quantization levels of the middle region, denotes the number of quantization bits, max denotes taking the maximum value;
[0016] The scaling factor of each region is determined by the length of the region and the number of quantization levels allocated to it.
[0017] Preferably, during the iterative search:
[0018] the number of quantization levels allocated to the region in the current quantization set , denotes the quantization level density, are the right and left floating-point boundaries of the region in each traversal case, is the updated ;
[0019] the number of quantization levels allocated to the region v , are the right and left floating-point boundaries of the region v in each traversal case, denotes the length of the region v, respectively correspond to the left region and the right region.
[0020] Preferably, for the region k, the dequantization unit performs dequantization on the data , respectively correspond to the left region, the middle region and the right region, denotes the left integer boundary of the region k, denotes the left floating-point boundary of the region k, the scaling factor of the region k , denotes the number of quantization levels of the region k, denotes whether the region k contains the boundary of the region, when , when , .
[0021] Preferably, the depth-first block-wise convolution unit performs the convolution operation in the traversal order as follows:
[0022] ① Traverse in the depth direction of the feature map, select the feature map block and the weight block along the depth direction for calculation;
[0023] ②According to the output channel direction of the convolution kernel, the next Tm convolution kernels are selected and ① is executed, Tm is a factor of M and also a factor of the output feature map dimension in the convolution kernel, M is the total number of the convolution kernels;
[0024] ③According to the width direction of the feature map, ① and ② are executed by moving SxTc each time, S represents the step length of the convolution layer, and Tc represents the factor of the width dimension in the output feature map;
[0025] ④According to the length direction of the feature map, ①, ② and ③ are executed by moving SxTr each time, and Tr represents the factor of the height dimension in the output feature map.
[0026] Preferably, the total throughput and the average computation-communication ratio under each parameter configuration are obtained by traversing the block configuration parameters (Tm, Tn, Tr, Tc) and refining the hierarchical adjustment parameters (Trr, Tcc) of each layer, so that the optimal parameter configuration meeting the requirements of the total throughput and the average computation-communication ratio is selected when the block configuration is performed, Tn is a factor of the input channel dimension in the input feature map, Trr and Tcc respectively represent the factors of the length and width dimensions in the output feature map selected when the i-th convolution layer is calculated, and have , 、 respectively represent the height and width of the output feature map when the i-th convolution layer is calculated, represents the stride of the i-th convolution layer, and min represents the minimum value.
[0027] Preferably, the total throughput GOPS and the average computation-communication ratio CTC are calculated as follows:
[0028] ,
[0029] ,
[0030] wherein, respectively represent the number of output channels and the number of input channels when the i-th convolution layer is calculated, K is the kernel size of the convolution layer, represents the maximum working frequency of the hardware, represents the clock cycle required by the i-th convolution layer, represents the computation-communication ratio of the i-th convolution layer, and L represents the number of convolution layers in the model.
[0031] Preferably, the system further comprises a memory reading unit, a memory writing unit, a control unit and a buffer for data storage;
[0032] The memory reading unit is responsible for controlling the starting address and the amount of data loaded each time and pre-processing the data loaded on the chip;
[0033] The memory write unit is responsible for controlling the start address and quantity of data write back and writing the quantized calculation result back to the DRAM;
[0034] The control unit receives the control signals sent by the CPU through the AXI bus, decodes these signals and then sends the decoded control signals to other functional units, thereby coordinating the data communication and calculation control between various components;
[0035] The buffer area includes an input buffer area, a weight buffer area and an output buffer area; the input buffer area stores the original input data, the weight buffer area is specially used to save weights and biases, and the output buffer area stores the calculation result from the calculation engine;
[0036] The calculation engine further includes a pooling unit for downsampling tasks, an up-sampling unit for up-sampling, an average pooling unit for average pooling calculation and a fully connected layer unit for fully connected layer calculation.
[0037] The FPGA-based underwater image recognition CNN high-energy efficiency calculation system provided by the application provides a general CNN accelerator architecture, which is provided with a calculation engine, a re-quantization unit for quantization processing and an inverse quantization unit for inverse quantization processing. The calculation engine is provided with a depth-first block-by-block convolution unit, adopts a depth-first and block-by-block processing method, is equipped with multiple parallel multiplication and accumulation units, simultaneously processes multiple data points, and performs convolution operation. The re-quantization unit is used for segmenting and adaptively quantizing the calculation result of the calculation engine to convert it into a reduced bit width format. The inverse quantization unit is used for inputting the quantized result into the calculation engine after inverse quantization. The segmenting and adaptive quantization not only compresses the CNN model, but also reduces the precision loss of underwater image quantization inference. By seamlessly integrating the depth-first block-by-block convolution and the designed segmenting and adaptive quantization method, the use rate of on-chip memory and the related memory resource power consumption are significantly reduced. The system also provides a model-level design space exploration method specially designed for the block convolution architecture, which can accurately evaluate the accelerator performance under different block configurations at the model level, and finally select the configuration with the highest energy efficiency for actual accelerator implementation. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a partial structure diagram of the FPGA-based underwater image recognition CNN high-energy efficiency calculation system;
[0039] Figure 2 is a principle diagram of the calibration process;
[0040] Figure 3 is the result of using MinMax uniform quantization on VGG16-Conv2d_3;
[0041] Figure 4 is the result of using MinMax uniform quantization on ResNet34-Conv2d_2;
[0042] Figure 5 is the result of using MinMax uniform quantization on YOLOv4-tiny-Conv2d_7;
[0043] Figure 6 is the result of using SAQ quantization on VGG16-Conv2d_3;
[0044] Figure 7 is the result of using SAQ quantization on ResNet34-Conv2d_2;
[0045] Figure 8 is the result of using SAQ quantization on YOLOv4-tiny-Conv2d_7;
[0046] Figure 9 is the complete structure diagram of the high energy efficiency computing system of the underwater image recognition CNN based on FPGA;
[0047] Figure 10 is the traversal sequence diagram of convolution calculation;
[0048] Figure 11 is the circuit diagram of the dequantization unit;
[0049] Figure 12 is the circuit diagram of the re-quantization unit. DETAILED DESCRIPTION
[0050] The embodiments of the present application will be described in detail below with reference to the drawings. The embodiments are presented only for the purpose of illustration and should not be understood as limiting the present application. The accompanying drawings are used for reference and illustration only and do not limit the scope of patent protection of the present application, because many changes can be made to the present application without departing from the spirit and scope of the present application.
[0051] The embodiment of the present application provides a high energy efficiency computing system of underwater image recognition CNN based on FPGA, as shown in Figure 1The structure diagram shows that it is provided with a calculation engine responsible for calculation, a weight quantization unit for quantization processing, and a dequantization unit for dequantization processing; the calculation engine is provided with a depth-first block-by-block convolution unit, which adopts a depth-first and block-by-block processing method, is equipped with multiple sets of parallel multiplication and accumulation units, simultaneously processes multiple data points, and performs convolution operation; the weight quantization unit is used for segmented adaptive (SAQ) quantization of the calculation result of the calculation engine to convert into a reduced bit width format; and the dequantization unit is used for dequantization of the quantized initial activation data in the opposite direction of the segmented adaptive quantization and input into the calculation engine for next round calculation.
[0052] The core idea of model quantization is to convert the weights or activation values represented by floating point numbers (FP32 or FP16) in the model into low-precision integers (such as INT8, INT4), thereby reducing the calculation cost and improving the calculation speed. Conversely, the dequantization of the model is to map the low-precision integers back to high-precision floating point data. The most common INT8 quantization is to map each 32-bit floating point data in the model to an 8-bit integer.
[0053] The example proposes a SAQ quantization method, which is customized for the unique data distribution characteristics of input activation in underwater environment. The SAQ quantization method divides the original floating point data X into left, middle and right regions, and independently quantizes each region. The goal is to assign the optimal number of quantization levels to each region. The method specifies the middle region as the quantization central region, and determines the optimal length, position and number of quantization levels of this region from a global perspective during the calibration process. It is worth noting that when the boundary of the quantization central region coincides with the boundary of the adjacent segment on either side, the data is effectively divided into two segments.
[0054] For the left region, the quantized data , represents the left floating point boundary of the left region, represents the scaling factor of the left region, is a rounding function, and min represents the minimum value.
[0055] For the middle region, the quantized data , represents the left floating point boundary of the middle region, represents the scaling factor of the middle region, represents the number of quantization levels of the left region.
[0056] For the right region, the quantized data , represents the left floating point boundary of the right region, represents the scaling factor of the right region, the number of quantization levels in the middle region, b denotes the number of quantization bits, and max denotes the maximum value.
[0057] The above process completes the quantization of the original data, and the boundaries of each quantization region also become integers. Therefore, the boundaries of the regions after dequantization will be different from the original floating-point boundaries. The dequantization process needs to be calculated based on the integer boundaries of each region and its corresponding floating-point boundaries. In the dequantization process, the left integer boundary and the left floating-point boundary of the left region are respectively: , denotes the original left floating-point boundary of the left region. The left integer boundary and the left floating-point boundary of the middle region are respectively: . The left integer boundary and the left floating-point boundary of the right region are respectively: , . For region k, the data after dequantization is: where correspond to the left region, the middle region and the right region respectively.
[0058] In the segmented adaptive (SAQ) quantization, the segmentation regions of the input data are not set arbitrarily, but are determined automatically through a calibration process. This calibration provides stable and effective prior knowledge for the quantization process, so that the quantization parameters can better adapt to the actual distribution characteristics of the data set. Specifically, this calibration process is based on the statistical characteristics of the calibration data. It first establishes an initial quantization concentrated region between the minimum value and the maximum value of the data in the calibration data set composed of part of the samples randomly extracted from the training set, and the length of this region is set to . Initially, the quantization level density in this region is set to 1. Considering that the activations in different layers of underwater images may exhibit large fluctuations and may have long-tailed distributions, the overall data range (the left and right boundaries converge to the middle) is gradually reduced to exclude extreme noise and outliers. Subsequently, within the reduced effective data range, the length of the quantization concentrated region is gradually increased to include more data distribution characteristics. At the same time, this region slides from left to right to search for the segmentation position that best matches the current data distribution. Finally, by adjusting the quantization level density in this region, this method achieves detailed characterization of the data, especially improving the quantization accuracy of the quantization concentrated region. After calibration, this region is usually distributed in the middle part of the data range, and then the entire data range is divided into three regions: left region, middle region, and right region.
[0059] The entire calibration process is as follows: Figure 2The combinations of the boundary reduction rate, the length of the quantization concentrated region, the sliding length, and the quantization level density are explored through multi-level nested iterations. The order of the multi-level nested iterations is: 1) increasing the quantization level density in the quantization concentrated region; 2) increasing the length of the quantization concentrated region; 3) sliding the quantization concentrated region to the right to change the position; 4) reducing the overall data minimum boundary; and 5) reducing the overall data maximum boundary. In the iterations, the quantization levels and the length of the non-quantization concentrated region change with the changes of the quantization concentrated region. For each combination, the reconstruction error (measured by the mean square error) of quantization and dequantization is calculated to evaluate the quantization performance. The best parameters are selected as the parameters that minimize the reconstruction error.
[0060] During the iterative search, the number of quantization levels allocated in the quantization concentrated region is given by:
[0061] ,
[0062] wherein, denotes the quantization level density, are the right and left floating-point boundaries of the quantization concentrated region in each traversal case, and is the updated .
[0063] The remaining quantization levels are proportionally allocated to the left region and the right region according to the length of the region, and the number of quantization levels allocated to the region v is:
[0064] ,
[0065] wherein, are the right and left floating-point boundaries of the region v in each traversal case, correspond to the left region and the right region, respectively, denotes the length of the region v.
[0066] The scaling factor of each region can be determined by the length of the region and the number of quantization levels allocated to it. The scaling factor of the region k is , denotes whether the region k contains the boundary of the region, when , when , .
[0067] After the traversal is completed, the combination that determines the minimum reconstruction error is determined as the segmented quantization parameters for the formal test. This data-driven calibration method can dynamically adapt to select the best segmented region and quantization parameter configuration, providing a robust and accurate reference for subsequent quantization inference on the test set, thereby significantly improving the performance and robustness of the quantized model in responding to input activations in underwater environments.
[0068] Figure 3 , Figure 4 , Figure 5 are the results of using MinMax uniform quantization in VGG16-Conv2d_3, ResNet34-Conv2d_2, and YOLOv4-tiny-Conv2d_7 layers, respectively. Figure 6 , Figure 7 , Figure 8 are the results of using SAQ quantization in VGG16-Conv2d_3, ResNet34-Conv2d_2, and YOLOv4-tiny-Conv2d_7 layers, respectively. The test data of the three models (VGG16, ResNet34, YOLOv4-tiny) are randomly sampled from the underwater image test sets of FishNet, Wildfish, and URPC datasets, respectively. Figures 3 to 8 The results in Table 1 show that when using the traditional uniform quantization method MinMax, the value distribution characteristics of the input activations tend to introduce greater quantization errors. In contrast, the SAQ method adaptively determines the optimal segmentation boundary and quantization density, effectively minimizing the quantization error.
[0069] The following section will provide a detailed explanation of how to deploy CNN on FPGA to complete the underwater image recognition task. This section will be divided into four parts: (1) the overall architecture of the computation, (2) the depth-first, block-wise convolution computation strategy, (3) the model-level design space exploration method tailored for the block-wise convolution architecture, and (4) the integration with the SAQ method.
[0070] Figure 9 The overall hardware architecture based on FPGA is shown, which aims to optimize the throughput and energy efficiency of CNN operations on FPGA. In order to achieve high computing capacity while reducing resource consumption, the system adopts a modular design method, providing flexible scalability to adapt to various CNN models and different image recognition tasks. For example, Figure 9As shown, the hardware architecture consists of a computation engine responsible for computation, a re-quantization unit for quantization processing, a de-quantization unit, a memory read unit, a memory write unit, a control unit, and buffers for data storage (input buffer and output buffer). The computation engine includes a depth-first block-wise convolution unit, a pooling unit for down-sampling tasks, an up-sampling unit for up-sampling, an average pooling unit for average pooling computation, and a fully connected layer unit for fully connected layer computation. The depth-first block-wise convolution unit adopts a depth-first and block-wise processing method, equipped with multiple sets of parallel multiply and accumulate (MAC) units, which can process multiple data points simultaneously, thus efficiently performing convolution operations. The re-quantization unit performs piecewise adaptive quantization on the computation results of the computation engine to convert them into a reduced bit-width format, thus facilitating efficient execution of subsequent quantization processes. The de-quantization unit is responsible for de-quantizing the initial activation data to provide the computation engine with original bit-width input data for the next round of computation. The memory read unit is responsible for controlling the starting address and quantity of data loaded each time and pre-processing the data loaded onto the chip (such as performing padding), and the memory write unit is responsible for controlling the starting address and quantity of data written back and writing the re-quantized computation results back to the DRAM.
[0071] In terms of data flow management, the system utilizes high-speed Advanced eXtensible Interface (AXI) buses. These buses establish efficient communication channels between the FPGA and external memory as well as the host CPU, which is crucial for transmitting large data sets required for image recognition tasks. The AXI bus provides low-latency and high-bandwidth access capabilities, enabling smooth and high-speed data flow to meet the real-time processing needs of CNNs.
[0072] To support these processing units, the system design employs three specialized types of buffers: input buffer, weight buffer, and output buffer. The input buffer stores the original input data, while the weight buffer is dedicated to storing weights and biases. The output buffer stores the computation results from the computation engine. Notably, dual instances of the input buffer, weight buffer, and output buffer are employed to improve the overall computational efficiency of the system.
[0073] In terms of control system, the CPU sends control signals to the control unit through the AXI bus. The control unit decodes these signals and subsequently sends the decoded control signals to other functional units, thereby coordinating data communication and computation control among various components. Overall, this design enables the FPGA architecture to effectively utilize its parallel processing capabilities, making it highly suitable for underwater image recognition tasks and providing strong technical support for various practical applications.
[0074] In traditional convolution calculation, the process heavily relies on efficient management and storage of intermediate results due to the need of accumulating input channel results, connecting data along height and width dimensions, and merging in the output channel direction. This is especially true when dealing with large CNNs, where the model's weight parameters and intermediate calculation results can consume a large amount of storage resources. Storing all these data in on-chip resources of FPGA, such as Block RAM (BRAM), is usually impractical. To solve this problem, this example proposes a block computation strategy, the core of which is to decompose the computation of the entire convolution layer into smaller convolution operations that can be repeatedly executed. This strategy can achieve lower resource occupation and power consumption without sacrificing calculation accuracy, providing a feasible solution for efficient convolution computation on resource-constrained FPGA platforms. First, this example selects an input feature map block with a size of (Tr×S+K-S)×(Tc×S+K-S)×Tn from the starting position of an input feature map with a size of H×W×N (height×width×channel number), where Tr and Tc are the factors in the height and width dimensions of the output feature map, S is the step size during convolution calculation, and K is the length or width of the convolution kernel. From M convolution kernels with a size of N×K×K, this example selects the first Tm convolution kernels, where Tm is a factor of M and also a factor of the output feature map dimension in the convolution kernel. For these Tm convolution kernels, this example selects data from the first Tn channels, where Tn is a factor of the input channel dimension in the input feature map. The total size of the selected convolution kernel block is Tm×Tn×K×K. Subsequently, these data are loaded into the on-chip input buffer and weight buffer (the bias term is not considered here) through the AXI bus. The total amount of data loaded onto the chip is: (Tr×S+K-S)×(Tc×S+K-S)×Tn+Tm×Tn×K×K. After that, the data in the buffer are gradually transferred to the processing units (PEs) for multiply-accumulate (MAC) operations. The results of these calculations are temporarily stored in the output buffer (Output Buffer). The obtained output size is Tr×Tc×Tm.
[0075] After completing the initial convolution block calculation, subsequent input blocks are sequentially selected from the input feature map. The method proposed in this example traverses in a depth-first manner. Compared with other traversal strategies, this depth-first method ensures that the storage requirements of intermediate results remain unchanged throughout the entire calculation process, thereby minimizing on-chip memory occupation. During depth traversal calculation, the results of block convolution between each input data block and the corresponding depth-aligned weight block are directly accumulated to the matching position in the output buffer. This method eliminates the need for additional storage of intermediate results. The accumulated data in the output buffer form the final output, which has a dimension of Tr×Tc×Tm. Subsequently, these results are transmitted back to the DRAM through the AXI bus.
[0076] Whenever the traversal along the depth direction is completed, the next set of Tm convolution kernels with size of N x K x K needs to be selected, and the above-mentioned block convolution calculation along the depth direction is repeated until all M output channels are completely traversed. At this time, the result size written back to the DRAM through the bus is Tr x Tc x M. The result size obtained after the calculation only differs from the final desired output size (i.e. × ×M, in height and width. Therefore, the traversal process is performed along the height and width directions of the input feature map, the position of the loaded feature map block in these two directions is constantly moved, and the above-mentioned calculation is repeated until the entire H x W region is covered, and the entire convolution calculation is completed. When the next input data block is selected along the height and width directions, the convolution operation needs to perform sliding calculation, resulting in data overlap between adjacent calculations. In order to ensure the accuracy and integrity of data collection, avoid errors or omissions, the step size of the subsequent input block moving in the same direction is S x Tr or S x Tc.
[0077] As shown in Figure 10 , the overall traversal order is: ①traverse according to the depth direction of the feature map, select the feature map block and the weight block along the depth direction for calculation, ②traverse according to the output channel direction of the convolution kernel, select the next Tm convolution kernel and perform ①, ③traverse according to the width direction of the feature map, move S x Tc each time and perform ①②, ④traverse according to the length direction of the feature map, move S x Tr each time and perform ①②③.
[0078] Using the above method, all the weights and activation data that need to be loaded onto the chip are effectively reduced to fixed-size blocks. In addition, intermediate results only need to be accumulated in the same BRAM memory area, without the need for continuous allocation of new memory space, thereby maximizing data reuse and significantly reducing memory and computing pressure on the chip. Thanks to the support of AXI bus, data can be transmitted quickly, further optimizing memory usage and computing performance. In order to alleviate the effective delay introduced by data exchange between each computing engine and off-chip memory through the AXI interface, additional buffers are introduced to realize alternating data transmission with existing buffers. This configuration allows one buffer to manage on-chip and off-chip data interaction, while the other buffer simultaneously transmits data to the PE array. This strategy preserves the high computing efficiency of processing elements (PEs) while effectively hiding memory communication delays. Therefore, this method improves the overall throughput of the computing circuit. This design not only improves the efficiency of the system, but also increases the scalability and flexibility of the entire platform, making the hardware more adaptable when processing large-scale data. By reducing the demand for on-chip storage, overall power consumption is also reduced, thereby prolonging the life of the device. Finally, this improved architecture provides greater flexibility and operability for deploying deep learning models, enabling support for more complex network structures and higher-dimensional data inputs.
[0079] Under the splicing computing architecture, achieving high energy efficiency requires a comprehensive exploration of the performance of accelerators with different block configurations. Due to different optimal blocking parameters, limiting exploration to a single layer is insufficient, which complicates the effective trade-off. Therefore, it is necessary to conduct model-level design space exploration to systematically evaluate the behavior of the accelerator throughout the network. The method of this example first specifies the on-chip memory size constraint. In the case where memory usage does not exceed the budget, different block parameter combinations (Tm, Tn, Tr, Tc) for all layers in the network are evaluated in detail. Since each buffer is implemented with double buffering, the actual memory usage can be represented as:
[0080] ,
[0081] where, , and represent the bit widths of weights, input activations, and output data, respectively. S and K are the stride and kernel size of the convolutional layer. For layers with the same kernel size but different strides, the output dimension is calculated at the maximum stride S to ensure compatibility.
[0082] Considering that the optimal block factor of Trand Tcmay vary greatly between layers, this example first unifies Tmand Tnto avoid hardware complexity and redundancy. Then, this example introduces layer-specific tuning parameters (Trr, Tcc), which represent the factors selected for the height and width dimensions of the output feature maps when computing a specific layer. To maximize memory utilization, for a convolutional layer with stride The number of clock cycles required to compute the i-th convolutional layer is approximated as:
[0083]
[0084] where is the ceiling function, respectively represent the number of output channels, input channels, height and width of the output feature maps when computing the i-th convolutional layer.
[0085] The throughput of the i-th convolutional layer is:
[0086]
[0087] where is the maximum operating frequency of the hardware.
[0088] The computation-to-communication ratio (CTC) represents the number of computation operations per unit of memory access during the computation of the current layer. The computation-to-communication ratio of the i-th convolutional layer is defined as:
[0089]
[0090] where respectively represent the number of memory accesses when loading the input, weights and writing back the output, respectively represent the memory footprint occupied when loading the input, weights and writing back the output, which are specifically:
[0091] Each layer must satisfy the constraints of the Roofline performance model:
[0092]
[0093] where is the number of DSP units in the target hardware platform, is the maximum memory transfer bandwidth of the target hardware platform. The parameter set that does not meet the roofline standard will be discarded, represents the maximum computing performance that the target hardware platform can achieve according to the actual hardware resource size at a certain working frequency.
[0094] By traversing the configuration of (Tm, Tn, Tr, Tc) and refining the local block parameters (Trr, Tcc) of each layer, this example obtains the comprehensive performance landscape of the accelerator at the model level. The total throughput and average CTC of L convolutional layers are defined as:
[0095] ,
[0096] ,
[0097] where L represents the number of convolutional layers in the model. Finally, the explored performance indicators are divided into several intervals, and representative configurations with significant CTC differences are selected in each interval for testing on the target hardware. This method can more comprehensively determine the high energy-efficient block parameter set suitable for actual implementation.
[0098] In the quantization stage, the piecewise adaptive quantization method quantizes the weight or activation data to low-bit-width integers during the model inference process. However, due to the different scaling factors between segments, these quantized integers cannot be directly used for multiplication or addition operations. Instead, a dedicated mapping circuit is needed to perform dequantization, converting the integers back to their original numerical value domain.
[0099] Figure 11 is the circuit diagram of the dequantization unit. As shown in Figure 11 , in the dequantization unit, this example implements this dequantization process by designing two circuit modules: the dequantization region identification circuit and the dequantization calculation circuit. In the dequantization region identification circuit, data is loaded onto the chip from off-chip memory by the memory read unit through the AXI bus, and then compared with the predefined integer segment boundaries. Each comparison produces two output signals, which are processed through an AND operation. The resulting signal is then input to a decoder to generate an address corresponding to a specific data segment. Subsequently, the dequantization calculation circuit uses this address to retrieve the relevant integer boundary, floating-point boundary, and scaling factor from a dedicated RAM module. Using these parameters, the circuit performs the dequantization process using two adders and a multiplier. This simplified method can efficiently and accurately dequantize data, facilitating the seamless integration of high-performance computing.
[0100] Due to the large number of MAC operations, the output data from the compute engine often experiences bit expansion, sometimes doubling or even further increasing. In order to maintain low bit-width data transmission and computation among all layers, these outputs must be re-quantized (de-quantized), i.e., remapped back to low bit-width representation. This example constructs a de-quantization unit with a similar circuit structure as the de-quantization unit, whose circuit is shown in FIG. 1. De-quantization includes a region identification circuit and a de-quantization computation circuit. Figure 12
[0101] In the de-quantization region identification circuit, the output data from the compute engine is compared with the floating-point boundary, and the comparison result is also processed by logical operations before being input to the decoder. Then, the decoder generates an address for accessing the re-quantization parameters stored in the memory. The de-quantization computation circuit obtains the parameters according to the generated address and performs the de-quantization process. The process of de-quantization also requires the floating-point boundary, the scaling factor, and the integer boundary, and then uses two adders and one multiplier to complete the de-quantization computation, but this process increases the rounding and range clipping of the intermediate results. A comparator is used to determine whether the quantized data exceeds the range of 0 to 255. The resulting signal is fed into a multiplexer to control the output of the quantized data. Finally, the re-quantized data is written back to the off-chip memory via the memory write unit through the AXI bus.
[0102] In the following, this example will be experimentally verified from the perspectives of software and hardware. In terms of software, this example compares the accuracy of SAQ with other quantization methods in multiple CNN models. In terms of hardware, this example implements several models and compares their energy efficiency with existing methods and other hardware platforms.
[0103] MinMax, EMA, OMSE, and Percentile are methods that increase the inference accuracy of quantization by setting different calibration processes on the basis of uniform quantization. MinMax takes the maximum and minimum values of the overall data as the boundary, EMA uses the method of exponential moving average to smooth the estimation of the boundary, OMSE uses exponential reduction to determine the boundary by finding the minimum reconstruction error, and Percentile uses the percentile strategy to smooth the estimation of the boundary. Although these methods enhance the calibration process to some extent and reduce the reconstruction error caused by non-uniform data distribution, Table 1 provides a detailed comparison of image recognition accuracy between SAQ and MinMax, EMA, OMSE, and Percentile. Under the same model and data set conditions, SAQ always has higher accuracy than traditional uniform quantization methods. This indicates that SAQ can effectively handle data with highly skewed or non-uniform distribution, thereby significantly reducing quantization error. While these uniform quantization methods still use uniformly distributed quantization levels, this limits their ability to effectively capture highly non-uniform data distribution.
[0104] Table 1: Comparison of SAQ with MinMax, EMA, OMSE and Percentile image recognition accuracy
[0105] ,
[0106] In Table 1, Num. Img represents the number of underwater images; Precision indicates the bit width of the weight and activation data; Accuracy is a performance indicator, which represents Top-1 accuracy for classification tasks and mAP@0.5 for object detection tasks.
[0107] In this example, the accelerator FPGA designed in this example is compared with three hardware platforms, GPU, CPU and embedded GPU (iGPU), in terms of energy efficiency and power consumption, as shown in Table 2. The experimental results in Table 2 show that the CNN accelerator designed using the method of this example achieves higher energy efficiency compared with these platforms.
[0108] Table 2: Comparison of energy efficiency of GPU, CPU, iGPU and FPGA
[0109] ,
[0110] In Table 2, GPU is NVIDIA GeForce RTX 4060 8G, CPU is Intel Core 14700KF, iGPU is embedded GPU (Jetson AGX Xavier), and FPGA is Xilinx Zynq UltraScale+MPSoC XCZU15EG.
[0111] In summary, the embodiment of the present application provides a high-energy-efficiency CNN computing system for underwater image recognition based on FPGA, which designs an underwater image recognition computing framework based on FPGA and optimizes from the software and hardware aspects. In terms of software, the embodiment proposes a segmented adaptive quantization (SAQ) method, which is customized for the input activation characteristics of different CNN layers in underwater images. This method not only reduces the memory usage in the calculation process, but also shows higher inference accuracy after quantization compared with existing quantization methods. In terms of hardware, the embodiment develops a deep-first, block-by-block convolution accelerator with a double-buffering mechanism, which reduces on-chip memory usage and related power consumption while maintaining high throughput. In addition, the embodiment designs a model-level design space exploration method to evaluate the accelerator performance under various block configurations across network layers, and finally selects the configuration with the highest energy efficiency for the actual accelerator implementation. The proposed accelerator achieves energy efficiencies of 79.63, 45.63 and 80.28 GOPS / W on VGG16, ResNet34 and YOLOv4-tiny, respectively.
[0112] The above embodiment is a preferred embodiment of the present application, but the embodiment of the present application is not limited to the above embodiment, and any change, modification, replacement, combination, simplification made without departing from the spirit and principle of the present application shall be an equivalent replacement manner and shall be included in the protection scope of the present application.
Claims
1. An FPGA-based high-energy-efficiency computing system for underwater image recognition CNN, characterized in that, The computing engine, the re-quantization unit for quantization processing, and the de-quantization unit for de-quantization processing are provided; the computing engine is provided with a depth-first block-by-block convolution unit which adopts a depth-first and block-by-block processing method, is equipped with multiple sets of parallel multiplication and accumulation units, simultaneously processes multiple data points, and performs convolution operation; the re-quantization unit is used for segment adaptive quantization of the calculation result of the computing engine to convert into a reduced bit width format; The de-quantization unit is used for de-quantization of the initial activation data after quantization and input into the computing engine for the next round of calculation; The re-quantization unit performs segment adaptive quantization, specifically: Original floating-point data X The system is divided into three regions: left, middle, and right. The middle region is designated as the quantization concentration region, and each region is subjected to independent uniform quantization. The optimal length, position, and number of quantization levels of the quantization concentration region are determined through a calibration process, and the remaining number of quantization levels are allocated to the left and right regions according to the proportion of the region length. The calibration process includes: The minimum value of data in the calibration dataset, which consists of a subset of samples randomly drawn from the training set. and maximum value Establish an initial quantization set region between them; Start iterative search, gradually narrow the effective data range between the minimum value and the maximum value of the data, gradually increase the length of the quantization concentrated area and the quantization level density, and slide the quantization concentrated area in the effective data range to search for the segment position that best matches the current data distribution; The quantization concentrated area with the minimum reconstruction error between quantization and de-quantization is taken as the final quantization concentrated area; During the iterative search: number of quantization levels assigned to the current quantization cluster , denotes the quantization level density, are the right and left floating point boundaries of the quantization cluster for each traversal case, is the updated ; quantization level number assigned to the region v , , are right and left floating point boundaries of the region v , denotes the length of the region v , corresponding to the left and right regions, respectively; area k scaling factor , Corresponding to the left, middle, and right regions respectively. Indicates the area k The number of quantization series, Indicates the area k Does it include the boundary of the region, when hour, ,when hour, .
2. The FPGA-based underwater image recognition CNN high-energy-efficiency computing system according to claim 1, characterized in that: For the left region, the quantized data , represents the left floating-point boundary of the left region, represents the scaling factor of the left region, is a rounding function, min represents taking the minimum value; For the middle region, quantized data , represents a left floating point boundary of the middle region, represents a scaling factor of the middle region, represents a number of quantization levels of the left region; For the right region, quantized data , represents the left floating point boundary of the right region, represents the scaling factor of the right region, represents the number of quantization levels of the middle region, represents the number of quantization bits, max represents taking the maximum value; The scaling factor of each region is determined by the length of the region and the number of quantization levels allocated to it.
3. The FPGA-based underwater image recognition CNN high-energy-efficiency computing system according to claim 2, characterized in that: For the region k , the inverse quantization unit performs inverse quantization on the data , represents a left integer boundary of the region k , represents a left floating point boundary of the region k .
4. The FPGA-based underwater image recognition CNN high-energy-efficiency computing system according to any one of claims 1 to 3, characterized in that, The traversal order of the convolution operation performed by the depth-first block-by-block convolution unit is: ①Traverse according to the depth direction of the feature map, select feature map blocks and weight blocks along the depth direction for calculation; ②According to the output channel direction of the convolution kernel, the next convolution kernel is selected and ① is executed, Tm Tm is a factor of the output feature map dimension in the convolution kernel, M M is the total number of convolution kernels; ③Traverse in the width direction of the feature map, and move S × Tc and perform ①②, S denotes the stride of the convolution layer, Tc denotes the factor in the width dimension of the output feature map; IV. Traversing in the length direction of the feature map, moving S X Tr and performing ①②③, Tr represents the factor in the high dimension of the output feature map.
5. The FPGA-based underwater image recognition CNN high-energy-efficiency computing system according to claim 4, characterized in that: Configure parameters by traversing the block ( Tm , Tn , Tr , Tc And refine the layer adjustment parameters for each layer. Trr , Tcc This allows us to obtain the total throughput and average computation-to-communication ratio for each parameter configuration, and thus select the optimal parameter configuration that meets the requirements of total throughput and average computation-to-communication ratio when configuring blocks. Tn The input feature map contains factors along the channel dimension. Trr , Tcc They represent the first i Each convolutional layer selects factors in the length and width dimensions of the output feature map during computation, and has , , They represent the first i The height and width of the output feature map calculated by each convolutional layer Indicates the first i The stride of each convolutional layer min This indicates taking the minimum value.
6. The FPGA-based underwater image recognition CNN high-energy efficiency computing system according to claim 5, characterized in that, Total throughput GOPS And average calculated communication ratio CTC Is calculated as follows: , , wherein, respectively represent the number of output channels, the number of input channels, i at the time of calculation of the K is the kernel size of the convolution layer, represents the maximum operating frequency of the hardware, represents the clock cycle required for the i convolution layer, represents the calculation communication ratio of the i convolution layer, L represents the number of convolution layers in the model.
7. The FPGA-based underwater image recognition CNN high-energy-efficiency computing system according to claim 6, characterized in that: The system further includes a memory reading unit, a memory writing unit, a control unit, and a buffer for data storage; The memory reading unit is responsible for controlling the starting address and quantity of data loaded each time and pre-processing the data loaded onto the chip; The memory writing unit is responsible for controlling the starting address and quantity of data written back and writing the re-quantized calculation result back to the DRAM; The control unit receives control signals sent by the CPU through the AXI bus, decodes these signals, and then sends the decoded control signals to other functional units, thereby coordinating data communication and calculation control between various components; The buffer includes an input buffer, a weight buffer, and an output buffer; the input buffer stores the original input data, the weight buffer is dedicated to saving weights and biases, and the output buffer stores the calculation result from the computing engine; The computing engine further includes a pooling unit for downsampling tasks, an up-sampling unit for up-sampling, an average pooling unit for average pooling calculation, and a fully connected layer unit for fully connected layer calculation.
Citation Information
Patent Citations
Efficient image recognition system based on embedded edge device
CN119540734A
Linear asymmetric quantization method based on bias and scale
CN120409561A