An image signal processor, model deployment method and system
By designing a computing engine and an on-chip partitioned image signal processor on an FPGA, and combining it with a genetic algorithm to optimize resource utilization, the problems of image quality degradation and FPGA resource limitations in traditional ISPs are solved, achieving efficient and low-power image processing.
Patent Information
- Application Number
- CN202511419951.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Traditional image signal processors are prone to error accumulation when running on ARM or x86 processors, leading to a decline in image quality. Furthermore, GPU-based methods are limited in their applicability to low-power, high-performance, and compact industrial applications, while the limited on-chip resources and lower computing frequency of FPGAs make it difficult to deploy ISP networks.
An image signal processor is implemented using an FPGA that does not require data exchange. The design includes a computing engine and an on-chip partitioning structure. Multi-layer convolution calculations are performed by dividing the input feature map into blocks. Pixel shuffling technology is used to reduce off-chip memory access, and a genetic algorithm is combined to optimize resource utilization through channel pruning.
It improves the computational efficiency of the image signal processor, reduces the number of off-chip memory accesses, maintains superior image quality, and achieves low-power, high-efficiency image processing capabilities, making it suitable for compact industrial applications.
Smart Images

Figure CN120894245B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image signal processor technology, specifically to an image signal processor, a model deployment method, and a system. Background Technology
[0002] Digital cameras are key components in various vision-based industrial scenarios. However, the inherent limitations of imaging environment and sensor technology often result in low image quality and problems such as white balance shift, image noise and color shift.
[0003] To address this, numerous researchers have developed Image Signal Processors (ISPs) and accompanying algorithms to convert RAW images to RGB images and enhance image quality. These traditional ISPs typically execute a series of sub-processing units on ARM or x86 processors, including color depigmentation, noise reduction, white balance, color space conversion, tone mapping, and color enhancement. Although these methods are widely used in modern cameras, they are prone to error accumulation in the processing pipeline, leading to a deterioration in image quality. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an image signal processor, a model deployment method and system to enhance the computational efficiency of the image signal processor while maintaining image quality.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] The first aspect of this invention discloses an image signal processor, which is implemented based on a field-programmable gate array (FPGA) that does not require data exchange. The FPGA includes at least a computing engine and an on-chip partitioned structure for data storage. The on-chip partitioned structure includes at least an input buffer, a first feature buffer, a second feature buffer, and an output buffer.
[0007] The input feature map input to the image signal processor is divided into multiple feature map blocks and cached in the input buffer;
[0008] The computing engine is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculation of the feature map block in the first feature buffer and the second feature buffer; perform pixel shuffling on the calculation result obtained after the last convolution calculation of the feature map block, and store the pixel shuffling result in the output buffer; and integrate the pixel shuffling results in the output buffer to form an output feature map.
[0009] Preferably, the computation engine includes a convolution engine, which is composed of multiple weighted dot product engine arrays, and each convolution kernel is executed by one of the weighted dot product engine arrays.
[0010] Preferably, each of the weighted dot product engine arrays consists of multiple weighted dot product engines and an input adder tree.
[0011] A second aspect of this invention discloses a model deployment method, the method comprising:
[0012] An initial population is generated, wherein each individual in the initial population corresponds to a convolutional neural network structure;
[0013] Individuals are selected from the initial population for breeding according to evaluation criteria, which include at least the number of DDR accesses and image quality.
[0014] Individuals are selected from the breeding results according to the evaluation index and subjected to multiple rounds of crossover and mutation until the iteration stopping condition is met, so as to obtain a set of convolutional neural network structures with optimal accuracy and delay tradeoff.
[0015] The convolutional neural network structure with optimal accuracy and delay tradeoff is deployed on a field-programmable gate array (FPGA) that does not require data exchange to obtain the image signal processor disclosed in the first aspect of the present invention.
[0016] The FPGA includes at least a computing engine and a full-on-chip partitioned structure for data storage, wherein the full-on-chip partitioned structure includes at least an input buffer, a first feature buffer, a second feature buffer, and an output buffer;
[0017] The input feature map input to the image signal processor is divided into multiple feature map blocks and cached in the input buffer;
[0018] The computing engine is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculation of the feature map block in the first feature buffer and the second feature buffer; perform pixel shuffling on the calculation result obtained after the last convolution calculation of the feature map block, and store the pixel shuffling result in the output buffer; and integrate the pixel shuffling results in the output buffer to form an output feature map.
[0019] Preferably, the convolutional neural network structure with optimal accuracy and latency tradeoff is deployed on a field-programmable gate array (FPGA) that does not require data exchange, to obtain the image signal processor disclosed in the first aspect of the present invention, comprising:
[0020] Using the sample dataset, the convolutional neural network structure with the optimal accuracy and latency tradeoff is trained until training is complete to obtain the neural network model;
[0021] The trained neural network model is deployed on a field-programmable gate array (FPGA) that does not require data exchange to obtain the image signal processor disclosed in the first aspect of the present invention.
[0022] Preferably, generating the initial population includes:
[0023] An initial population is generated using a predefined sparse standard random pruning filter.
[0024] Preferably, individuals are selected from the initial population for reproduction according to evaluation indicators, including:
[0025] According to the evaluation criteria, individuals were selected from the initial population for breeding using a tournament selection method.
[0026] A third aspect of this invention discloses a model deployment system, the system comprising:
[0027] A generation module is used to generate an initial population, wherein each individual in the initial population corresponds to a convolutional neural network structure;
[0028] A breeding module is used to select individuals from the initial population for breeding according to evaluation indicators, which at least include DDR access quantity and image quality.
[0029] An iterative module is used to select individuals from the breeding results according to the evaluation index for multiple rounds of crossover and mutation until the iteration stopping condition is met, so as to obtain a set of convolutional neural network structures with optimal accuracy and delay tradeoff.
[0030] A deployment module is used to deploy the convolutional neural network structure with optimal accuracy and latency tradeoff to a field-programmable gate array (FPGA) that does not require data exchange, so as to obtain the image signal processor disclosed in the first aspect of the present invention.
[0031] The FPGA includes at least a computing engine and a full-on-chip partitioned structure for data storage, wherein the full-on-chip partitioned structure includes at least an input buffer, a first feature buffer, a second feature buffer, and an output buffer;
[0032] The input feature map input to the image signal processor is divided into multiple feature map blocks and cached in the input buffer;
[0033] The computing engine is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculation of the feature map block in the first feature buffer and the second feature buffer; perform pixel shuffling on the calculation result obtained after the last convolution calculation of the feature map block, and store the pixel shuffling result in the output buffer; and integrate the pixel shuffling results in the output buffer to form an output feature map.
[0034] Preferably, the deployment module is specifically used to: train the convolutional neural network structure with optimal accuracy and latency tradeoff using a sample dataset until training is completed to obtain a neural network model; and deploy the trained neural network model to a field-programmable gate array (FPGA) that does not require data exchange to obtain the image signal processor disclosed in the first aspect of the present invention.
[0035] Preferably, the generation module is specifically used to: generate an initial population using a predefined sparsity standard random pruning filter.
[0036] Based on the above embodiments of the present invention, an image signal processor, model deployment method, and system are provided. The image signal processor is implemented on an FPGA that does not require data exchange. The FPGA includes at least a computing engine and a full-on-chip partitioned structure for data storage. The input feature map of the image signal processor is divided into multiple feature map blocks and cached in an input buffer. The computing engine is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculations on the feature map blocks in a first feature buffer and a second feature buffer; perform pixel shuffling on the calculation results obtained after the last convolution calculation on the feature map blocks, and store the pixel shuffling results in an output buffer; and integrate the pixel shuffling results in the output buffer to form an output feature map. This eliminates the number of off-chip memory accesses and improves resource utilization, enhancing the computational efficiency of the image signal processor while maintaining superior image quality. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0038] Figure 1 This is a structural block diagram of an image signal processor provided in an embodiment of the present invention;
[0039] Figure 2 An example diagram of the overall architecture of an FPGA that does not require data exchange, provided for an embodiment of the present invention;
[0040] Figure 3 This is a partial example diagram of the on-chip partitioning structure provided in an embodiment of the present invention;
[0041] Figure 4 This is an example diagram of the on-chip data flow provided in an embodiment of the present invention;
[0042] Figure 5 An example diagram of the first part of a computing engine for an ISP provided in an embodiment of the present invention;
[0043] Figure 6 Example diagram of the second part of the ISP-oriented computing engine provided in the embodiments of the present invention;
[0044] Figure 7 A flowchart of a model deployment method provided in an embodiment of the present invention;
[0045] Figure 8 This is an overall flowchart of a model deployment method provided in an embodiment of the present invention;
[0046] Figure 9 Examples of qualitative analysis of different ISP performance provided in embodiments of the present invention;
[0047] Figure 10 This is a structural block diagram of a model deployment system provided in an embodiment of the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0050] Digital cameras are key components in various vision-based industrial scenarios, such as explosion testing, PCB solder joint inspection, and quality monitoring on high-speed production lines. However, inherent limitations of the imaging environment and sensor technology often result in low image quality, leading to problems such as white balance shift, image noise, and color shift. Furthermore, images acquired by complementary metal-oxide-semiconductor (CMOS) sensors are typically in RAW format, which is incompatible with most computer vision algorithms.
[0051] To address this, numerous researchers have developed image signal processors (ISPs) and accompanying algorithms to convert RAW images to RGB images and enhance image quality. These traditional ISPs typically execute a series of sub-processing units on ARM or x86 processors, including color depigmentation, noise reduction, white balance, color space conversion, tone mapping, and color enhancement.
[0052] Demosaicing involves interpolating a single-channel raw color filter array (CFA) image into a multi-channel full-color image. A noise reduction unit removes noise and improves the signal-to-noise ratio. White balance adjusts the RGB channels to neutralize color biases from different light sources. Color correction modifies color values using a correction matrix. While these methods are widely used in modern cameras, they are prone to error accumulation in the processing pipeline, leading to image quality degradation. Traditional ISP sub-processing units are typically hand-designed for specific tasks and applications, thus requiring significant manual debugging and optimization when facing various image processing tasks.
[0053] With the development of deep learning (DL) and the advancement of computer hardware, an increasing number of methods have been developed for image signal processing based on GPU platforms. These methods are able to learn hidden statistical information from data and jointly solve multiple tasks, achieving significant success in improving image quality and reducing human intervention. However, GPU platforms typically consume a lot of power and occupy a considerable amount of space, which limits their applicability in industrial applications that require low power consumption, high performance, and compact size, such as high-speed cameras and drones.
[0054] Research has shown that Field-Programmable Gate Arrays (FPGAs), with their high throughput, low latency, and low power consumption, are widely used in industrial applications, and many researchers have made outstanding contributions in this field. Increasingly, research is focusing on using FPGAs to accelerate CNNs (Convolutional Neural Networks) due to their high parallelism and low power consumption.
[0055] For example, FPGA-based neural networks are used to estimate the degradation of organic light-emitting diodes (OLEDs); optimal control of cascaded H-bridge multilevel converters is achieved by deploying lightweight neural networks on FPGAs; and FPGAs are used to accelerate CNNs for plant disease identification, resulting in a low-power, high-precision, and fast plant disease identification terminal, which is well-suited for real-time plant disease identification in the field. A lightweight dual-stream defect detection network, accelerated using FPGAs, achieves competitive accuracy and speed in photovoltaic cell defect detection. Overall, applying deep learning-based FPGA methods in industry can not only achieve lower power consumption but also obtain better quality.
[0056] FPGA-based accelerators have emerged as a promising alternative for hardware applications implementing CNNs on edge devices. These devices stand out due to their compact size, low power consumption, and superior parallel computing capabilities, making them well-suited for energy-efficient parallel computing. However, the limited on-chip resources and relatively low computing frequency of FPGAs pose significant challenges to architecture design, as follows:
[0057] 1) Units that implement complex ISP network operations (such as convolution and pixel shuffling) on FPGAs are typically inefficient and difficult to design;
[0058] 2) On-chip memory capacity is usually insufficient to store all intermediate feature maps, requiring frequent access to off-chip memory, which can lead to computational stagnation;
[0059] 3) ISP networks typically have high computational and memory access requirements, making them difficult to deploy on resource-constrained FPGAs.
[0060] To address these challenges, this solution proposes an image signal processor, a model deployment method, and a system. The image signal processor is implemented based on an FPGA that eliminates the need for data exchange. The FPGA includes at least a computing engine and a fully on-chip partitioned structure for data storage, eliminating the number of off-chip memory accesses and improving resource utilization. This enhances the computational efficiency of the image signal processor while maintaining superior image quality.
[0061] Overall, this solution proposes a high-speed, low-power ISP based on FPGA and deep learning, which features high inference speed and low power consumption. To fully utilize the parallelism of the FPGA and reduce the number of off-chip memory accesses, a data-exchange-free FPGA accelerator (called EFA) is designed. This FPGA includes at least a computing engine for the ISP and a fully on-chip partitioned structure for data storage.
[0062] The computation engine developed a weighted dot product engine (also known as the weighted dot product engine or WDE for short) and a pixel shuffle (also known as pixel recombination) fused with the output buffer, thereby improving the computational efficiency of convolution, activation and pixel shuffle.
[0063] It should be noted that the on-chip partitioning structure in this solution refers to allocating a portion of the FPGA resources for data storage. This portion of resources is entirely on-chip within the FPGA chip and does not occupy off-chip storage resources.
[0064] This solution's on-chip partitioning structure divides the input feature map into several blocks based on the size of the on-chip memory, thereby reducing off-chip memory access. Furthermore, to reduce the computational cost and memory footprint of the ISP network in the FPGA, this solution also presents an FPGA-oriented channel pruning method (FCP). When deploying the ISP network (i.e., the subsequently deployed neural network model) on the FPGA based on FCP, the number of off-chip memory accesses is analyzed, and channel pruning based on a genetic algorithm is used to reduce redundant filters in the ISP network, thereby reducing the computational and memory requirements of deep learning-based methods and improving the efficiency of vision-based industrial systems.
[0065] The following describes the solution in detail through various embodiments.
[0066] See Figure 1 The diagram shows a structural block diagram of an image signal processor provided by an embodiment of the present invention. The image signal processor is implemented based on an FPGA that does not require data exchange. The FPGA includes at least a computing engine 100 and an on-chip partition structure 200 for data storage. The on-chip partition structure 200 includes at least an input buffer, a first feature buffer, a second feature buffer, and an output buffer.
[0067] Specifically, the input feature map of the input image signal processor is divided into multiple feature map blocks and cached in the input buffer.
[0068] The computation engine 100 (hereinafter referred to as the computation engine) is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculation of the feature map block in the first feature buffer and the second feature buffer; shuffle the calculation results obtained after the last convolution calculation of the feature map block into pixels, and store (cache) the pixel shuffling results in the output buffer; and integrate the pixel shuffling results in the output buffer to form the output feature map (RGB image).
[0069] In some specific embodiments, the computation engine includes a convolution engine, which consists of multiple arrays of weighted dot engines (WDEs), with each convolution kernel executed by one array of weighted dot engines (i.e., a WDE array). Each array of weighted dot engines consists of multiple weighted dot engines and an input adder tree.
[0070] As can be seen from the above, an FPGA that does not require data exchange (i.e., EFA) mainly consists of two parts: an on-chip partitioned structure and a computing engine. The following sections will provide a detailed explanation of these two parts.
[0071] Before detailing the on-chip partitioning structure and computing engine, let's first briefly explain the overall architecture of an FPGA (EFA) that does not require data exchange.
[0072] like Figure 2 As shown in the example diagram of the overall architecture of the FPGA without data exchange, the processing system (PS) shifts the RAW image captured by the CMOS sensor into a 4-channel matrix (example only), partitions it according to the on-chip partitioning structure, and stores it in DDR (double data rate). The PS then transfers the input blocks (feature map blocks) to the on-chip input buffer (i.e., Figure 2 The input buffer in the computation engine performs inference (including convolution calculations and pixel shuffling) on the ISP network (i.e., the subsequently deployed neural network model). The computation results (intermediate results) obtained during convolution calculations are temporarily stored in the first or second feature buffer until the final layer is processed, at which point they are reloaded into the computation engine. Afterward, the PS writes the result block back to DDR. The first and second feature buffers are then used for this purpose. Figure 2 Feature cache in the middle.
[0073] in, Figure 2 In this context, MUX stands for multiplexer, DEMUX for demultiplexer, MAC (MultiplyAccumulate) for multiply-accumulate, and FIFO (First In First Out) for first-in-first-out.
[0074] (1) Explanation of the partition structure on the whole chip:
[0075] It should be noted that for FPGAs with limited on-chip memory, fully storing the input image and feature map requires frequent access to external memory, which interrupts the continuous operation of the computing engine and cannot fully utilize the parallelism of the FPGA.
[0076] Therefore, this solution divides the input feature map into multiple equally sized feature map blocks to reduce the consumption of the on-chip input buffer on the FPGA. For example... Figure 3As shown in the partial example diagram of the full-chip partitioning structure, the input feature map is divided into parts of size r×f. x ×n feature patches, where f x and n represent the size of the input feature map and the number of input channels, respectively, and r represents the block size. By adjusting r, a feature map block can be completely contained in the input buffer.
[0077] The computation engine processes each feature map patch sequentially, using a convolution kernel of size k×k×n. As the window slides, it can generate a convolution kernel of size r×f. x ×m output blocks, where m is the number of output channels. It's important to note that to ensure the output block length is r, the input buffer length must be r+k-1.
[0078] During the processing of each feature map patch by the computing engine, after the convolution calculation of one feature map patch is completed, the next feature map patch is not processed immediately. Instead, the calculation result is re-input into the computing engine for subsequent convolution calculations until the final output is obtained.
[0079] For example Figure 4 The caching strategy shown is illustrated in the on-chip data flow example diagram during network inference. Layer 1 to Layer x-1 are convolutional layers, and Layer x is pixel shuffling. In the first convolutional layer computation, as shown... Figure 4 As shown by the red line, the feature map is read from the input buffer for convolution calculation, and the calculation results of the calculation engine are cached in buffer 0 (Buffer0, i.e., the first feature buffer).
[0080] like Figure 4 As shown by the brown lines, in the second convolutional layer calculation, the result (intermediate result) of the first convolutional layer calculation is read from buffer 0 for convolution calculation, and the result of the second convolutional layer calculation is cached in buffer 1 (Buffer1, i.e., the second feature buffer). For the third convolutional layer calculation, the result of the second convolutional layer calculation is read from buffer 1 for convolution calculation, and the result of the third convolutional layer calculation is cached in buffer 0. This process continues, performing multiple convolutional calculations for each feature map patch.
[0081] In summary, subsequent convolution calculations only involve buffer 0 and buffer 1. Therefore, this scheme uses only 3 buffers (input buffer, first feature buffer, and second feature buffer) to store all intermediate results.
[0082] Ultimately, as Figure 4As shown by the blue lines, pixel shuffling is performed by a finite state machine (FSM) after the final convolutional layer computation is completed. Specifically, after the final convolutional layer computation (Layer x-1), the FSM (Layer x) performs pixel shuffling.
[0083] It should be noted that pixel shuffling transforms the channel dimension of the CNN output features into a spatial dimension, outputting the final RGB image. Each feature map block will obtain a corresponding pixel shuffling result, and the pixel shuffling results of all feature map blocks are integrated together to form the entire output feature map.
[0084] Furthermore, the input buffer is only needed in the first convolutional computation, which allows subsequent convolutional computations and input buffer updates to be performed simultaneously. This approach mitigates the impact of off-chip access on the performance of the computing engine.
[0085] The above is a description of the partition structure on the entire chip.
[0086] (2) Explanation of the computing engine for ISPs:
[0087] like Figure 5 and Figure 6 As shown in the example diagram of the ISP-oriented computing engine, the computing engine of this scheme includes a WDE-based convolution engine (to improve convolution efficiency) and a pixel shuffling fused with the output buffer (to improve pixel shuffling performance).
[0088] For the weighted dot product engine: the dot product is an indispensable step in the convolution calculation in CNN, which can be represented by formula (1).
[0089] (1);
[0090] In formula (1), the vector This represents the input feature map within the current convolution window. The parameters represent the convolution kernel, m is the number of elements in the convolution window, and x... i For vectors The i-th element, w i For vectors The i-th element in the convolution; for example, if the convolution window is 3*3 in size, then m is 3*3=9. and Both are one-dimensional vectors with 9 elements.
[0091] like Figure 5As shown in the weighted dot product unit illustration, in the design of this scheme's weighted dot product engine, a dedicated buffer stores the weights of each network layer, with a size of c×k×k (e.g., 3×3×3). A buffer-based shift register converts the row-stream input activations to column format, allowing the activation window to be updated column-by-column in a single cycle. The updated activation values and their corresponding weights are input into a multiply-accumulate (MAC) array, which consists of 9 multipliers that perform two integer multiplications per cycle, followed by a 9-input adder tree to efficiently sum these products. A multiplexer controls the flow of parameters to the MAC array during the computation of different convolutional kernels.
[0092] also, Figure 2 The demonstration also showcases convolution operations using WDEs (i.e., convolution operations on feature map patches). First, a WDR array is constructed from n WDEs and an input adder tree. The convolution engine consists of m WDE arrays, with each convolution kernel executed by one WDE array. Activation inputs are loaded from the input buffer into the WDE arrays, along with weights. Then, the WDE arrays perform dot product operations, feeding the results into the adder tree to obtain the output features. After each WDE array operation, the output features are stored in an on-chip buffer awaiting processing in the next layer.
[0093] Regarding pixel shuffling in the fused output buffer: It should be noted that in the Deep Learning Processing Unit (DPU), the pixel shuffling layer is usually deployed on the PS, which significantly reduces the computation speed.
[0094] like Figure 6 As shown in the three-state finite state machine, the pixel shuffling design of the output buffer fusion scheme transforms a 1×12 input vector into a 3×2×2 matrix, where each 3×1×1 segment represents the resulting image pixel quantized to 8 bits.
[0095] Due to the need for line-by-line output and the spatial distribution of pixels across two image lines, this scheme uses a depth of f. x Two FIFOs with different bit widths (48 bits and 192 bits respectively) are used for buffering. The first two pixels and the remaining pixels are sequentially routed to the corresponding FIFOs. For example, Figure 6 As shown in the transformed 3×2×2 matrix, "the first two pixels" refers to the top two pixels, and "the remaining pixels" refers to the bottom two pixels.
[0096] Since the width of the AXI (Advanced eXtensible Interface) bus is 128 bits, which is less than the output bit width of the FIFO, this solution uses a three-state finite state machine to handle this mismatch.
[0097] like Figure 6 As shown, in state 0, the first 128 bits of data from the FIFO are transmitted to the AXI bus, and the remaining 64 bits are temporarily stored in a buffer. In state 1, the first 64 bits of data from the FIFO are merged with the data in the buffer and then output, while the subsequent 128 bits of data are buffered for future use. In the final state (state 2), the FIFO stops outputting, and the buffered data is completely output.
[0098] It's worth noting that data from the first FIFO is output directly through the FSM, while data from the second FIFO is output only after the first FIFO has finished outputting. The average output rate of the FSM is 192 × 3 ÷ 2 bits per cycle, while the input rate of the FIFO is 48 × 2 bits per cycle, ensuring that the FSM achieves throughput balance between the FIFO and AXI interfaces. In summary, this process efficiently outputs eight pixels from the FIFO to the DDR through three states.
[0099] The above is an explanation of the computing engine for ISPs.
[0100] To improve the performance of CNNs on FPGA hardware, this solution proposes a pruning method constrained by FPGA hardware, which combines DDR access analysis with genetic algorithm (GA) to balance model accuracy and computational efficiency while taking FPGA limitations into account.
[0101] Here is a detailed explanation of DDR access analysis:
[0102] The overall inference latency (Lat) of executing a CNN on an FPGA can be divided into computation latency (Lat). comp ) and DDR access latency (Lat DDR Although computation may pause due to operations such as off-chip memory access, the computation engine's cycle time is fixed and Lat can be calculated using formula (2). comp .
[0103] (2);
[0104] In formula (2), f x and f y T represents the length and width of the input feature map, respectively. i and T o C represents the block factor of the input channel and the block factor of the output channel, respectively. i and Co These represent the number of input and output channels, respectively. Step is the convolution stride, and k is the convolution kernel size.
[0105] DDR access latency (Lat) DDR This typically includes DDR access during image input. img The number of DDR accesses during inference D i D img Usually only related to the image size S x S y and S z Related, can be represented as D img =S x ×S y ×S z .
[0106] Therefore, the most important factor affecting latency is DDR access during inference. Based on the above information about EFA, assuming the block size is r, the number of accesses is D. i It can be expressed as formula (3).
[0107] (3);
[0108] The above is a description of DDR access analysis.
[0109] See Figure 7 This invention also provides a flowchart of a model deployment method, which includes:
[0110] Step S701: Generate the initial population.
[0111] In the specific implementation step S701, an initial population is generated through a predefined sparse standard random pruning filter. Each individual in the initial population corresponds to a convolutional neural network structure (CNN structure).
[0112] It should be noted that the pruning process first defines an initial population, which is a set of CNN structures. Each CNN result is represented as a gene sequence. Specifically, each gene in the genetic algorithm represents a CNN structure, and the difference between different genes is that they activate different filters.
[0113] Step S702: Select individuals from the initial population for breeding according to the evaluation criteria.
[0114] It should be noted that when evaluating the fitness of each individual in the initial population, an evaluation index can be used to assess the fitness (reflecting the individual's performance). This evaluation index should at least include the number of DDR accesses (i.e., D in the DDR access analysis above). iThe dual-objective evaluation of image quality (PSNR) guides the selection process toward a structure that achieves the best balance between performance and resource usage.
[0115] In the specific implementation step S702, individuals are selected from the initial population for breeding according to the evaluation indicators using a tournament selection method to obtain the corresponding breeding results. These breeding results include individuals with different evaluation indicators (different PSNR and different D). i ( ) individuals.
[0116] In the process of selecting individuals from the initial population for breeding using the tournament selection method, two individuals are randomly selected, and their fitness scores (based on PSNR and D) are used to determine their fitness scores. i Individuals with higher evaluation scores are selected. This approach facilitates the selection of CNN architectures that are more suitable for FPGA deployment.
[0117] Step S703: Select individuals from the breeding results according to the evaluation index and perform multiple rounds of crossover and mutation until the iteration stopping condition is met, so as to obtain a set of convolutional neural network structures with optimal accuracy and delay tradeoff.
[0118] It should be noted that the iteration stopping condition is: the genetic algorithm runs to the predetermined number of generations, or the performance (determined by PSNR and D) is satisfactory. i The assessment indicates that the improvement is trending towards plateauing.
[0119] In the specific implementation step S703, individuals are selected from the breeding results according to the evaluation index for multiple rounds of crossover and mutation until the iteration stopping condition is met, so as to obtain a set of Pareto-optimized convolutional neural network structures with optimal accuracy and delay tradeoff.
[0120] Specifically, after selecting individuals with better performance from the breeding results, the selected individuals are crossbred and mutated, and then individuals with better performance are selected again for crossbring and mutation. This process is repeated until the genetic algorithm reaches a predetermined number of generations or the performance improvement tends to plateau, finally resulting in a set of "pareto-optimized convolutional neural network structures with optimal accuracy and delay tradeoffs".
[0121] It should be noted that the crossover process generates offspring by combining genes from two parent architectures, potentially inheriting the advantageous traits of both. The mutation process, on the other hand, randomly alters one gene in an individual; in this scheme, this means adjusting the number of active filters in the layer, introducing diversity into the population.
[0122] Steps S701 to S703 above constitute the "channel pruning method based on genetic algorithm".
[0123] Step S704: Deploy the convolutional neural network structure with optimal accuracy and delay tradeoff onto an FPGA that does not require data exchange to obtain an image signal processor.
[0124] In the specific implementation step S704, the sample dataset is used to train a "convolutional neural network structure with optimal accuracy and delay tradeoff" until the training is completed, so as to obtain the neural network model.
[0125] For example, train a CNN architecture with optimal accuracy and latency tradeoff using the Zurich RAW to RGB (ZRR) dataset until training is complete to obtain a neural network model.
[0126] The trained neural network model is deployed to the above-described embodiments of the present invention. Figure 1 The aforementioned FPGA (EFA) that does not require data exchange, thus pertaining to the above embodiments of the present invention. Figure 1 The image signal processor (ISP) mentioned.
[0127] In practical applications, an EFA was implemented using specified software to generate a bitstream, which was then deployed on a specific FPGA model. The FPGA's PS (Power Supply) section contains a quad-core processor operating at 1333MHz, while the PL (Power Processing) section has a base clock of 100MHz. Both the PL and PS are equipped with 4GB of DDR memory, operating at 2400MHz. The PL is dedicated to implementing the accelerator, while tasks such as file reading and display are handled by the PS. When using automated FPGA design tools, this solution utilizes three DPU models with a parallelism of 4096.
[0128] For network training, the ISP network was trained on a single GPU using the ZRR dataset, which contains 47,000 images for training and 1,204 images for validation, all with a resolution of 488×488.
[0129] The proposed ISP network was trained using the Adam optimizer, with parameters β1 and β2 set to 0.9 and 0.999 respectively, and the learning rate set to 1×10⁻⁶ during training. -3 .
[0130] like Figure 8 The overall flowchart of a model deployment method is shown. From an overall perspective, this solution includes the following steps:
[0131] Step S801: Design an FPGA that does not require data exchange (EFA).
[0132] The FPGA without data exchange (EFA) includes a compute engine for the ISP and a fully on-chip partitioned structure, improving the inference speed of the ISP. EFA significantly reduces the number of off-chip memory accesses and increases inference speed. For details on the EFA, please refer to the above-described embodiments of the present invention. Figure 1 The content.
[0133] Step S802: Construct a channel pruning method for FPGA.
[0134] In the specific implementation step S802, delay analysis and DDR access analysis of the EFA structure are performed to obtain the above formulas (2) and (3). Then, formulas (2) and (3) are used for channel pruning based on genetic algorithm to reduce the amount of computation and memory access in the ISP network.
[0135] Step S803: Train the ISP network using the ZRR dataset and implement the ISP on the FPGA using EFA.
[0136] It should be noted that the execution principle of steps S801 to S803 can be found in the description of the various embodiments of the present invention above, and will not be repeated here.
[0137] To verify that the performance of the ISP proposed in this solution is superior to that of other ISPs, the following experimental data is used for illustration:
[0138] The performance of various ISPs was compared on the ZRR dataset, including GPU-based ISPs, CPU-based ISPs, and the FPGA-based ISP of this solution.
[0139] like Figure 9 As shown in the example graphs illustrating the qualitative analysis of different ISP performance, the FPGA-based ISP proposed in this scheme achieves superior image processing results.
[0140] Table 1 below presents the specific quantitative results. The proposed ISP achieves satisfactory results in shadow detail, specular control, and color reproduction, with the fewest number of parameters (4.7k), the lowest power consumption (23.0W), and the fastest inference speed (1754FPS), while consuming the least computational resources (0.23GFLOPS). Furthermore, this ISP achieves a frame rate of 76.3 frames per watt, which is 4.7 times higher than other ISPs.
[0141] Table 1:
[0142]
[0143] In comparison with GPUs and CPUs, as shown in Table 2 below, the ISP deployment results on CPU and GPU platforms demonstrate that our proposed ISP (implemented on FPGA) achieves a latency of 0.57 milliseconds and a frame rate of 1754 FPS, while consuming only 23.0W. In contrast, the TITAN Xp and GTX 1080 Ti consume 166.2W and 125.3W respectively, far exceeding the design of our proposed solution, and their frame rates are also lower. Although the Intel E5-2680 and I5-13400 have frequencies of 2500MHz and 4600MHz respectively, their inference speeds are still slower, and their power consumption is several times higher than that of our proposed solution.
[0144] Table 2:
[0145]
[0146] In the comparison with FPGA automation tools, Tables 3 and 4 below show the comparison between the proposed EFA and the automated FPGA design tool DPU on various CNN configurations. The results show that when the network contains only convolutional layers, the proposed EFA significantly outperforms the DPU. In networks with three convolutional layers, the proposed EFA runs 12.9 times faster than the DPU. Even in more complex networks with ten convolutional layers, the proposed EFA is still 5.4 times faster.
[0147] Table 3:
[0148]
[0149] Table 4:
[0150]
[0151] Table 5 below illustrates the effectiveness of hardware constraint pruning. As the number of network channels decreases, the number of DDR accesses gradually decreases, reaching 0 when pruned to 16 channels. Simultaneously, network latency decreases from 1.94 milliseconds with 128 channels to 0.67 milliseconds with 16 channels, and the number of frames increases from 515 to 1492. With the model size reduction, network parameters decrease from 166.16k to 4.65k, and GPU inference power consumption decreases from nearly 250W to 166W. Furthermore, the network's PSNR only decreases by 0.59%. Although the network performs better in terms of parameter count, latency, and power consumption when pruned to 8 channels, the PSNR decreases by 2.47%, therefore, this solution ultimately selects a 16-channel network for FPGA deployment.
[0152] Table 5:
[0153]
[0154] In summary, this solution has the following beneficial effects:
[0155] 1. This solution proposes a high-speed ISP for dedicated cameras. This is an FPGA-supported ISP based on deep learning, which can complete the image signal processing process with ultra-low power consumption, low cost and fast inference speed.
[0156] 2. A non-data exchange-free FPGA accelerator (EFA) is proposed, which includes a computation engine for ISP and a fully on-chip partitioned structure, improving the inference speed of the ISP process. With the help of EFA, the number of off-chip memory accesses can be significantly reduced and inference speed can be improved.
[0157] 3. A network pruning method for FPGA hardware is proposed, which reduces the number of parameters and computational load of the ISP network.
[0158] 4. A network was deployed on the FPGA via EFA, and an FPGA-based ISP was implemented. This is an FPGA accelerator for a deep learning-based ISP method, providing an inference speed of 1743 FPS, 4.7 times more energy efficient than existing ISPs, while maintaining comparable image quality.
[0159] 5. Extensive experiments were conducted on both synthetic and real haze images. The results show that this method achieves better dehazing performance and generalization compared to existing methods.
[0160] Corresponding to the model deployment method provided in the above embodiments of the present invention, see also... Figure 10 The present invention also provides a structural block diagram of a model deployment system, which includes: a generation module 1001, a reproduction module 1002, an iteration module 1003, and a deployment module 1004;
[0161] The generation module 1001 is used to generate an initial population, where each individual in the initial population corresponds to a convolutional neural network structure.
[0162] Specifically, the generation module 1001 is used to generate an initial population through a predefined sparsity standard random pruning filter.
[0163] The breeding module 1002 is used to select individuals from the initial population for breeding according to evaluation indicators, which include at least the number of DDR accesses and image quality.
[0164] Specifically, the breeding module 1002 is used to select individuals from the initial population for breeding according to the evaluation criteria and using a tournament selection method.
[0165] The iteration module 1003 is used to select individuals from the breeding results according to the evaluation index for multiple rounds of crossover and mutation until the iteration stopping condition is met, so as to obtain a set of convolutional neural network structures with optimal accuracy and delay tradeoff.
[0166] Deployment module 1004 is used to deploy a convolutional neural network structure with optimal accuracy and latency tradeoff to the FPGA disclosed in the above embodiments that does not require data exchange, so as to obtain the image signal processor disclosed in the above embodiments.
[0167] Specifically, the deployment module 1004 is used to: train a convolutional neural network structure with optimal accuracy and latency tradeoff using a sample dataset until training is complete, so as to obtain a neural network model; and deploy the trained neural network model to the FPGA disclosed in the above embodiments that does not require data exchange, so as to obtain the image signal processor disclosed in the above embodiments.
[0168] In summary, this invention provides an image signal processor, a model deployment method, and a system. The image signal processor is implemented based on an FPGA that eliminates the need for data exchange. The FPGA includes at least a computing engine and a fully on-chip partitioned structure for data storage. The computing engine is used to: sequentially perform multi-layer convolution calculations on each feature map block in the input buffer, and alternately cache the calculation results obtained during the multi-layer convolution calculations on the feature map blocks in a first feature buffer and a second feature buffer; shuffle the calculation results obtained after the last convolution calculation on the feature map blocks, and store the pixel shuffling results in the output buffer; and integrate the pixel shuffling results in the output buffer to form an output feature map. This eliminates the number of off-chip memory accesses and improves resource utilization, enhancing the computational efficiency of the image signal processor while maintaining superior image quality.
[0169] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0170] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0171] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image signal processor, characterized by, The image signal processor is implemented based on a field programmable gate array (FPGA) without data exchange, the FPGA at least comprising a computing engine and a full-chip partition structure for data storage, the full-chip partition structure at least comprising an input buffer, a first feature buffer, a second feature buffer and an output buffer; wherein input feature maps of the image signal processor are divided into a plurality of feature map blocks and buffered in the input buffer; the computing engine is configured to sequentially perform multi-layer convolution calculation on each feature map block in the input buffer, and alternately buffer the calculation results obtained in the process of performing multi-layer convolution calculation on the feature map block in the first feature buffer and the second feature buffer; perform pixel shuffling on the calculation results obtained by completing the last convolution calculation on the feature map block, and store the pixel shuffling results in the output buffer; and integrate each pixel shuffling result in the output buffer to form an output feature map.
2. The image signal processor of claim 1, wherein, The computing engine comprises a convolution engine composed of a plurality of weighted dot product engine arrays, and each convolution kernel is executed by one of the weighted dot product engine arrays.
3. The image signal processor of claim 2, wherein, Each of the weighted dot product engine arrays is composed of a plurality of weighted dot product engines and an input adder tree.
4. A model deployment method characterized by, The method comprises: generating an initial population, each individual of the initial population corresponding to a convolutional neural network structure; selecting individuals from the initial population for breeding according to an evaluation index, the evaluation index at least comprising DDR access quantity and image quality; selecting individuals from the breeding results for multiple rounds of crossover and mutation according to the evaluation index until an iteration stop condition is met, to obtain a group of convolutional neural network structures with optimal accuracy and delay trade-off; deploying the convolutional neural network structure with optimal accuracy and delay trade-off to a field programmable gate array (FPGA) without data exchange, to obtain the image signal processor of any one of claims 1 to 3; the FPGA at least comprising a computing engine and a full-chip partition structure for data storage, the full-chip partition structure at least comprising an input buffer, a first feature buffer, a second feature buffer and an output buffer; wherein input feature maps of the image signal processor are divided into a plurality of feature map blocks and buffered in the input buffer; the computing engine is configured to sequentially perform multi-layer convolution calculation on each feature map block in the input buffer, and alternately buffer the calculation results obtained in the process of performing multi-layer convolution calculation on the feature map block in the first feature buffer and the second feature buffer; perform pixel shuffling on the calculation results obtained by completing the last convolution calculation on the feature map block, and store the pixel shuffling results in the output buffer; and integrate each pixel shuffling result in the output buffer to form an output feature map.
5. The method of claim 4, wherein, deploying the convolutional neural network structure with optimal accuracy and delay trade-off to a field programmable gate array (FPGA) without data exchange, to obtain the image signal processor of any one of claims 1 to 3, comprising: training the convolutional neural network structure with the optimal precision and delay trade-off until training is completed using a sample data set to obtain a neural network model; deploying the trained neural network model to an FPGA without data exchange to obtain the image signal processor of any one of claims 1-3.
6. The method of claim 4, wherein, generating an initial population, including: generating the initial population by randomly pruning filters according to a predefined sparsity criterion.
7. The method of claim 4, wherein, selecting individuals from the initial population for reproduction according to an evaluation index, including: selecting individuals from the initial population for reproduction according to an evaluation index using a tournament selection method.
8. A model deployment system, characterized by, The system includes: a generating module configured to generate an initial population, each individual of the initial population corresponding to a convolutional neural network structure; a reproduction module configured to select individuals from the initial population for reproduction according to an evaluation index, the evaluation index including at least the number of DDR accesses and image quality; an iteration module configured to select individuals from the reproduction results for multiple rounds of crossover and mutation according to the evaluation index until an iteration stop condition is met to obtain a set of convolutional neural network structures with the optimal precision and delay trade-off; a deployment module configured to deploy the convolutional neural network structure with the optimal precision and delay trade-off to an FPGA without data exchange to obtain the image signal processor of any one of claims 1-3. The FPGA includes at least a computing engine and an on-chip partition structure for data storage, the on-chip partition structure including at least an input buffer, a first feature buffer, a second feature buffer, and an output buffer. The input feature map input into the image signal processor is divided into multiple feature map blocks and cached in the input buffer. The computing engine is configured to sequentially perform multi-layer convolution calculation on each feature map block in the input buffer, alternately cache the calculation results obtained during the multi-layer convolution calculation of the feature map blocks in the first feature buffer and the second feature buffer, perform pixel shuffling on the calculation results obtained after the last convolution calculation of the feature map blocks, store the pixel shuffling results in the output buffer, and integrate the pixel shuffling results in the output buffer to form an output feature map.
9. The system of claim 8, wherein, The deployment module is specifically configured to train the convolutional neural network structure with the optimal precision and delay trade-off until training is completed using a sample data set to obtain a neural network model, and deploy the trained neural network model to an FPGA without data exchange to obtain the image signal processor of any one of claims 1-3.
10. The system of claim 8, wherein, The generating module is specifically configured to generate the initial population by randomly pruning filters according to a predefined sparsity criterion.
Citation Information
Patent Citations
Data processing method and data processing equipment
CN104252338A
Image processing method and device, computer equipment and computer readable storage medium
CN115456858A