System and method for improving image processing efficiency based on wavelet transform
By rearranging and storing the image data after wavelet transformation in embedded devices, the problem of excessive computing power consumption caused by address jumping during inverse transformation of wavelet transformation in embedded devices is solved, and efficient image processing and efficient operation of MWCNN network on the embedded side is achieved.
Patent Information
- Application Number
- CN202111459758.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-02
AI Technical Summary
In embedded devices, when the deep learning-based image recovery algorithm runs in low-power, low latency, and limited resources, there is a problem of excessive computing power consumption, especially unnecessary consumption due to address jumps during inverse wavelet transform.
By wavelet transforming the image data in the input cache and rearranging it, the data of the corresponding points in the image frame are stored in the output cache in sequence according to the image sequence number to avoid address jumps and realize continuous address reading.
This method effectively reduces computing power consumption and improves image processing efficiency. It is especially suitable for embedded devices with low computing power, and realizes the efficient operation of MWCNN network on the embedded end.
Smart Images

Figure CN114066713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and in particular to a system and method for improving image processing efficiency based on wavelet transform. Background Art
[0002] In recent years, with the development of deep learning technology, many scholars have introduced neural networks into image restoration algorithms and achieved good results. With the large-scale deployment of various deep learning technology-based devices in scenarios such as smart cities, smart medical care, unmanned driving, and military security, image restoration technology based on deep learning will also be applied to these devices. However, these application scenarios have strict requirements on the power consumption, latency, resource utilization, and real-time performance of the equipment. Therefore, it is urgent to study the implementation of deep learning-based image restoration algorithms in low-power, low-latency, and resource-limited devices.
[0003] With the development of convolutional neural networks, some scholars have proposed the MWCNN network, which is based on the UNet architecture and uses wavelet transform (DWT) and inverse transform to replace pooling and deconvolution. This network has excellent performance in image denoising, single-frame image super-resolution and JPEG image artifact removal, and has obvious improvements in performance and speed. However, this network runs on a PC or server, which is not very friendly to some specific scenarios.
[0004] MWCNN deployed on embedded devices requires a neural network accelerator for acceleration. The accelerator module structure is as follows: Figure 1 The image processing method flow is as follows Figure 2 Figure 3 As shown, the image data to be calculated is stored in the input cache of the accelerator module; the computing unit rearranges, segments, and convolves the image data; the forward reasoning operation after WMCNN block division uses pipeline processing: the intermediate results of the calculation of each part are stored in the output cache and used in subsequent operations.
[0005] The way to store in RAM is as follows Figure 4 As shown, the image data of each block is stored according to the original image sequence. First, the point data X11-X1n of the first frame image J1 is stored in sequence, and then the point data X21-X2n of the second frame image J2 is stored in sequence, and then the third frame image J3 and the fourth frame image J4 are processed in the same way.
[0006] When the wavelet transform needs to be inversely transformed, four points in the inversely transformed image are obtained by calculation from the same position of the four frames of images, such as the first points P11, P21, P31 and P41 in each frame of image. Figure 4When reading the same point in four frames of images, there is an address jump. When reading data through the AXI interface, the address and the number of data to be read must be configured. Since it is not possible to read continuously in the manner of increasing addresses, it is necessary to read some data first and then jump to the next frame to read some data, which will definitely lead to unnecessary computing power consumption.
[0007] In embedded devices, computing power is already very limited, and excessive consumption will inevitably reduce image processing capabilities. Summary of the invention
[0008] The purpose of the present invention is to provide a system and method for improving image processing efficiency based on wavelet transform, so as to solve the problems existing in the above-mentioned prior art.
[0009] The method for improving image processing efficiency based on wavelet transform of the present invention performs wavelet transform on the image data in the input buffer and then rearranges them, and stores the data of the points at corresponding positions in each frame of the image in the output buffer in sequence according to the image sequence number.
[0010] The number of images is four frames.
[0011] First, the first point data of the first frame image to the fourth frame image is stored, and then the second point data of the first frame image to the fourth frame image is stored, and the corresponding point data of the first frame image to the fourth frame image is stored in sequence until the nth point data of the first frame image to the fourth frame image is stored; each frame image contains n points.
[0012] The system for improving image processing efficiency based on wavelet transformation of the present invention utilizes the method to store data after wavelet transformation.
[0013] The system and method for improving image processing efficiency based on wavelet transform described in the present invention have the advantage that the image is stored in the form of pixels, and can be directly read in a manner of continuously increasing addresses during subsequent inverse transformation. There is no address jump, which effectively reduces computing power consumption and is particularly suitable for use in embedded devices with low computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the accelerator module structure of WMCNN in the prior art;
[0015] Figure 2 It is a schematic diagram of image processing based on wavelet transform in the prior art;
[0016] Figure 3 It is a schematic diagram of the structural change of image data during wavelet transformation in the prior art;
[0017] Figure 4 It is a schematic diagram of data storage after wavelet transformation in the prior art.
[0018] Figure 5 It is a schematic diagram of data storage after wavelet transformation in the method of the present invention. DETAILED DESCRIPTION
[0019] First, the trained MWCNN is quantized and pruned on the PC side; a neural network accelerator is designed on the FPGA side to accelerate the network's forward reasoning. The structure of the accelerator module is also as follows Figure 1 shown.
[0020] The system and method for improving image processing efficiency based on wavelet transform of the present invention are as follows Figure 5 As shown:
[0021] The image data and weight data to be calculated are stored in the input buffer of the accelerator module;
[0022] Rearrange the data, that is, divide each frame of the image into small blocks of 32*32*1 for calculation. The calculation process of forward reasoning is also the same Figure 2 As shown. The result of the first convolution (32*32*4) is stored in the first block of the output buffer, because it will be added to the subsequent convolution result and then convolved again to get the output, so this is also the reason for block convolution, there is a residual structure.
[0023] Convolution parallel operation is used for convolution operation, which is equivalent to accelerating the convolution operation. The DWT operation and convolution operation are pipelined: the intermediate results of the calculation of each part are stored in the output cache and used in subsequent operations.
[0024] After dividing a frame of image into blocks for operation, the amount of data in the intermediate cache can be reduced. The convolution and wavelet transform (DWT) are pipelined, and two rows of data are obtained from the previous operation, and the wavelet transform operation can be performed, which reduces the interaction with RAM and saves time. The corresponding pixel point data is stored in sequence according to the order of each frame of the image, avoiding the storage method of jumping space, greatly improving the efficiency of image processing.
[0025] The specific storage method is as follows Figure 5 As shown, the data X11 of the first pixel point P11 of the first frame image J1 is first stored, and then the data X21 of the first pixel point P21 of the second frame image J2 is stored, followed by the data X31 of the first pixel point P31 of the third frame image J3, and finally the data X41 of the first pixel point P41 of the fourth frame image J4 is stored, completing the cycle of the first pixel point.
[0026] Perform the next pixel loop: first store the data X12 of the second pixel P12 of the first frame image J1, then store the data X22 of the second pixel P22 of the second frame image J2, then store the data X32 of the second pixel P32 of the third frame image J3, and finally store the data X42 of the second pixel P42 of the fourth frame image J4 to complete the loop for the second pixel.
[0027] The third and fourth pixels are subsequently stored in a loop according to the same storage logic until the loop of the nth and last pixel in each frame image is completed: first, the data X1n of the nth pixel P1n of the first frame image J1 is stored, and then the data X2n of the nth pixel P2n of the second frame image J2 is stored, followed by the data X3n of the nth pixel P3n of the third frame image J3, and finally the data X4n of the nth pixel P4n of the fourth frame image J4 is stored, completing the loop of one pixel.
[0028] The accelerator module of the present invention can be implemented in FPGA, and the intermediate cache data can be cached by BRAM, which is connected to BRAM via AXI interface. When it is necessary to obtain the inverse transformed picture, the pixels in four frames of images need to be calculated. When reading data through the AXI interface, the address and the number of read data need to be configured. Due to the unique storage method of the present invention, the data can be read directly in a manner of continuous increase of the address without address jump.
[0029] Beneficial effects: By dividing the image into blocks, the memory of the intermediate cache is reduced; the convolution and wavelet transform are pipelined, which reduces the interaction between the calculation module and the cache module and reduces the operation time; finally, a new image storage method based on inverse wavelet transform is proposed, which greatly reduces the time of inverse wavelet transform. For the calculation process of convolution, wavelet transform, and inverse wavelet transform of the entire acceleration module, this system has accelerated it and greatly saved resources, making it possible for the MWCNN network to run on the embedded end.
[0030] For those skilled in the art, various other corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all of these changes and deformations should fall within the protection scope of the claims of the present invention.
Claims
1. A method for improving image processing efficiency based on wavelet transform, characterized in that: First, the trained MWCNN is quantized and pruned on the PC side; a neural network accelerator is designed on the FPGA side to accelerate the network's forward reasoning; the image data and weight data to be calculated are stored in the input cache of the accelerator module; convolution operations are performed in parallel, and DWT operations and convolution operations are pipelined; the image data in the input cache is wavelet transformed and rearranged, and the data of the points at the corresponding positions in each frame of the image are stored in the output cache in sequence according to the image sequence number; The number of images is four frames; First, the first point data of the first frame image to the fourth frame image is stored, and then the second point data of the first frame image to the fourth frame image is stored, and the corresponding point data of the first frame image to the fourth frame image is stored in sequence until the nth point data of the first frame image to the fourth frame image is stored; each frame image contains n points.
2. A system for improving image processing efficiency based on wavelet transform, characterized in that: The data after wavelet transformation is stored using the method as claimed in claim 1.
Citation Information
Patent Citations
Neural network processing system for reducing IO overhead based on wavelet transform
CN108665062A
Convolutional neural network accelerator supporting sparse pruning based on FPGA design
CN111242277A