Convolution Input Data Determination, Convolution Operation, Image Denoising, Method and System
By designing a convolutional input data determination method in a CNN-based image denoising algorithm, data multiplexing is achieved using BT and BL cache structures, the problems of large calculation volume and high power consumption in practical applications are solved, and the image denoising effect with high efficiency and low power consumption is achieved.
Patent Information
- Application Number
- CN202111530167.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In practical applications, image denoising algorithms based on CNN are limited by huge parameters and calculations, resulting in high inferred delay and high power consumption, making it difficult to meet the real-time and low power consumption requirements.
By designing a method for determining convolution input data, the two data cache structures BT and BL are used to cache the convolution results before the current operation result to determine the convolution input data of the next layer of convolution operation, realizing data multiplexing and reducing repeated calculations.
This method greatly improves the computing efficiency of the accelerator, reduces computing power consumption, and simplifies the data processing process, so as to achieve efficient image denoising under low resource consumption.
Smart Images

Figure CN114298926B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and specifically relates to a method and system for determining convolution input data, performing convolution operations, image denoising, an accelerator, and an image processing chip. Background Art
[0002] Image denoising is an important image preprocessing task. Compared with the image before denoising, a noiseless image can improve the performance of other computer vision tasks and provide better visual effects at the same time. Compared with traditional image denoising algorithms, the image denoising algorithm based on CNN (Convolutional Neural Network) can achieve great performance improvement in both denoising quantization metrics and intuitive perception.
[0003] Although the image denoising algorithm based on CNN can achieve good denoising effects, it requires a huge number of parameters and a large amount of computation compared with traditional algorithms, and is very vulnerable to limitations in actual application scenarios. If such an algorithm is deployed on a CPU platform, the huge amount of computation will bring unacceptable inference latency and cannot meet the requirements of real-time tasks. If such an algorithm is deployed on a GPU (Graphics Processing Unit) platform, the huge power consumption requirements cannot meet the low power consumption requirements of embedded applications. Many CNN accelerators are implemented using FPGA (Field Programmable Gate Array) to accelerate the image denoising algorithm based on CNN, but there is currently no hardware accelerator design for the image denoising algorithm based on CNN. Moreover, the image denoising algorithm based on CNN is very different from traditional image classification algorithms. Its input image resolution is the same as the output image, so more intermediate results will be generated. Under the previous hardware accelerator design methods for classification networks, the hardware implementation of the denoising algorithm requires a large off-chip data interaction requirement. Therefore, it is necessary to design a neural network hardware accelerator with low power consumption, low latency, and low resource consumption according to the characteristics of the image denoising algorithm. In this hardware accelerator design, it is necessary to design and implement a good data reuse scheme to avoid additional data transmission or recalculation processes. Summary of the Invention
[0004] In view of this, this application provides a method and system for determining convolution input data, performing convolution operations, image denoising, an accelerator, and an image processing chip, so as to provide a convolution input data determination scheme during the image denoising algorithm based on CNN, simplify the data processing process, and improve the data processing efficiency.
[0005] A method for determining convolution input data provided by this application includes:
[0006] Determine the convolution operation input data corresponding to the current operation result;
[0007] Store the k-s row convolution results before the current operation result into BT; k is the dimension of the convolution input data, s is the step size of a single slide in the convolution operation, and BT is a preset first data cache structure;
[0008] Store the k-s column convolution results before the current operation result in BT into BL, so as to determine the convolution input data of the next layer of convolution operation according to BT and BL; BL is a preset second data cache structure.
[0009] Optionally, the current operation result is the operation result of the i-th row and the j-th column, the size of the convolution input data is k×k, the size of BT is (k-s)×(IMGW + k-s), the size of BL is k×(k-s), and IMGW is the number of columns of the image to be processed;
[0010] Storing the k-s row convolution results before the current operation result into BT includes: storing the convolution results of the operation results of the i-(k-s) to i-s rows into BT;
[0011] Storing the k-s column convolution results before the current operation result in BT into BL includes: storing the operation results of the j-(k-s) to j-s columns in BT into BL.
[0012] Optionally, determining the convolution input data of the next layer of convolution operation according to BT, BL and the current operation result includes:
[0013] Determine the k-s column data of the convolution input data according to BL;
[0014] Determine k-s data of one column other than the k-s columns in the convolution input data according to the data in the j-th column of BT;
[0015] Determine the other undetermined data in the convolution input data according to the current operation result.
[0016] Optionally, when obtaining the operation result of the first row, the method of storing data into BT includes:
[0017] Determine the first storage position corresponding to the operation result of the first row in BT, write the operation results of each row in the first row into the corresponding storage position in turn, and write fill values into other storage positions of BT.
[0018] Optionally, after storing the k-s column convolution results before the current operation result in BT into BL, the method for determining the convolution input data further includes:
[0019] Replace the oldest column of data in BL with the column of data where the current operation result is located in the convolution input data.
[0020] Optionally, when obtaining the operation results of one row, the updating method of the BT includes:
[0021] Replacing the oldest s rows of data in the BT with the latest s rows of operation results.
[0022] Optionally, when the current operation result is located in the first column, the updating method of the BL includes:
[0023] Obtaining the corresponding column of the current operation result in the BT to obtain the first storage column;
[0024] Obtaining the corresponding column of the current operation result in the BL to obtain the second storage column;
[0025] Determining the storage position of the current operation result in the second storage column to obtain the second storage position;
[0026] Writing the data of the first storage column into the other storage positions in the second storage column except the second storage position, writing padding values into the other storage columns in the BL except the second storage column, and writing the second storage position into the current operation result.
[0027] Optionally, the number of rows of the image to be processed is IMGH, and the number of columns is IMGW; the size of the entire convolution operation result for the next layer of convolution operation is (IMGH + k - s) × (IMGW + k - s), including the feature map located in the middle position and the padding values located around, and the size of the feature map is IMGH × IMGW.
[0028] This application also provides a convolution operation method, including:
[0029] Determining convolution input data by using any of the above convolution input data determination methods;
[0030] Performing a convolution operation on a convolution object by using the convolution input data.
[0031] This application also provides an image denoising method, including:
[0032] Performing at least one convolution operation on the image to be processed by using any of the above convolution operation methods to remove the noise of the image to be processed.
[0033] Optionally, the performing at least one convolution operation on the image to be processed by using any of the above convolution operation methods includes:
[0034] Vertically splitting the image to be processed to obtain a plurality of sub-images;
[0035] Perform at least one convolution operation on each sub-image using any of the above convolution operation methods to obtain denoised sub-images corresponding to each sub-image;
[0036] Stitch together the denoised sub-images to obtain a denoised image corresponding to the image to be processed.
[0037] This application also provides a convolution input data determination system, including:
[0038] A first determination module, configured to determine convolution operation input data corresponding to a current operation result;
[0039] A first storage module, configured to store the convolution results of k - s rows before the current operation result into BT; k is the dimension of the convolution input data, s is the step size of one sliding in the convolution operation, and BT is a preset first data cache structure;
[0040] A second storage module, configured to store the convolution results of k - s columns before the current operation result in BT into BL, so as to determine the convolution input data of the next layer of convolution operation according to BT and BL; BL is a preset second data cache structure.
[0041] This application also provides an accelerator, including an acceleration circuit; the acceleration circuit is configured to execute any of the above convolution input data determination methods or any of the above convolution operation methods.
[0042] This application also provides an image processing chip, including any of the above accelerators.
[0043] In the above convolution input data determination, convolution operation, image denoising, method and system, accelerator, and image processing chip of the present application, by determining the convolution input data corresponding to the current operation result, using BT to cache the k-s row convolution results before the current operation result, and using BL to cache the k-s column convolution results before the current operation result, to determine the convolution input data for the next layer of convolution operation according to the data cached in BT and BL respectively; wherein the data cached in BT and BL can be reused during multiple convolution input data determination processes, and the corresponding data reuse module stores sufficient intermediate results for calculating each pixel to be calculated, and the data output result of each layer is only related to the values stored in the reuse module during the calculation of the previous layer, avoiding repeated calculations of the neural network depth level when calculating pixel outputs at different positions of the output layer, enabling on-chip data caching of neural network data, avoiding frequent off-chip data interactions when obtaining intermediate results of the upper layer, greatly improving the operation efficiency of the accelerator, and reducing the operation power consumption. In addition, the above convolution input data determination method can also use Input to prepare the convolution input data required for the next convolution operation, that is, store the convolution input data in Input, and establish the corresponding relationships between Input and BT and BL respectively through a sliding window to quickly and orderly determine the convolution input data, which can further improve the determination efficiency of the convolution input data, thereby improving the operation efficiency of the corresponding accelerator and image processing chip and reducing the power consumption of the corresponding accelerator and image processing chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 is a schematic flowchart of a method for determining convolution input data according to an embodiment of the present application;
[0046] Figure 2 is a schematic diagram of the sliding direction of a sliding window according to an embodiment of the present application;
[0047] Figure 3a 、 Figure 3b and Figure 3c are schematic diagrams of a sliding window according to an embodiment of the present application;
[0048] Figure 4a 、 Figure 4b 、 Figure 4c and Figure 4d are schematic diagrams of the cache data update process according to an embodiment of the present application;
[0049] Figure 5a 、Figure 5b , Figure 5c and Figure 5d are schematic diagrams of the cache data update process in an embodiment of the present application;
[0050] Figure 6 is a schematic diagram of the system structure for determining convolution input data in an embodiment of the present application. Detailed implementation manners
[0051] The CNN-based image denoising algorithm often includes multiple convolution operations. The convolution kernel of each convolution operation, which is the convolution input data, is determined based on the result of the previous convolution operation, easily generating a large number of intermediate results, and the involved parameter quantity and computational complexity are extremely large.
[0052] To address the above problems, the convolution input data determination, convolution operation, image denoising, method and system, accelerator, and image processing chip provided by the present application use BT to cache the k-s row convolution results before the current operation result, and use BL to cache the k-s column convolution results before the current operation result, so as to determine the convolution input data of the next layer of convolution operation according to the data cached in BT and BL respectively, and the current operation result; among them, the data cached in BT and BL can be reused during multiple convolution input data determination processes. The corresponding data reuse module stores sufficient intermediate results for each pixel to be calculated. The data output result of each layer is only related to the values stored in the reuse module during the calculation of the previous layer, avoiding the repeated calculation of the neural network depth level when calculating the pixel outputs at different positions of the output layer, enabling on-chip data caching of neural network data, avoiding frequent off-chip data interaction when obtaining the upper-layer intermediate results, greatly improving the operation efficiency, and reducing the operation power consumption.
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. Without conflict, the following various embodiments and their technical features can be combined with each other.
[0054] The first aspect of the present application provides a method for determining convolution input data. Referring to Figure 1 as shown, the method for determining convolution input data includes:
[0055] S110, determining the convolution operation input data corresponding to the current operation result;
[0056] S120, store the convolution results of the k-s rows before the current operation result into BT; k is the dimension of the convolution input data, s is the step size of one sliding in the convolution operation, and BT is a preset first data cache structure;
[0057] S130, store the convolution results of the k-s columns before the current operation result in BT into BL, so as to determine the convolution input data of the next layer of convolution operation according to BT and BL; BL is a preset second data cache structure. Optionally, after this step, the convolution input data of the next layer of convolution operation can be determined according to BT, BL and the current operation result. Optionally, a cache result Input can be set, and the convolution input data is buffered by Input to determine the convolution input data according to the data buffered in Input.
[0058] The image denoising algorithm based on CNN includes multiple convolution operations. The convolution kernel of each convolution operation, which is the convolution input data, is determined according to the result of the previous convolution operation; the result obtained by each convolution operation can also be called a feature map. The above convolution input data can be characterized as a sliding window that can slide on the convolution operation result. The size of this sliding window is k×k. Starting from the first row and first column of the convolution operation result, it slides to the right with a step size of s, and after traversing the operation results of one row, it continues to slide to the right from the first operation result of the next row. The corresponding sliding direction can be referred to Figure 2 as shown; where each sliding determines the convolution input data once, which is used as the input for the subsequent convolution operation. The obtained convolution result can continue to be used as the basis for determining the subsequent convolution input data, and so on, until the denoising of the image to be processed is achieved or the number of convolution times reaches the preset number threshold. The values of the sliding window dimension k and the sliding step size s can be set according to factors such as the denoising effect required by the image to be processed. For example, the value range of k can be set from 3 to 10, and s can be set to 1, etc., to improve the operation efficiency on the basis of ensuring the value effect.
[0059] The second data cache structure BL is used to cache the data reused by adjacent sliding windows during the horizontal sliding of the convolution kernel (product of input data). Its design is mainly for convolution calculations with a convolution kernel size of k×k and a sliding step of s. At this time, when the convolution kernel slides horizontally, the amount of data reused by two adjacent sliding windows is k rows × (k - s) columns × the number of channels of the current feature map. The size of BL corresponding to one layer of convolution calculation is the amount of reused data. The first data cache structure BT is used to cache the data reused by two vertically adjacent sliding windows. Its design is for the row-first convolution calculation method, that is, when performing convolution operations, the convolution kernel slides horizontally to complete the convolution operation, and after completing one row of convolution operations, it slides down one step and starts a new row of convolution operations from the leftmost side of the feature map. Therefore, the interval of the data reused by two vertically adjacent sliding windows is larger than the interval of the data reused by horizontally adjacent sliding windows. This interval affects the column size of the entire convolution operation result after padding values are added to the feature map. Therefore, the size of the BT cache structure is (k - s) rows × (the number of columns of the feature map + the number of columns of the padded values) × the number of channels of the current feature map. Input is used to prepare the data of k rows × k columns × the number of channels of the current feature map required for one convolution operation. It can obtain the data required for the current convolution operation from BL, BT, and the calculation results of the previous layer. The multiply-accumulate unit can directly take values from this structure to complete the current convolution operation. The size of Input is k rows × k columns × the number of channels of the current feature map. Among them, the left (k - s) columns of data can be directly loaded from BL, the first (k - s) rows of data in the right s columns need to be loaded from the corresponding positions in BT, and the last row of data in the rightmost s columns needs to be obtained from the calculation results of the previous layer. In one example, if k = 3 and s = 1, the corresponding sliding window can be referred to Figure 3a as shown, such as Figure 3b shown. After a sliding window calculation is completed, the data in the rightmost column of the 3×3×the number of channels of the current feature map sliding window is the latest column of data and is also the data that needs to be cached. During the subsequent update of BL, this column of data can be stored in BL to achieve horizontal update of BL. When the sliding window slides to the right for convolution operations, the data cached in BL is the data in the left two columns of the 3×3×the number of channels of the current feature map sliding window required for this operation. The data in the rightmost column of the sliding window can be obtained from BT and the latest calculation results. As Figure 3c shown, BT is used to cache the data of the previous 2 rows before the current calculation result. After a sliding window calculation is completed, the lower right corner of the 3×3×the number of channels of the current feature map sliding window is the latest data (the current calculation result) and is also the data that needs to be cached. The data cached on the left side of this position in this row is the updated cached data for reuse in the next row of convolution operations. Figures 3a to 3cAmong them, the size of the Input structure is 3 rows × 3 columns × the number of channels of the current feature map. The data in the left two columns can be directly loaded from BL. The first two rows of data in the rightmost column need to be loaded from the corresponding positions in BT. The last row of data in the rightmost column is the current operation result and can be obtained from the calculation result of the previous layer.
[0060] For the above method for determining convolution input data, after determining the convolution operation input data corresponding to the current operation result, use BT to cache the k - s row convolution results before the current operation result, and use BL to cache the k - s column convolution results before the current operation result, so as to determine the convolution input data for the next layer of convolution operation according to the data cached in BT and BL and the current operation result; among them, the data cached in BT and BL can be reused during multiple processes of determining convolution input data, which can reduce the amount of intermediate parameters generated during the convolution operation, thereby reducing the computational amount. In addition, the above method for determining convolution input data can also use Input to prepare the convolution input data required for the next convolution operation, that is, store the convolution input data into Input, and establish the corresponding relationships between Input and BT and BL respectively through a sliding window to quickly and orderly determine the convolution input data, which can further improve the efficiency of determining convolution input data.
[0061] In one embodiment, the current operation result is the operation result of the i-th row and the j-th column, the size of the convolution input data is k×k, the size of BT is (k - s)×(IMGW + k - s), the size of BL is k×(k - s), and IMGW is the number of columns of the image to be processed;
[0062] The storing the k - s row convolution results before the current operation result into BT includes: storing the convolution results of the operation results from the (i - (k - 1))-th row to the i - s-th row into BT;
[0063] The storing the k - s column convolution results before the current operation result in BT into BL includes: storing the operation results of the (j - (k - s))-th column to the j - s-th column in BT into BL.
[0064] Specifically, the determining the convolution input data for the next layer of convolution operation according to BT, BL and the current operation result includes: determining the k - s column data of the convolution input data (such as the first k - s column data) according to BL; determining the k - s data of the column other than the k - s columns in the convolution input data according to the data in the j-th column of BT (such as the first k - s data of the k-th column data); determining the other undetermined data in the convolution input data according to the current operation result (such as the k-th data of the k-th column data).
[0065] When obtaining the operation result of the i-th row and j-th column in this embodiment, the convolution result of the operation results of the i-(k-s) to i-s rows is stored in BT, so that BT can be reused in the determination process of the convolution input data when obtaining the operation result of the i-th row. The operation results of the j-(k-s) to j-s columns in BT are stored in BL to determine a part of the convolution input data based on BL, which can ensure the accuracy and orderliness in the determination process of the convolution input data on the basis of simplifying the relevant operation process.
[0066] In one embodiment, when obtaining the operation result of the first row, the method of storing data in the BT includes: determining the first storage location corresponding to the operation result of the first row in the BT, writing each operation result of the first row into the corresponding storage location in sequence, and writing fill values into other storage locations of the BT. The fill value can be a value such as 0 that does not affect the effective convolution operation after being filled.
[0067] In an example, if k = 3, s = 1, the fill value is 0, and the size of BT is 2×(IMGW + 2), as Figure 4a shown, when obtaining the operation result of the first row, the first storage location corresponding to each operation result in the operation result of the first row can be determined first, each operation result is written into the corresponding storage location in sequence, and then written into other storage locations of BT to determine the initial stored data of BT. As Figure 4a shown, the first storage location is located in the middle of the second row in BT. In other examples, the first storage location can also be located in the middle of the first row in BT. Optionally, when obtaining the operation result of the first row, only each operation result needs to be written into BT, and BL and Input do not need to be updated.
[0068] The operation result of the first row includes the first operation result of the first row. Optionally, when obtaining the first operation result F11 of the first row, the value of BT is updated. Specifically, it can be referred to Figure 4b shown, determine that the storage location corresponding to the first operation result F11 is the storage location (0,1), write 0 into the 3 storage locations (0,0), (1,0), and (1,1) of BT, and write the new value obtained from the output of F11 into the (0,1) location. BL remains unchanged, and the convolution operation of F12 and the write operation of the corresponding Input are not performed. Optionally, when obtaining the intermediate operation result of the first row, such as the col1-th operation result F12, the value of BT is updated, as Figure 4cAs shown, write the col1-th operation result F12 to the position (0, col1 + 1) of BT, write 0 to the position (1, col1 + 1) of BT, and do not perform read and write operations on BL and Input. Optionally, when obtaining the last operation result of the first row, update BT, write the last operation result to the corresponding storage position, and fill zeros in the upper right corner of BT; for example, when obtaining the operation result of the point (0, IMGW), update the value of BT, which can be as Figure 4d As shown, write 0 to the positions (0, IMGW + 1) and (1, IMGW + 1) of BT, and do not perform read and write operations on BL and Input.
[0069] In one embodiment, after storing the k - s column convolution results before the current operation result in the BT into BL, the convolution input data determination method further includes: replacing the oldest column of data in BL with the column of data where the current operation result is located in the convolution input data. In this embodiment, during the process of updating BL, replacing the oldest column of data in BL with the column of data where the current operation result is located in Input can reduce data transfer during the update process, thereby improving the corresponding operation efficiency.
[0070] In one embodiment, when obtaining a row of operation results, the method for updating BT includes: replacing the oldest s rows of data in BT with the latest s rows of operation results. In this embodiment, during the process of updating BT, covering the oldest row of data in BT with the latest s rows of operation results can reduce data transfer during the update process, thereby improving the corresponding operation efficiency.
[0071] In one embodiment, when the current operation result is located in the first column, the method for updating BL includes:
[0072] Obtain the corresponding column of the current operation result in BT to obtain the first storage column;
[0073] Obtain the corresponding column of the current operation result in BL to obtain the second storage column;
[0074] Determine the storage position of the current operation result in the second storage column to obtain the second storage position (such as the last position in the second storage column);
[0075] Write the data of the first storage column to the other storage positions in the second storage column except the second storage position, write fill values (such as 0) to the other storage columns in BL except the second storage column, and write the second storage position with the current operation result.
[0076] In one example, if k = 3, s = 1, the padding value is 0, and the coordinates corresponding to the current operation result F13 are (row2, 0). When updating BL, as Figure 5a shown, 0 can be written to the positions (0, 1), (1, 1), and (2, 1) of BL, the current operation result F13 is written to the position (2, 0) of BL, and the data BT1 stored at ((row2 + 0) % 2, 1) and the data BT2 stored at ((row2 + 1) % 2, 1) in BT are written to the positions (0, 0) and (1, 0) of BL respectively (where % represents the modulo operation). Optionally, BT is updated later, and the position (2, 0) in BL is written to the position (row2, 1) of BT, and no read or write operations are performed on Input.
[0077] In one embodiment, the number of rows of the image to be processed (effective feature map) is IMGH, and the number of columns is IMGW; the size of the entire convolution operation result for the next layer of convolution operation is (IMGH + k - s) × (IMGW + k - s); it includes a feature map located in the middle position and padding values located around it, and the size of the feature map is IMGH × IMGW. In this embodiment, whole rows and whole columns of padding values can be filled around the effective feature map according to the dimension of the convolution kernel and the sliding step size to trigger the subsequent convolution operation in multiple convolution operations, using the calculation result output by the previous convolution layer as a drive to update the data cache structures of BL, BT, and Input; after the corresponding update is completed, the Input structure can provide data support for the multiply-accumulate calculation of this layer of convolution layer.
[0078] In one example, during the operation of the image denoising algorithm, in order to ensure that the resolution of the feature maps of each layer is equal, zero padding is performed around the feature maps before convolution operation. For different positions of the feature maps, the zero-padding situations are different. The pixel data of the zero-padded feature map is stored in BT, BL, and Input. Therefore, the update methods of these three cache results are different according to the different corresponding regions of the single-pixel output, where the region division of the feature map is as Figure 2 shown. And the overall behavior of the data reuse unit between layers is the same. If k = 3, s = 1, the padding value is 0, and the feature map resolution is IMGH × IMGW, for the feature maps of the same channel label in two adjacent layers: the feature map of the previous layer is denoted as F1, and the feature map of the next layer is denoted as F2. After a certain 3×3 region in F1 is convolved with the convolution kernel, the value of a single pixel is output, and this pixel is passed into the calculation module of the next layer F2 after being processed by the data reuse unit. Logically, the result after convolution of F1 is input to the corresponding position of F2 in the form of a single pixel. In terms of hardware implementation, this process is manifested as updating the storage spaces of BL, BT, and Input corresponding to the feature map of this channel. During the convolution process, the order in which different positions in F2 obtain the input data from F1, that is, the trajectory of the convolution kernel sliding on F1 is asFigure 2 as shown
[0079] Optionally, for the first data in the last line after padding with zeros, the value of BL can be updated first. As Figure 5b shown, write 0 to the positions (0,1), (1,1), (2,1) of BL, write 0 to the position (2,0) of BL, and write the data BT3 stored in ((IMGH + 0) % 2, 1) of BT and the data BT4 stored in ((IMGH + 1) % 2, 1) of BT to the positions (0,0), (1,0) of BL respectively. Then update BT, and write the value 0 at the position (2,0) of BL to the position (IMGH, 1) of BT. Do not perform read and write operations on Input.
[0080] Optionally, for the operation result at the middle position of the next layer F2, Input, BT, and BL can be updated sequentially. As Figure 5c shown, for the operation result at the point (row3, col3), at this time, BL, BT, and the new value F14 output from F1 can form a complete 3×3 data for convolution. First, update Input. The values of the first and second columns are updated through the values stored in BL (such as Figure 5c shown as BL1 to BL6). Write the values at the positions (0, (col3 + 0) % 2), (1, (col3 + 0) % 2), (2, (col3 + 0) % 2) of BL to the positions (0,0), (1,0), (2,0) of Input respectively, and write the values at the positions (0, (col3 + 1) % 2), (1, (col3 + 1) % 2), (2, (col3 + 1) % 2) of BL to the positions (0,1), (1,1), (2,1) of Input respectively. For the third column in Input, write the values at the positions ((row3 + 0) % 2, col3 + 1), ((row3 + 1) % 2, col3 + 1) of BT (such as Figure 5c shown as BT5 and BT6) to the positions (0,2), (1,2) of Input. Write the new value F14 output from F1 to the position (2,2) of Input. Optionally, then update BT, and write the new value F14 at the position (2,2) of Input, that is, the value output from F1, to the position (row3 % 2, col3 + 1) of BT. Finally, update BL, and write the values at the positions (0,2), (1,2), (2,2) of Input to the positions (0, col3 % 2), (1, col3 % 2), (2, col3 % 2) of BL respectively. The data in Input will be output to the convolution layer one by one in the form of single pixels for calculation.
[0081] Optionally, for the operation results of other boundary regions in the entire convolution operation result, Input, BT, and BL can be updated sequentially, the data of the corresponding regions can be stored, and the zero-padding operations for the right boundary, lower boundary, and lower right corner can be performed respectively: As Figure 5d shown, if the coordinate corresponding to the operation result is (row4, col4), at this time, the values of BL, BT, and the new value output from F1 can form a complete 3×3 data for convolution. First, update Input. The values of the first column and the second column are updated through the values stored in BL (as Figure 5d shown, BL1 to BL6). Write the values at positions (0, (col4 + 0) % 2), (1, (col4 + 0) % 2), and (2, (col4 + 0) % 2) in BL to positions (0, 0), (1, 0), and (2, 0) in Input respectively, and write the values at positions (0, (col4 + 1) % 2), (1, (col4 + 1) % 2), and (2, (col4 + 1) % 2) in BL to positions (0, 1), (1, 1), and (2, 1) in Input respectively. For the third column in Input, write the values at positions ((row4 + 0) % 2, col4 + 1) and ((row4 + 1) % 2, col4 + 1) in BT to positions (0, 2) and (1, 2) in Input respectively (as Figure 5d shown, BT5 and BT6). Write 0 to position (2, 2) in Input. Then update BT, and write the new value obtained from the output of F1 at position (2, 2) in Input to position (row4 % 2, col4 + 1) in BT. Finally, update BL, and write the values at positions (0, 2), (1, 2), and (2, 2) in Input to positions (0, col4 % 2), (1, col4 % 2), and (2, col4 % 2) in BL respectively. The data in Input will be output to the convolution layer pixel by pixel for calculation.
[0082] The above method for determining convolution input data, after determining the convolution operation input data corresponding to the current operation result, uses BT to cache the k-s row convolution results before the current operation result, and uses BL to cache the k-s column convolution results before the current operation result, so as to determine the convolution input data for the next layer of convolution operation according to the data cached by BT and BL respectively; among them, the data cached by BT and BL can be reused during multiple processes of determining convolution input data. The corresponding data reuse module stores sufficient intermediate results for each pixel to be calculated, and the data output result of each layer is only related to the values stored in the reuse module during the calculation of the previous layer, avoiding repeated calculations of the neural network depth level when calculating pixel outputs at different positions of the output layer, enabling on-chip data caching of neural network data, avoiding frequent off-chip data interactions when obtaining intermediate results of the upper layer, greatly improving the accelerator operation efficiency, and reducing the operation power consumption. In addition, the above method for determining convolution input data can also use Input to prepare the convolution input data required for the next convolution operation, that is, store the convolution input data in Input, and establish the corresponding relationships between Input and BT and BL respectively through a sliding window, so as to quickly and orderly determine the convolution input data, which can further improve the efficiency of determining convolution input data.
[0083] In a second aspect, the present application provides a convolution operation method, which can run on an acceleration module and / or an accelerator of a corresponding image processing chip, including:
[0084] Determine the convolution input data by using the method for determining convolution input data described in any of the above embodiments;
[0085] Perform a convolution operation on a convolution object by using the convolution input data.
[0086] For the above convolution operation method, the convolution input data is determined by using the method for determining convolution input data described in any of the above embodiments, and then a convolution operation is performed on a convolution object by using the convolution input data, which can form an accelerator of a corresponding image processing chip, simplify the convolution operation process, and improve the corresponding image processing efficiency. In addition, when the above convolution operation method uses a fused-layer architecture to implement a hardware accelerator based on a convolutional neural network, this module can cache the data required for its calculation for adjacent sliding windows, avoid additional repeated calculations or off-chip data interaction processes, accelerate the neural network inference process, and reduce power consumption.
[0087] In a third aspect, the present application provides an image denoising method, including: performing at least one convolution operation on a to-be-processed image by using the convolution operation method described in any of the above embodiments to remove noise from the to-be-processed image.
[0088] The data reuse process of the above image denoising method mainly includes two data cache structures, BL and BT, and an input data structure (Input). Since the data reuse interval in BT is relatively large, the width of the BT structure needs to be the same as the width of the processed image. To reduce the hardware resource consumption of the BT structure, an image vertical splitting method can be used. In one embodiment, performing at least one convolution operation on the image to be processed using the convolution operation method described in any of the above embodiments includes:
[0089] Vertically splitting the image to be processed to obtain a plurality of sub-images;
[0090] Performing at least one convolution operation on each sub-image respectively using the convolution operation method described in any of the above embodiments to obtain denoised sub-images corresponding to each sub-image;
[0091] Stitching the denoised sub-images to obtain a denoised image corresponding to the image to be processed.
[0092] In this embodiment, the image to be processed is vertically split to obtain a plurality of sub-images with relatively small widths, and then at least one convolution operation is performed on each sub-image respectively using the convolution operation method described in any of the above embodiments, which can effectively reduce the cache space required by BT and simplify the data operation process corresponding to BT in the convolution operation process, thereby further improving the operation efficiency; after obtaining the denoised sub-images corresponding to each sub-image, the denoised sub-images are stitched to obtain a denoised image corresponding to the image to be processed, which can ensure the effect of the obtained denoised image.
[0093] The above image denoising method vertically splits the image to be processed (the original noisy image), performs the entire neural network calculation, and then performs column-by-column stitching, reducing the data reuse interval, that is, reducing the width of BT, and can further reduce resource consumption.
[0094] The fourth aspect of this application provides a convolution input data determination system. Referring to Figure 6 as shown, the above convolution input data determination system includes:
[0095] A first determination module 110, configured to determine the convolution operation input data corresponding to the current operation result;
[0096] A first storage module 120, configured to store the convolution results of k - s rows before the current operation result into BT; k is the dimension of the convolution input data, s is the step size of one sliding in the convolution operation, and BT is a preset first data cache structure;
[0097] A second storage module 130, configured to store the k-s column convolution results before the current operation result in the BT into the BL, so as to determine the convolution input data for the next layer of convolution operation according to the BT and the BL; the BL is a preset second data cache structure.
[0098] For the specific limitations on the convolution input data determination system, reference can be made to the limitations on the convolution input data determination method in the above text, which will not be elaborated here. Each module in the above convolution input data determination system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the operation module of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the operation module of the computer device can call and execute the operations corresponding to the above modules.
[0099] The fifth aspect of the present application provides a convolution operation system, including:
[0100] A second determination module, configured to determine the convolution input data by using the convolution input data determination system described in any of the above embodiments;
[0101] An operation module, configured to perform a convolution operation on a convolution object by using the convolution input data.
[0102] For the specific limitations on the convolution operation system, reference can be made to the limitations on the convolution operation method in the above text, which will not be elaborated here. Each module in the above convolution operation system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the operation module of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the operation module of the computer device can call and execute the operations corresponding to the above modules.
[0103] The sixth aspect of the present application provides an accelerator, including an acceleration circuit; the acceleration circuit is configured to execute the convolution input data determination method described in any of the above embodiments or the convolution operation method described in any of the above embodiments.
[0104] The above accelerator can effectively accelerate the corresponding image denoising algorithm, has a good data reuse module, can avoid the processes of additional data transmission, recalculation, and interaction with the corresponding off-chip data, and has the advantages of low power consumption, low latency, and low resource consumption.
[0105] The seventh aspect of the present application provides an image processing chip, including the accelerator described in any of the above embodiments.
[0106] The above-mentioned image processing chip designs a data reuse module for the hardware implementation of the image denoising algorithm, including data cache structures such as BL, BT, and Input. The feature image pixel data of each channel is stored on these three storage units, which can realize on-chip data caching of neural network data, avoid frequent off-chip data interaction when obtaining upper-layer intermediate results, greatly improve the operation efficiency of the accelerator, and reduce the operation power consumption.
[0107] In addition, the data reuse module therein stores sufficient intermediate results for calculation for each pixel to be calculated. The data output result of each layer is only related to the value stored in the reuse module during the calculation of the previous layer, avoiding repeated calculations of the neural network depth level when calculating the pixel outputs at different positions of the output layer. It can also vertically divide the original noise map, perform the entire neural network calculation, and then splice by columns, reducing the data reuse interval, that is, reducing the width of BL, and further reducing resource consumption.
[0108] Although the present application has been shown and described with respect to one or more implementations, those skilled in the art will envision equivalent variations and modifications based on reading and understanding this specification and the drawings. The present application includes all such modifications and variations and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above components, the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (i.e., it is functionally equivalent), even if it is not structurally equivalent to the disclosed structure that performs the functions in the exemplary implementations of this specification shown herein.
[0109] That is, the above description is only an embodiment of the present application, and thus does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made using the content of the specification and drawings of the present application, such as the mutual combination of technical features between various embodiments, or direct or indirect application in other related technical fields, is similarly included within the patent protection scope of the present application.
[0110] In addition, for structural elements with the same or similar characteristics, the present application may use the same or different reference numerals for identification. Furthermore, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include one or more features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0111] In this application, the term "exemplary" is used to mean "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or superior to other embodiments. The foregoing description is presented to enable any person skilled in the art to make and use this application. In the foregoing description, various details are set forth for purposes of explanation. It will be apparent to those skilled in the art that the application may be practiced without these specific details. In other instances, well-known structures and processes are not set forth in detail to avoid obscuring the description of this application with unnecessary detail. Accordingly, this application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
Claims
1. A method for determining convolution input data, characterized in that, the method for determining convolution input data includes: determining convolution operation input data corresponding to the current operation result; the current operation result is the operation result of the data in the i-th row and j-th column of the convolution input data of the current layer convolution operation; storing the convolution results of k - s rows before the current operation result into BT; k is the dimension of the convolution input data, s is the step size of one sliding in the convolution operation, and BT is a preset first data cache structure; storing the convolution results of k - s columns before the current operation result in the BT into BL, so as to determine the convolution input data of the next layer of convolution operation according to the BT and the BL; BL is a preset second data cache structure.
2. The method for determining convolution input data according to claim 1, characterized in that, the size of the convolution input data is k×k, the size of BT is (k - s)×(IMGW + k - s), the size of BL is k×(k - s), and IMGW is the number of columns of the image to be processed; the storing the convolution results of k - s rows before the current operation result into BT includes: storing the convolution results of the i - (k - s) to i - s rows of operation results into BT; the storing the convolution results of k - s columns before the current operation result in the BT into BL includes: storing the operation results of the j - (k - s) to j - s columns in the BT into BL.
3. The method for determining convolution input data according to claim 2, characterized in that, the determining the convolution input data of the next layer of convolution operation according to the BT, the BL and the current operation result includes: determining k - s column data of the convolution input data according to the BL; determining k - s data of one column other than the k - s columns of the convolution input data according to the data in the j-th column of the BT; determining other undetermined data in the convolution input data according to the current operation result.
4. The method for determining convolution input data according to claim 1, characterized in that, when obtaining the operation result of the first row, the method for storing data into the BT includes: determining the first storage position corresponding to the operation result of the first row in the BT, writing the operation results of each row of the first row into the corresponding storage position in turn, and writing fill values into other storage positions of the BT.
5. The method for determining convolution input data according to claim 1, characterized in that, after storing the convolution results of k - s columns before the current operation result in the BT into BL, the method for determining convolution input data further includes: replacing the oldest column of data in the BL with a column of data where the current operation result is located in the convolution input data.
6. The method for determining convolution input data according to claim 1, characterized in that, when obtaining the operation result of one row, the updating method of the BT includes: replacing the oldest s rows of data in the BT with the latest s rows of operation results.
7. The method for determining convolution input data according to claim 1, characterized in that, when the current operation result is located in the first column, the updating method of the BL includes: Obtain the corresponding column of the current operation result in the BT to obtain a first storage column; Obtain the corresponding column of the current operation result in the BL to obtain a second storage column; Determine the storage position of the current operation result in the second storage column to obtain a second storage position; Write the data of the first storage column into other storage positions in the second storage column except the second storage position, write padding values into other storage columns in the BL except the second storage column, and write the second storage position into the current operation result.
8. The method for determining convolution input data according to claim 2, wherein, the number of rows of the image to be processed is IMGH; the size of the entire convolution operation result for the next-layer convolution operation is (IMGH + k - s) × (IMGW + k - s), including a feature map located at the middle position and padding values located around, and the size of the feature map is IMGH × IMGW.
9. A convolution operation method, wherein, comprises: Determine convolution input data by using the method for determining convolution input data according to any one of claims 1 to 8; Perform a convolution operation on a convolution object by using the convolution input data.
10. An image denoising method, wherein, comprises: Perform at least one convolution operation on the image to be processed by using the convolution operation method according to claim 9 to remove noise from the image to be processed.
11. The image denoising method according to claim 10, wherein, the performing at least one convolution operation on the image to be processed by using the convolution operation method according to claim 9 comprises: Vertically segment the image to be processed to obtain a plurality of sub-images; Perform at least one convolution operation on each sub-image by using the convolution operation method according to claim 9 to obtain denoised sub-images corresponding to the respective sub-images; Stitch the denoised sub-images to obtain a denoised image corresponding to the image to be processed.
12. A system for determining convolution input data, wherein, comprises: A first determination module for determining convolution operation input data corresponding to a current operation result; The current operation result is the operation result of the data in the i-th row and j-th column of the convolution input data of the current-layer convolution operation; A first storage module for storing the convolution results of k - s rows before the current operation result into the BT; k is the dimension of the convolution input data, s is the step size of one sliding in the convolution operation, and the BT is a preset first data cache structure; A second storage module for storing the convolution results of k - s columns before the current operation result in the BT into the BL to determine the convolution input data of the next-layer convolution operation according to the BT and the BL; the BL is a preset second data cache structure.
13. An accelerator, wherein, comprises an acceleration circuit; the acceleration circuit is used to execute the method for determining convolution input data according to any one of claims 1 to 8 or the convolution operation method according to claim 9.
14. An image processing chip, wherein, comprises the accelerator according to claim 13.
Citation Information
Patent Citations
Convolution operation method and device, computer apparatus, and computer readable storage medium
CN109543139A
Data processing method, related equipment and computer storage medium
CN110728351A