Sensor end image denoising method, device and equipment based on DoS system and medium
By adopting the DoS system-based image denoising method on the sensor side, and using simplified EDU network and quantitative perception training, the problems of resource limitation on the sensor side and DNN algorithm calculation complexity are solved, and the low-power and efficient image denoising effect is achieved, providing a new technical path for intelligent processing on the sensor side.
Patent Information
- Application Number
- CN202510043328.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
Due to the power and area resource limitations of the sensor end, the traditional DNN-based denoising algorithm has high computational complexity and is difficult to deploy directly, resulting in the denoising module being unable to move upstream of the ISP. The system needs to consume more resources when processing noise, and the denoising effect has not been significantly improved.
Using the sensor-side image denoising method based on the DoS system, the DoS macro unit adopts a pipeline mode in the preset EDU network to perform image denoising processing on the original image data. The EDU network reduces the number of layers and concat connections by the initial U-Net network, and quantizes the weights of the network using quantization perception training, introducing a lookup table module to replace inter-layer operations.
The denoising task was successfully moved up to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source, realizing a low-power and efficient image denoising solution, providing a new technical path for intelligent processing at the sensor end.
Smart Images

Figure CN119991477A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing and computing architecture, and in particular to a sensor-side image denoising method, device, equipment and medium based on a DoS system. Background Art
[0002] Noise is one of the main factors that lead to image quality degradation, and has a significant impact on the accuracy of subsequent high-level tasks (such as classification, segmentation, etc.). Therefore, denoising has become one of the core issues in low-level image processing.
[0003] In recent years, with the rapid development of deep neural network (DNN) technology, DNN-based denoising algorithms have made significant progress. With the powerful feature extraction capabilities of deep learning, these algorithms can effectively model complex noise distributions and significantly improve image denoising effects. However, since DNN-based denoising algorithms usually require high computing power and large storage space, the denoising module is often placed after the sensor end for execution. Although this deployment method simplifies the design complexity, it increases the processing complexity and leads to more resource overhead.
[0004] The core computing method of DNN is Multiply-And-Accumulate (MAC), which generates a large amount of data flow. The von Neumann architecture used in traditional digital integrated circuits requires frequent data movement between the memory and the central processing unit, which severely limits the energy efficiency and processing speed of the system. In contrast, the analog computing-in-memory (CIM) architecture can significantly reduce data movement by integrating the computing unit with the storage unit. The CIM architecture has the advantages of simple structure, low energy consumption, and high parallelism. It has shown significant advantages in terms of area, energy efficiency, and speed. Therefore, it has been widely studied in the field of near-sensor computing (NSC).
[0005] However, existing DNN denoising algorithms usually have high computational complexity and are difficult to deploy directly in the CIM (Compute-in-Memory) architecture. This technical bottleneck has resulted in no effective solution to combine DNN denoising algorithms with CIM to achieve efficient image denoising at the sensor end. Summary of the invention
[0006] The present application provides a sensor-side image denoising method, apparatus, device and medium based on a DoS system, thereby solving the problem that sensors are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to be directly deployed on the edge chip, resulting in the denoising module being unable to be moved upstream of the ISP, and the system needing to consume more resources when processing noise, while the denoising effect is not significantly improved.
[0007] The first aspect of the present application provides a sensor-side image denoising method based on a DoS system, wherein the DoS system includes a DoS macro unit, and includes the following steps: receiving raw image data collected by a sensor; based on a preset EDU network, using the DoS macro unit to perform image denoising on the raw image data using a preset pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of an initial U-Net network, quantizing the weights of the initial U-Net network using quantization-aware training, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
[0008] Optionally, the preset EDU network includes only twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network retains only two concat connections.
[0009] Optionally, the DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit uses a preset pipeline mode to perform image denoising on the original image data based on a preset EDU network, the method includes: using the weight loader to load the weight of the current layer from the off-chip memory during the chip initialization phase, and storing the weight in the processing unit; using the processing unit to complete the simulated MAC calculation, and transmitting the MAC calculation result to the lookup table module; using the lookup table The module loads boundary values and output values during the chip initialization phase, and uses the data bus to transfer the inter-layer operation output result from the previous layer or the feature map of the off-chip memory to the input buffer for storage, uses the processing unit to perform multiplication-accumulation MAC calculation on the weights of the current layer and the feature map, and passes the MAC calculation result to the lookup table module, so that the lookup table module can perform inter-layer operations on the MAC calculation result according to the boundary value and the output value; and uses the data bus to transfer the inter-layer operation output result of the lookup table module to the next layer or the off-chip memory.
[0010] Optionally, the above-mentioned sensor-side image denoising method based on the DoS system further includes: determining whether a transposed convolution operation instruction is received; if the transposed convolution operation instruction is received, using the insertion filling module to insert zeros around the input buffer data point by point to perform a transposed convolution operation.
[0011] Optionally, the above-mentioned sensor-side image denoising method based on the DoS system also includes: determining whether a maximum pooling operation instruction is received; if the maximum pooling operation instruction is received, using the maximum pooling module to perform maximum pooling on the feature map after the multiplication-accumulation MAC calculation is completed, and passing the feature map after maximum pooling to the lookup table module.
[0012] The second aspect of the present application provides a sensor-side image denoising device based on a DoS system, including: a receiving module, used to receive raw image data collected by a sensor; a processing module, used to perform image denoising on the raw image data based on a preset EDU network, using the DoS macro unit and a preset pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of an initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
[0013] Optionally, the preset EDU network includes only twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network retains only two concat connections.
[0014] Optionally, the DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit uses a preset pipeline mode to perform image denoising on the original image data based on a preset EDU network, the processing module is further used to: use the weight loader to load the weight of the current layer from the off-chip memory during the chip initialization phase, and store the weight in the processing unit; use the processing unit to complete the simulated MAC calculation, and transmit the MAC calculation result to the lookup table module; use the The lookup table module loads boundary values and output values during the chip initialization phase, and uses the data bus to transfer the inter-layer operation output result from the previous layer or the feature map of the off-chip memory to the input buffer for storage, uses the processing unit to perform multiplication-accumulation MAC calculation on the weights of the current layer and the feature map, and passes the MAC calculation result to the lookup table module, so that the lookup table module can perform inter-layer operations on the MAC calculation result according to the boundary value and the output value; and uses the data bus to transfer the inter-layer operation output result of the lookup table module to the next layer or the off-chip memory.
[0015] Optionally, the above-mentioned sensor-end image denoising device based on the DoS system further includes: a first judgment module, used to determine whether a transposed convolution operation instruction is received; a filling module, used to, if the transposed convolution operation instruction is received, use the inserted filling module to insert zeros around the input buffer data point by point to perform a transposed convolution operation.
[0016] Optionally, the above-mentioned sensor-end image denoising device based on the DoS system also includes: a second judgment module, used to determine whether a maximum pooling operation instruction is received; a pooling module, used to, if the maximum pooling operation instruction is received, use the maximum pooling module to perform maximum pooling on the feature map after the multiplication-accumulation MAC calculation is completed, and pass the feature map after maximum pooling to the lookup table module.
[0017] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the sensor-side image denoising method based on the DoS system as described in the above embodiment.
[0018] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the sensor-side image denoising method based on the DoS system as described in the above embodiment.
[0019] In the above implementation, the original image data collected by the sensor is received; based on the preset EDU network, the DoS macro unit is used to perform image denoising on the original image data in pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network. Thus, the problem that sensors are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to deploy directly on edge chips, resulting in the denoising module being unable to be moved upstream of the ISP, and the system needs to consume more resources when processing noise, but the denoising effect is not significantly improved. The denoising task is successfully moved to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source. This method realizes a low-power and efficient image denoising solution, and provides a new technical path for intelligent processing on the sensor end.
[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0022] Figure 1 It is a schematic diagram of the U-Net network structure in the related technology;
[0023] Figure 2 A flowchart of a sensor-side image denoising method based on a DoS system according to an embodiment of the present application;
[0024] Figure 3 Schematic diagram of the Nb2Nb network structure according to an embodiment of the present application;
[0025] Figure 4 A schematic diagram of an EDU network structure according to an embodiment of the present application;
[0026] Figure 5 A schematic diagram of concat storage overhead according to an embodiment of the present application;
[0027] Figure 6 A schematic diagram of a DoS hardware system framework according to an embodiment of the present application;
[0028] Figure 7 (a) is a schematic diagram of the structure of the traditional transposed convolution. Figure 7 (b) Schematic diagram of implementing transposed convolution function with insertion padding and traditional convolution;
[0029] Figure 8 A schematic diagram of a hardware system-level resource consumption indicator of an EDU according to an embodiment of the present application;
[0030] Fig. 9 This is an example diagram of a sensor-side image denoising device based on a DoS system according to an embodiment of the present application;
[0031] Fig.10 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0033] The following describes the sensor-side image denoising method, device, equipment and medium based on the DoS system of the embodiment of the present application with reference to the accompanying drawings. In view of the fact that the sensors mentioned in the above background technology are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to be directly deployed on the edge chip, resulting in the denoising module being unable to be moved upstream of the ISP, and the system needs to consume more resources when processing noise, but the denoising effect has not been significantly improved. The present application provides a sensor-side image denoising method based on the DoS system, in which the original image data collected by the sensor is received; based on the preset EDU network, the DoS macro unit is used to perform image denoising on the original image data in a pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network. As a result, the problem that sensors are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to deploy directly on edge chips, resulting in the denoising module being unable to be moved upstream of the ISP. The system needs to consume more resources when processing noise, but the denoising effect is not significantly improved is solved. The denoising task is successfully moved to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source. This method realizes a low-power and efficient image denoising solution, and provides a new technical path for intelligent processing on the sensor end.
[0034] Before specifically introducing the sensor-side image denoising method based on the DoS system of the present application, the related technical solutions are briefly introduced.
[0035] Technical solution of prior art 1:
[0036] Through a U-shaped network structure, such as Figure 1 As shown in the figure, it consists of a symmetrical encoder-decoder, and uses skip connections (concatenation, concat) to fuse high-resolution features with context information, thereby significantly improving the accuracy and robustness of the network. U-Net performs well in tasks such as equal resolution and super resolution, and is widely used in the ISP field due to its efficiency and ease of use.
[0037] Technical solution of the second prior art:
[0038] Since it is difficult and expensive to collect enough noisy images and their corresponding clean images, this work proposes a self-supervised learning method based on the Noise2Noise method that relies only on noisy images to train a denoising network. Specifically, this method estimates a noise-free image using two or more noisy images, and its core network uses the U-Net structure as the backbone network, which significantly reduces the denoising task's dependence on labeled data.
[0039] Technical solution of prior art three:
[0040] This solution proposes a 384kb SRAM-CIM unit based on 28nm process, designed for AI edge chips, capable of 8-bit precision MAC operations and providing 20-bit output close to full precision. The architecture introduces a segmented bit line charge sharing scheme, a source injection local multiplication unit, and a priority hybrid ADC, which significantly improves the reliability of the system while optimizing energy efficiency and performance. The CIM unit achieves high energy efficiency (22.75TOPS / W), has good scalability, and exhibits strong anti-variation capabilities under process deviation conditions, making it very suitable for energy-efficient AI edge application scenarios.
[0041] During image acquisition, noise runs through the entire ISP process, including the propagation of optical signals in the natural environment, the conversion from optical to electronic signals, and all subsequent electronic processing stages. According to the Freese noise coefficient formula, that is, formula (1), in a multi-stage system, the noise generated by the components close to the front end has the most significant impact on the overall system, where F i and G i They represent the noise coefficient and available power gain of the i-th stage respectively. In addition, if the noise is not processed in the early stage, as it propagates in the system, the denoising problem will become more complicated and the resource consumption will increase significantly. Therefore, the image denoising module should be deployed as close to the sensor as possible, or even integrated with the sensor.
[0042]
[0043] However, sensors are usually limited by power and area resources, and the high computational complexity of traditional DNN-based denoising algorithms makes it difficult to deploy these algorithms directly on edge chips. Therefore, the denoising module cannot be moved upstream of the ISP, causing the system to consume more resources (such as power, time, etc.) when processing noise, but the denoising effect has not been significantly improved.
[0044] In order to solve the above problems, this application provides a design method for deploying a denoising algorithm on the sensor side through the collaborative design of algorithm and hardware, thereby improving the efficiency and effect of the denoising task.
[0045] Specifically, this application designs a simplified Neighbor2Neighbor (Nb2Nb) encoder-decoder (Encoder-Decoder U-Net, EDU) algorithm. While retaining the performance of the original model, the algorithm significantly reduces resource consumption by streamlining the network structure, and is highly coupled with the computing and storage integration (CIM) architecture to optimize the deployment efficiency on edge devices.
[0046] In addition, the present application proposes a flexible and scalable near-sensor CIM unit, called DoS unit (Denoiseon Sensor), which supports multiple non-MAC operations, including plug-in zero filling, maximum pooling, and look-up table module operations (Look-up Table, LUT). This design expands the functionality of the traditional CIM architecture, enabling it to fully support the operators required by the EDU algorithm.
[0047] By combining the above algorithm with hardware design, this application successfully moved the denoising task to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source. This method realizes a low-power and efficient image denoising solution, providing a new technical path for intelligent processing on the sensor end.
[0048] Specifically, Figure 2 A schematic flow chart of a sensor-side image denoising method based on a DoS system provided in an embodiment of the present application.
[0049] like Figure 2 As shown, the sensor-side image denoising method based on the DoS system, the DoS system includes a DoS macro unit, and includes the following steps:
[0050] In step S201, raw image data collected by a sensor is received.
[0051] In step S202, based on the preset EDU network, the DoS macro unit adopts the preset pipeline mode to perform image denoising on the original image data to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
[0052] Optionally, in some embodiments, the preset EDU network includes only twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network retains only two concat connections.
[0053] Optionally, in some embodiments, the DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit uses a preset pipeline mode to perform image denoising on the original image data based on a preset EDU network, it includes: using the weight loader to load the weight of the current layer from the off-chip memory in the chip initialization stage, and storing the weight in the processing unit; using the processing unit to complete the simulated MAC calculation, and transmitting the MAC calculation result to the lookup table module; using the lookup table module to load the boundary value and output value in the chip initialization stage, and using the data bus to transfer the inter-layer operation output result from the previous layer or the feature map of the off-chip memory to the input buffer for storage, using the processing unit to perform a multiplication-accumulation MAC calculation on the weight and feature map of the current layer, and passing the MAC calculation result to the lookup table module, so that the lookup table module can perform an inter-layer operation on the MAC calculation result according to the boundary value and the output value; using the data bus to transmit the inter-layer operation output result of the lookup table module to the next layer or the off-chip memory.
[0054] Specifically, this application makes the following four major improvements to the Nb2Nb network to obtain the preset EDU network:
[0055] (1) Reduce the number of layers
[0056] The traditional Nb2Nb network architecture is as follows Figure 3 As shown in the figure. Although its overall structure is relatively simple, its depth (number of layers) is still relatively large. From the perspective of hardware deployment, each layer of convolution calculation will bring additional computing, storage requirements, data movement, and running time overhead, resulting in high resource consumption of the original Nb2Nb network, which is difficult to adapt to the restrictions of edge chips.
[0057] Through the algorithm and hardware co-design experiment, this application found that several intermediate small layers in the Nb2Nb network can be deleted while maintaining high accuracy. This streamlined design significantly reduces the demand for computing and storage resources. Therefore, the optimized EDU network architecture is as follows Figure 4 As shown, it only includes twelve layers of convolution, three layers of maximum pooling, and three layers of transposed convolution, which improves the efficiency of hardware implementation and makes it more suitable for deployment in resource-constrained edge devices.
[0058] (2) Reduce concat times
[0059] Although the streamlined Nb2Nb network architecture is relatively simple, the concat operation introduces additional storage overhead, especially in hardware deployment, where storage becomes a significant resource bottleneck. Figure 5 As shown in the figure, in a network with three layers of convolution, assuming there is a concat operation, the size of the input image is H×W×C, and each layer of convolution requires one clock cycle. Due to the existence of the concat operation, the first feature map of the input image needs to be stored from T1 to T4 before it can be released. Therefore, before the data stream is connected, the total storage requirement will accumulate and eventually reach H×W×C×3 (where 3 represents the storage time). Figure 3 As shown in Figure 1, the traditional Nb2Nb network contains multiple concat operations, resulting in a large storage space requirement. This application analyzes in detail the impact of each concat operation on algorithm accuracy and hardware resource consumption through algorithm and hardware co-design. In the EDU network, only the two concat connections that are most critical to performance are retained, and other concat operations are deleted, thereby significantly reducing storage space requirements and optimizing hardware resource allocation. Figure 5 shown.
[0060] (3) Weight Quantization
[0061] Compared with digital computing, analog computing has significant advantages in high efficiency and low power consumption. The analog CIM architecture performs MAC calculations through analog circuits, providing high-speed, low-energy computing capabilities. However, analog computing faces challenges such as signal margin and is less flexible than digital computing, resulting in all numerical values in the calculation having to be represented by low-bit-width integers. For DNN-based denoising algorithms, this means that the feature maps, weights, biases and other numerical values in the network must be represented by low bit widths.
[0062] This application uses quantization-aware training (QAT) to quantize the EDU network and optimize all values to 4 bits. This improvement enables the EDU network to be efficiently deployed on edge chips that simulate the CIM architecture, which not only significantly improves computing efficiency, but also effectively reduces power consumption, and further optimizes the performance and adaptability of hardware implementation.
[0063] (4) Look-up table module (LUT)
[0064] In the traditional DNN-based denoising algorithm, some computational operations (such as activation functions, rescaling, or batch normalization (Batch Normalization, BN)) are usually required between layers. In traditional digital circuit systems, high-precision calculators (such as floating-point calculators or high-bit-width integer calculators) are usually used to handle inter-layer calculations. However, this approach has two significant disadvantages: first, additional high-precision calculators are required, which increases hardware complexity; second, additional data movement and memory access are introduced, further exacerbating resource consumption. For resource-constrained edge chips, these two disadvantages are unacceptable.
[0065] Based on the mapping characteristics of inter-layer calculations, this application abstracts them into a simple mapping relationship from the integer output of the previous layer to the integer input of the next layer, without paying in-depth attention to the specific calculation process. To this end, it is proposed to replace all inter-layer operations with a look-up table module (LUT). Since both ends of this mapping are composed of discrete and finite values, LUT can seamlessly replace inter-layer operations while ensuring that the calculation accuracy is not affected.
[0066] By introducing LUT, the need for high-precision calculators is eliminated and a large amount of data movement and memory access is effectively avoided. This design significantly reduces the use of hardware resources, including computational complexity and storage requirements, making it more suitable for low-power deployment scenarios of edge devices.
[0067] Therefore, the present application simplifies the Neighbor2Neighbor (Nb2Nb) network, making the network more concise and lower in resource consumption. The optimized network is called Edge Denoise U-Net (ie, EDU network).
[0068] Furthermore, hardware system design based on CIM.
[0069] Based on the above algorithm design, this application proposes a CIM-based hardware system, called DoS system, which can efficiently support all operations in the EDU architecture. The overall framework of the DoS system is as follows: Figure 6 shown.
[0070] In the DoS system, the images collected by the sensor are transmitted to a series of DoS macro units (Macro) through the bus, and the denoised image is generated as the output after network inference. In order to optimize performance, the DoS system adopts a pipeline mode instead of a layer-by-layer mode. The advantage of the pipeline mode is that it allows part of the calculation results of the previous layer to be directly passed to the input buffer (IB) of the next layer, without the need to store the intermediate results in the off-chip memory. This not only significantly reduces storage requirements and energy consumption, but also avoids frequent interactions with off-chip memory. At the same time, the pipeline mode can also start the calculation of the next layer when the calculation of the previous layer is not completed, thereby improving the overall reasoning speed.
[0071] Among them, the DoS macro unit is the core component of the system, responsible for performing single-layer operations based on convolution. The advantage of this modular design is that it reduces storage requirements while giving the system a high degree of flexibility and scalability. By flexibly combining DoS macro units with different functions, the system can efficiently support a variety of convolution-based neural network architectures.
[0072] exist Figure 6 In the figure, the black arrows indicate the data flow of the main functions inside the DoS macrocell. In addition, the system also contains two optional data flows (indicated by blue dashed lines and cyan dashed lines) to expand the functionality to support all operator types in the EDU.
[0073] The main data flow of the DoS macro unit completes convolution and inter-layer calculations. In the chip initialization stage, the weight loader (Weight Loader, WL) loads the weight of the current layer from the off-chip memory through the data bus and stores it in the storage unit of the processing unit (Processing Engines, PEs). The processing unit completes the analog MAC calculation and transmits the MAC calculation result to the lookup table module. At the same time, the boundary value and corresponding output value of the current layer are transmitted to the lookup table module (LUT) using the data bus to provide support for subsequent inter-layer operations. In the inference stage, the input buffer (IB) uses the data bus to receive the output from the previous layer or the feature map of the off-chip memory. PEs is based on the charge sharing CIM design, performs MAC calculations, and the MAC calculation results are passed to the LUT. The lookup table module performs inter-layer operations on the MAC calculation results according to the boundary value and output value. The lookup table module LUT of this application adopts the range-addressable lookup table module (Range-Addressable LUT, RALUT) design, which can efficiently complete inter-layer operations such as activation functions and rescaling. The output of the LUT is transmitted to the next layer or off-chip memory via the data bus.
[0074] Optionally, in some embodiments, the above-mentioned sensor-side image denoising method based on the DoS system further includes: determining whether a transposed convolution operation instruction is received; if a transposed convolution operation instruction is received, using an inserted filling module to insert zeros around the input buffer data point by point to perform a transposed convolution operation.
[0075] Optionally, in some embodiments, the above-mentioned sensor-side image denoising method based on the DoS system further includes: determining whether a maximum pooling operation instruction is received; if a maximum pooling operation instruction is received, using the maximum pooling module to perform maximum pooling on the feature map after the multiplication-accumulation MAC calculation is completed, and passing the feature map after maximum pooling to the lookup table module.
[0076] By analyzing the EDU network architecture (such as Figure 4 As shown in the figure, it can be found that in addition to convolution and inter-layer calculation, the EDU network also includes transposed convolution and maximum pooling operations. For transposed convolution, this application introduces an insert padding module, which inserts zeros around the input buffer data point by point, so that the conventional convolution operation can realize the function of transposed convolution. The traditional transposed convolution is as follows: Figure 7 As shown in (a), the transposed convolution is realized by inserting the padding module and the traditional convolution. Figure 7 As shown in (b), this not only allows the existing convolution modules to be reused without the need to design additional dedicated transposed convolution hardware, but also significantly reduces hardware area and power consumption, thereby achieving efficient resource utilization of transposed convolution.
[0077] For the maximum pooling operation, since it always appears between MAC calculation and inter-layer calculation, this application designs an optional maximum pooling module after the processing unit PEs completes the MAC calculation, which is used to downsample the feature map. The downsampled feature map will be passed to the LUT, thereby further reducing the resolution and optimizing the complexity of subsequent calculations.
[0078] In summary, the DoS macro unit reuses the convolution module without the need to design dedicated transposed convolution or maximum pooling hardware, which greatly reduces the hardware area and power consumption of the overall system while maintaining flexibility and efficiency, allowing the EDU algorithm to be efficiently deployed on low-resource hardware.
[0079] The value of this application in actual deployment is as follows:
[0080] By combining the EDU algorithm with the DoS hardware architecture, significant experimental results have been achieved in image denoising and hardware performance.
[0081] Figure 8The hardware performance improvement of the DoS system after algorithm optimization is demonstrated. Through the coordinated optimization design of the algorithm and hardware, the power consumption and area of the hardware system are reduced by 31.7% and 36.0% respectively, and the energy efficiency is improved by 57.0%, while the algorithm performance (PSNR / SSIM) is only reduced by 1%. In addition, off-chip memory access, as one of the main sources of hardware power consumption and time overhead, has a significant impact on system performance. Through the design optimization of EDU, the concat connection in the network is reduced, and the number of off-chip memory accesses and storage requirements are significantly reduced, thereby further improving the adaptability and overall efficiency of hardware resources.
[0082] Image denoising, as a low-level task, usually occurs in the initial stage of the image processing system. Its main purpose is to provide high-quality data input for subsequent high-level tasks (such as classification, segmentation, etc.). In order to verify whether the denoising effect of EDU on edge devices is sufficient to support subsequent tasks, an experimental evaluation was conducted. After adding Gaussian noise with a standard deviation of σ=25 to the CIFAR-10 dataset, the EDU algorithm was applied to denoise it, and the classification performance of the official pre-trained ResNet-20 was evaluated on the clean CIFAR-10, noisy CIFAR-10, and denoised CIFAR-10 datasets, respectively. The comparison of the inference accuracy of Resnet-20 on Cifar-10 using different denoisers is shown in Table 1.
[0083] The experimental results in Table 1 show that although EDU will cause certain performance degradation when deployed to edge devices (especially during quantization), after completing the underlying denoising task, it can effectively improve the performance of subsequent high-level tasks (in this experiment, the high-level task is image classification). This proves that the combination of EDU and DoS system can not only meet the resource constraints of the edge environment, but also provide reliable data support for subsequent tasks.
[0084] Table 1
[0085] Dataset Top 1 accuracy (%) Top 5 accuracy (%) Clean CIFAR-10 91.73 99.66 Noise CIFAR-10 (Gaussian noise with σ=25) 25.20 80.28 Denoising CIFAR-10 (Nb2Nb) 80.47 98.95 Denoising CIFAR-10 (EDU) 78.36 98.88
[0086] In summary, the combination of the EDU algorithm and the DoS hardware architecture achieves efficient and low-power image denoising, while providing high-quality input data for subsequent tasks (such as super-resolution, image reconstruction, etc.). This verifies the practical feasibility and significant advantages of this application in edge device deployment.
[0087] Therefore, this application is designed for edge devices through a DNN-based denoising algorithm and CIM hardware architecture, which can efficiently complete the denoising task at the sensor end; a simplified U-Net encoder-decoder model is designed, which significantly reduces the computing resource requirements by reducing the number of network layers and concat connections, so that it can adapt to the resource limitations of edge devices; through a near-sensor CIM hardware architecture, it supports a variety of optional operation modules, including plug-in padding, maximum pooling and lookup table modules, and flexibly adapts to operations such as convolution, transposed convolution, activation function and rescaling; a modular design is adopted to add or delete DoS Macro units enable flexible expansion of network depth and functionality and support a variety of convolutional neural network architectures. Pipeline design optimizes data flow to avoid off-chip memory access on edge devices, greatly reducing storage requirements, energy consumption, and time overhead, while significantly improving computing efficiency. Plug-in padding modules are used to efficiently implement transposed convolution functions, and optional maximum pooling modules are added after convolution to reduce additional hardware resource usage and improve resource utilization. For DNN denoising algorithms, an optimization method based on quantization-aware training is proposed to compress weights, feature maps, and other numerical values into low-bit-width integer representations, thereby improving computing efficiency and reducing system energy consumption.
[0088] In summary, the beneficial effects of this application are as follows:
[0089] (1) Efficient edge deployment: The practical feasibility of the DNN-based denoising algorithm on edge devices was verified, demonstrating excellent resource efficiency and deployment performance;
[0090] (2) Noise processing close to the sensor: Moving the denoising module close to the sensor effectively reduces the complexity of noise propagation and achieves higher quality denoising with less resources.
[0091] (3) Resource utilization optimization: By reducing the need for off-chip memory access, optimizing memory management and data flow transmission, the system power consumption and hardware area occupied are greatly reduced, and the overall energy efficiency is improved;
[0092] (4) Flexible expansion capability: The modular design supports flexible expansion of network depth and functionality, and is compatible with a variety of convolution operations and inter-layer operators. It is widely applicable to a variety of convolutional neural network architectures, significantly enhancing the applicability of the system.
[0093] (5) Low power consumption and high energy efficiency: Combining the analog CIM architecture and quantitative optimization algorithm, it significantly reduces energy consumption, while improving computing parallelism and enhancing the adaptability of edge devices;
[0094] (6) Enhance the performance of subsequent tasks: Through efficient denoising, high-quality input data is provided for subsequent tasks such as classification, segmentation, and image reconstruction, thereby improving the overall performance of high-level tasks;
[0095] (7) Innovative advantages of combining algorithms with hardware: Combining the EDU algorithm with the hardware pipeline mode can fully utilize resource advantages and improve the computing efficiency and response speed of edge devices.
[0096] According to the sensor-side image denoising method based on the DoS system proposed in the embodiment of the present application, the original image data collected by the sensor is received; based on the preset EDU network, the DoS macro unit is used to adopt the pipeline mode to perform image denoising on the original image data to obtain the processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network. Thus, the problem that sensors are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to be directly deployed on edge chips, resulting in the denoising module being unable to be moved upstream of the ISP, and the system needs to consume more resources when processing noise, but the denoising effect has not been significantly improved. The denoising task is successfully moved to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source. This method realizes a low-power and efficient image denoising solution and provides a new technical path for intelligent processing at the sensor end.
[0097] Next, a sensor-side image denoising device based on a DoS system according to an embodiment of the present application is described with reference to the accompanying drawings.
[0098] Fig. 9 4 is a block diagram of a sensor-side image denoising device based on a DoS system according to an embodiment of the present application.
[0099] like Fig. 9 As shown, the sensor-side image denoising device 10 based on the DoS system includes: a receiving module 100 and a processing module 200.
[0100] Among them, the receiving module 100 is used to receive the original image data collected by the sensor; the processing module 200 is used to perform image denoising on the original image data based on the preset EDU network, using the DoS macro unit and adopting the preset pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
[0101] Optionally, in some embodiments, the preset EDU network includes only twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network retains only two concat connections.
[0102] Optionally, in some embodiments, the DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit uses a preset pipeline mode to perform image denoising on the original image data based on a preset EDU network, the processing module is also used to: use the weight loader to load the weight of the current layer from the off-chip memory during the chip initialization phase, and store the weight in the processing unit; use the processing unit to complete the simulated MAC calculation, and transmit the MAC calculation result to the lookup table module; use the lookup table module to load the boundary value and output value during the chip initialization phase, and use the data bus to transfer the inter-layer operation output result from the previous layer or the feature map of the off-chip memory to the input buffer for storage, use the processing unit to perform a multiplication-accumulation MAC calculation on the weight and feature map of the current layer, and pass the MAC calculation result to the lookup table module, so that the lookup table module can perform an inter-layer operation on the MAC calculation result according to the boundary value and the output value; use the data bus to transmit the inter-layer operation output result of the lookup table module to the next layer or the off-chip memory.
[0103] Optionally, in some embodiments, the above-mentioned sensor-end image denoising device 10 based on the DoS system further includes: a first judgment module, used to determine whether a transposed convolution operation instruction is received; a filling module, used to insert zeros around the input buffer data point by point using the inserted filling module to perform a transposed convolution operation if a transposed convolution operation instruction is received.
[0104] Optionally, in some embodiments, the above-mentioned sensor-side image denoising device 10 based on the DoS system further includes: a second judgment module, used to determine whether a maximum pooling operation instruction is received; a pooling module, used to use the maximum pooling module to perform maximum pooling on the feature map after the multiplication-accumulation MAC calculation is completed if the maximum pooling operation instruction is received, and pass the feature map after maximum pooling to the lookup table module.
[0105] It should be noted that the aforementioned explanation of the embodiment of the sensor-side image denoising method based on the DoS system is also applicable to the sensor-side image denoising device based on the DoS system of this embodiment, and will not be repeated here.
[0106] According to the sensor-side image denoising device based on the DoS system proposed in the embodiment of the present application, the original image data collected by the sensor is received; based on the preset EDU network, the DoS macro unit is used to adopt the pipeline mode to perform image denoising on the original image data to obtain the processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network. Thus, the problem that sensors are usually limited by power and area resources, and the traditional DNN-based denoising algorithm has high computational complexity and is difficult to be directly deployed on edge chips, resulting in the denoising module being unable to be moved upstream of the ISP, and the system needs to consume more resources when processing noise, but the denoising effect is not significantly improved. The denoising task is successfully moved to the sensor end, reducing the impact of noise propagation on the subsequent processing of the system from the source. This method realizes a low-power and efficient image denoising solution and provides a new technical path for intelligent processing at the sensor end.
[0107] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0108] A memory 1001 , a processor 1002 , and a computer program stored in the memory 1001 and executable on the processor 1002 .
[0109] When the processor 1002 executes the program, the sensor-side image denoising method based on the DoS system provided in the above embodiment is implemented.
[0110] Furthermore, the electronic device further comprises:
[0111] The communication interface 1003 is used for communication between the memory 1001 and the processor 1002 .
[0112] The memory 1001 is used to store computer programs that can be executed on the processor 1002 .
[0113] The memory 1001 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0114] If the memory 1001, the processor 1002 and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001 and the processor 1002 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0115] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can communicate with each other through an internal interface.
[0116] The processor 1002 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0117] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned sensor-side image denoising method based on a DoS system.
[0118] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0119] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "special" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0120] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0121] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or N wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in another suitable manner if necessary, and then stored in a computer memory.
[0122] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0123] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0124] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0125] The computer-readable storage medium mentioned above may be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A sensor-side image denoising method based on a DoS system, characterized in that: The DoS system includes a DoS macro unit, comprising the following steps: Receive raw image data collected by the sensor; Based on the preset EDU network, the DoS macro unit is used to adopt a preset pipeline mode to perform image denoising on the original image data to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of the initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
2. The method according to claim 1, characterized in that: The preset EDU network only includes twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network only retains two concat connections.
3. The method according to claim 1, characterized in that The DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit performs image denoising on the original image data in a preset pipeline mode based on a preset EDU network, the DoS macro unit includes: Using the weight loader to load the weight of the current layer from the off-chip memory during the chip initialization phase, and storing the weight in the processing unit; Using the processing unit to complete the simulated MAC calculation, and transmitting the MAC calculation result to the lookup table module; The lookup table module is used to load boundary values and output values in the chip initialization stage, and the output result of the inter-layer operation from the previous layer or the feature map of the off-chip memory is transferred to the input buffer and then stored by the data bus, and the processing unit is used to perform a multiplication-accumulation MAC calculation on the weight value of the current layer and the feature map, and the MAC calculation result is passed to the lookup table module, so that the lookup table module is used to perform an inter-layer operation on the MAC calculation result according to the boundary value and the output value; The data bus is used to transmit the inter-layer operation output result of the lookup table module to the next layer or an off-chip memory.
4. The method according to claim 3, characterized in that: Also includes: Determine whether a transposed convolution operation instruction is received; If the transposed convolution operation instruction is received, the insertion filling module is used to insert zeros around the input buffer data point by point to perform a transposed convolution operation.
5. The method according to claim 3, characterized in that: Also includes: Determine whether a maximum pooling operation instruction is received; If the maximum pooling operation instruction is received, the maximum pooling module is used to perform maximum pooling on the feature map after the product-accumulation MAC calculation, and the feature map after maximum pooling is passed to the lookup table module.
6. A sensor-side image denoising device based on a DoS system, characterized in that: include: A receiving module, used for receiving raw image data collected by the sensor; The processing module is used to perform image denoising on the original image data based on a preset EDU network and using the DoS macro unit in a preset pipeline mode to obtain a processed image, wherein the preset EDU network is obtained by reducing the number of layers and concat connections of an initial U-Net network, using quantization-aware training to quantize the weights of the initial U-Net network, and introducing a lookup table module to replace all inter-layer operations in the initial U-Net network.
7. The device according to claim 6, characterized in that The preset EDU network only includes twelve layers of convolution, three layers of maximum pooling and three layers of transposed convolution, and the preset EDU network only retains two concat connections.
8. The device according to claim 6, characterized in that The DoS macro unit includes a weight loader, a lookup table module, an input buffer, a plug-in filling module, a maximum pooling module, a processing unit and a data bus. When the DoS macro unit performs image denoising on the original image data in a preset pipeline mode based on a preset EDU network, the processing module is further used to: Using the weight loader to load the weight of the current layer from the off-chip memory during the chip initialization phase, and storing the weight in the processing unit; Using the processing unit to complete the simulated MAC calculation, and transmitting the MAC calculation result to the lookup table module; The lookup table module is used to load boundary values and output values in the chip initialization stage, and the output result of the inter-layer operation from the previous layer or the feature map of the off-chip memory is transferred to the input buffer and then stored by the data bus, and the processing unit is used to perform a multiplication-accumulation MAC calculation on the weight value of the current layer and the feature map, and the MAC calculation result is passed to the lookup table module, so that the lookup table module is used to perform an inter-layer operation on the MAC calculation result according to the boundary value and the output value; The data bus is used to transmit the inter-layer operation output result of the lookup table module to the next layer or an off-chip memory.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the sensor-side image denoising method based on the DoS system as described in any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the sensor-side image denoising method based on a DoS system as described in any one of claims 1 to 5.