Mobile device image denoising method based on double-branch residual sparse network
Through the serial and parallel structure of the dual-branch residual sparse network, combined with the residual sparse module and the attention-guided residual sparse module, the complexity and performance balance of the image denoising model on small mobile devices is solved, and efficient image denoising effect and detail retention are achieved.
Patent Information
- Application Number
- CN202510415977.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing image denoising models are difficult to balance between performance and model complexity, making it difficult to scale and apply to multiple image processing scenarios on small mobile devices.
The image denoising method based on a double-branch residual sparse network is adopted. Through the upper and lower branch network and feature fusion module with a series-parallel structure, combined with the residual sparse module and the attention-guided residual sparse module, the hybrid expansion convolution and attention mechanism is used to capture multi-scale features and reduce the calculation amount.
While reducing the number of network parameters, it maintains excellent denoising performance and improves detailed retention capabilities, and is suitable for mobile devices with limited resources.
Smart Images

Figure CN120339105A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and an image denoising method based on a double-branch residual sparse network with a series-parallel structure, which is applicable to resource-constrained mobile devices such as unmanned aerial vehicles and mobile robots. Background Art
[0002] Image denoising is an important task in the field of image processing, aiming to reduce or eliminate noise in images while preserving image details and features as much as possible. With the rapid development of mobile communication technology and digital imaging sensors, mobile phones and cameras have become the most commonly used imaging devices in daily life, and the demand for high-quality images is increasing continuously. Therefore, it is particularly important to develop a denoising technology that can effectively process noisy images at low computational cost and obtain good visual reproduction quality.
[0003] In this context, significant progress has been made in image denoising technology. For example, the Block-Matching and 3D Filtering (BM3D) algorithm generates clearer images by grouping similar image blocks, applying 3D filtering to reduce noise, and averaging the results while preserving details. The BM3D-Net convolutional neural network inspired by the BM3D algorithm models by unfolding the BM3D computational flow and introducing "extraction" and "aggregation" layers. In addition, there are many other denoising algorithms, such as the Weighted Nuclear Norm Minimization algorithm, the multi-channel WNNM algorithm, the block sparse collaborative low-rank algorithm based on sparse and collaborative low-rank matrix decomposition, the three-dimensional magnetic resonance imaging denoising model, etc.
[0004] Although traditional denoising models perform well under specific conditions, they are usually limited by a large number of parameters. Denoising models based on deep neural networks improve performance by learning complex features and patterns, usually require fewer hyperparameters, and provide faster inference speeds, thus achieving better results in various denoising tasks. For example, the Deep Convolutional Neural Network improves denoising performance by learning the noise distribution in images and the residuals during the denoising process. The Fast and Flexible Denoising Network achieves efficient image denoising with low computational cost and a small model size.
[0005] However, as the depth of the network increases, the number of parameters and the operation time also increase accordingly, making it difficult to deploy the model in practical applications. To solve this problem, researchers have enhanced the representation ability of the denoising network by broadening the network structure, such as the Deep Convolutional Neural Network and the Batch Renormalization Deep Network. In addition, the role of the attention mechanism in image denoising has become increasingly prominent, enabling the model to better understand the image content, eliminate noise, and thus improve the image quality and clarity.
[0006] Although the existing deep learning networks perform well in image denoising, there are still some deficiencies. For larger network models, although the denoising effect is good, they have higher hardware requirements for devices, long running times, and large memory occupancy. For smaller network models, although the running speed and space occupancy can meet the device requirements, the denoising performance is relatively insufficient. Therefore, researching a network model that can meet the performance requirements and has a relatively low complexity is a challenging direction. Summary of the Invention
[0007] The purpose of the present invention is to propose a mobile device image denoising method based on a dual-branch residual sparse network to solve problems in the prior art such as the difficulty in balancing denoising performance and model complexity, the difficulty in expanding on small mobile devices, and the inability to be well applicable to various image processing scenarios.
[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0009] A mobile device image denoising method based on a dual-branch residual sparse network, comprising the following steps:
[0010] S1. Data preparation and preprocessing:
[0011] Use a mobile device to capture images in multiple scenarios, perform data annotation, enhancement on the images, and construct a mobile device image dataset. Then, preprocess the images in the dataset, and the preprocessing includes normalization, grayscale conversion, resolution adjustment, and data segmentation;
[0012] S2. Construct a dual-branch residual sparse denoising network based on a series-parallel structure:
[0013] The dual-branch residual sparse denoising network includes an upper-branch network, a lower-branch network, and a feature fusion module. Among them, the upper-branch network is formed by connecting five residual sparse modules (RSB) in series through upsampling and downsampling, which is used to gradually extract image features and capture multi-scale information; the lower-branch network is composed of five attention-guided residual sparse modules (ARSB) connected in series; the feature fusion module is used to integrate the feature information of the upper-branch network and the lower-branch network, perform final residual learning, and output the denoised image.
[0014] S3. Design the loss function L of the dual-branch residual sparse denoising network as:
[0015]
[0016] Where and respectively represent the image after model denoising and the corresponding clean image; N represents the total number of images.
[0017] S4. The mobile device images obtained in S1 are divided into a training set and a test set through data segmentation. The training set images are input into the constructed dual-branch residual sparse denoising network, and the gradient of the loss function is calculated using the backpropagation method until the loss function tends to be stable, obtaining a trained dual-branch residual sparse denoising network model.
[0018] S5. Input the noisy images in the test set into the trained dual-branch residual sparse denoising network to obtain denoised images.
[0019] Preferably, each of the residual sparse modules (RSB) includes a standard convolution and a dilated convolution. The input feature and the output feature are added through a residual connection to avoid the problem of gradient disappearance and improve the training stability.
[0020] The attention-guided residual sparse module (ARSB) adds a channel attention and a pixel attention mechanism on the basis of the residual sparse module, enabling the network to focus on the key regions of the image and improving the denoising effect.
[0021] The residual sparse module (RSB) and the attention-guided residual sparse module (ARSB) use hybrid dilated convolution to increase the receptive field without increasing the computational amount; the dilated convolution expands the coverage range by introducing intervals in the convolution kernel without increasing the parameters, improving the receptive field and feature extraction ability of the model without increasing the computational amount, and realizing model lightweight.
[0022] Preferably, the construction method of the upper-branch network in S2 specifically includes:
[0023] 1) Use a standard convolution module composed of a standard convolution layer with a convolution kernel size of 3×3 and a stride of 1, batch normalization, and a ReLU activation function.
[0024] 2) Use a dilated convolutional layer with a convolutional kernel size of 3×3, a stride of 1, and dilation rates of 2 and 3, batch normalization, and ReLU activation function to form a dilated convolutional module;
[0025] 3) Connect a standard convolutional module, a dilated convolutional module with a dilation rate of 2, a standard convolutional module, and a dilated convolutional module with a dilation rate of 3 in sequence, and make a residual connection with the input to form a residual sparse module;
[0026] 4) Connect five residual sparse modules in sequence. The first three residual sparse modules are connected through downsampling, and the last three residual sparse modules are connected through upsampling to form an upper branch network;
[0027] The function of the upper branch network is expressed as follows:
[0028]
[0029] Among them, represents the output of each RSB module, i ∈ 1, 2, 3, 4, 5; O shallow represents the input of the upper branch network; RSB represents the residual sparse module; Down and Up represent downsampling and upsampling respectively; represents the output of the upper branch network of the deep feature extraction module.
[0030] Preferably, in the upper branch network, the residual sparse module expands the receptive field of the upper branch network by combining standard convolution and dilated convolution operations to capture multi-scale feature information, and the specific content is as follows:
[0031] First, extract image features through a standard convolutional layer, and then use dilated convolution (dilation rates of 2 and 3) to increase the receptive field to obtain more extensive context information while ensuring computational efficiency;
[0032] Apply a residual connection between the input and output of the module to alleviate the gradient vanishing problem in the deep network, promote information flow, and accelerate the training and convergence of the network. Its function is expressed as follows:
[0033] O RSB = RSB(f rsb ) + f rsb = DBR3(CBR(DBR2(CBR(f rsb )))) + f rsb
[0034] Among them, f rsbRSB represents the input of the module; CBR represents the standard convolution module; DBR2 represents the dilated convolution module with a dilation rate of 2; DBR3 represents the dilated convolution module with a dilation rate of 3. The number of channels of both the standard convolution and the dilated convolution is 64, and the convolution kernel is 3×3.
[0035] Preferably, the construction method of the lower branch network in S2 specifically includes:
[0036] 1) Connect the standard convolution module and the dilated convolution module with a dilation rate of 2 and perform a residual connection to form an internal residual sparse module;
[0037] 2) Connect the internal residual sparse module, the channel attention mechanism, the pixel attention mechanism, the standard convolution module, and the dilated convolution module with a dilation rate of 3 in sequence, and perform a residual connection with the input to form an attention-guided residual sparse module;
[0038] 3) Connect five attention-guided residual sparse modules in sequence to form the lower branch network;
[0039] The function of the lower branch network is expressed as follows:
[0040]
[0041] Among them, represents the output of each ARSB module, i ∈ 1, 2, 3, 4, 5; O shallow represents the input of the upper branch network; ARSB represents the attention-guided residual sparse module; represents the output of the lower branch network of the deep feature extraction module.
[0042] Preferably, the attention-guided residual sparse module in the lower branch network specifically includes the following:
[0043] First, connect the standard convolution module and the dilated convolution module with a dilation rate of 2 in sequence and form an internal residual sparse module through a residual connection; use the standard convolution to extract local features, while the dilated convolution captures more extensive context information by increasing the receptive field;
[0044] Next, connect the internal residual sparse module with the channel attention mechanism, the pixel attention mechanism, the standard convolution module, and the dilated convolution module with a dilation rate of 3 in sequence to form an attention-guided residual sparse module; in the attention-guided residual sparse module, the channel and pixel attention mechanisms adaptively assign weights to different features to enhance the network's attention to important features;
[0045] Finally, the output of the module is added to the initial input through a residual connection to ensure the efficient flow of information and strengthen the gradient propagation during training; the attention-guided residual sparse module combines multi-scale feature extraction, the attention mechanism, and residual connections to improve the performance and stability of the lower-branch network in complex tasks. Its functional representation is as follows:
[0046] O ARSB = ARSB(f arsb ) + f arsb = DBR3(CBR(CAB(PAB(DBR2(CBR(f arsb ))))) + f arsb ) + f arsb
[0047] where f arsb represents the input of the ARSB module; CAB represents the channel attention module; PAB represents the pixel attention module.
[0048] Preferably, the feature fusion module in S2 specifically includes the following:
[0049] Feature fusion input: The outputs of the upper-branch network and the lower-branch network are added together as the input of the feature fusion module for feature fusion;
[0050] Module composition: The feature fusion module consists of a standard convolution module, an attention-guided residual sparse module, two standard convolutions, and a Sigmoid activation function;
[0051] Feature fusion process:
[0052] 1) The input of the feature fusion module is sequentially passed through the standard convolution module, the attention-guided residual sparse module, and the first standard convolution;
[0053] 2) The above output is concatenated with the normalized noise image and sequentially passed through the Sigmoid activation function and the second standard convolution;
[0054] 3) The output of the first standard convolution is multiplied by the output of the second standard convolution, and finally, a residual is taken with the original noise image to obtain the final denoised image;
[0055] The functional representation of the feature fusion process is as follows:
[0056]
[0057] where f rb represents the input of the RB module; cat represents the concatenation operation; C represents the standard convolution; I n represents the input noise image; Sig represents the Sigmoid function.
[0058] Preferably, in S4, the gradient of the loss function is calculated using the backpropagation method to train the model. The specific implementation steps are as follows:
[0059] 1) Set the dual-branch residual sparse denoising network to the training mode;
[0060] 2) Input the noisy images in the training set into the dual-branch residual sparse denoising network, perform forward propagation, and calculate and output the fused denoised images;
[0061] 3) Calculate the loss between the model output and the original clean image according to the loss function;
[0062] 4) Calculate the gradient of the loss function according to the backpropagation method;
[0063] 5) Use the Adam optimizer to update the model parameters according to the gradient change of the loss function. The learning rate of the optimizer is set to 0.001, and the hyperparameters are β1 = 0.9 and β2 = 0.99;
[0064] 6) Repeat the above process until the loss function reaches stability, and obtain the trained dual-branch residual sparse denoising network model.
[0065] Based on the above method, a mobile device image denoising system applying this method is further proposed, including:
[0066] The Residual Sparse Block (RSB) contains standard convolutional blocks and dilated convolutional blocks to capture local and global features of the image. At the same time, batch normalization and ReLU activation functions are used to enhance the non-linear expression ability of the network, and the gradient vanishing problem in the training of deep networks is alleviated through residual connections, accelerating the training process and improving the learning efficiency of the model for image features.
[0067] The Attention-guided Residual Sparse Block (ARSB) contains standard convolutions and dilated convolutions to capture image features, achieves a balance between network depth and width through a combination of dilated convolutions and residual connections, effectively avoids the limitations of standard convolutions and the inefficiency of dilated convolutions, and at the same time introduces channel attention and pixel attention mechanisms to enhance the network's attention to important features, thereby improving the denoising effect.
[0068] The Residual Block (RB) contains a convolutional layer, an attention-guided module, and a skip connection. The convolutional layer is used to extract image features, the attention-guided module identifies and enhances important features in the image while suppressing noise, and the skip connection ensures that the input is directly passed to the output, enabling the network to learn the residual between the input and the output, thereby avoiding the gradient vanishing problem in the training of deep networks.
[0069] Compared with the prior art, the present invention has the following advantages:
[0070] (1) Balance between lightweight and high performance: Through a unique dual-branch series-parallel structure and optimized network design, the present invention significantly reduces the number of network parameters while maintaining excellent denoising performance.
[0071] (2) Enhancement of detail preservation and denoising ability: The present invention introduces residual sparse blocks and attention-guided residual sparse blocks, combines hybrid dilated convolution and attention mechanism, can extract the structural and texture information of images more comprehensively, and better preserves details during the denoising process. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings involved in the embodiments are briefly introduced below. Obviously, the accompanying drawings in the following description are only schematic illustrations of some embodiments of the present invention, and those skilled in the art can also construct other forms of drawings based on these drawings without creative efforts.
[0073] Figure 1 It is the overall flowchart of the mobile device image denoising method based on the dual-branch residual sparse network proposed by the present invention;
[0074] Figure 2 It is the schematic diagram of the dual-branch residual sparse denoising network structure constructed in Embodiment 1 of the present invention;
[0075] Figure 3 For the present invention Figure 2 It is the structural diagram of the residual sparse module in the present invention;
[0076] Figure 4 For the present invention Figure 2 It is the structural diagram of the attention-guided residual sparse module in the present invention;
[0077] Figure 5 For the present invention Figure 2 It is the block diagram of the feature fusion module structure in the present invention;
[0078] Figure 6 It is the comparison diagram of the denoising effects of the present invention and several existing denoising methods in a high-noise scenario. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.
[0080] The present invention proposes a mobile device image denoising method based on a dual-branch residual sparse network to balance denoising performance and model complexity. The present invention adopts a series-parallel structure, including a residual sparse module (RSB) and an attention-guided residual sparse module (ARSB), and is connected through upsampling and downsampling operations to effectively capture multi-scale information of the image. The RSB and ARSB modules use hybrid dilated convolutions to increase the receptive field without increasing the computational amount. Dilated convolution expands the coverage range by introducing gaps in the convolution kernel without increasing the number of parameters, thereby improving the receptive field and feature extraction ability of the model without increasing the computational amount, and thus realizing model lightweight. The purpose of the present invention is to remove various types of noise while retaining image details and edge information to be suitable for deployment and application on mobile devices.
[0081] The following further describes the embodiments and effects of the present invention in detail with reference to the accompanying drawings and specific examples.
[0082] Embodiment 1:
[0083] Refer to Figure 1 , the present invention proposes a mobile device image denoising method based on a dual-branch residual sparse network, and its implementation steps are as follows:
[0084] Step 1, acquisition and preprocessing of the dataset.
[0085] To improve the performance of the model of the present invention in the image denoising task, the Waterloo Exploration Database is used for training. This database contains 4,744 natural images. To improve the utilization rate of computing resources, according to the receptive field size of the invention model, the training images are randomly cropped into image patches of 128×128 size. For real noise elimination evaluation, the medium-sized SIDD (Smartphone Image Denoising Dataset) dataset is selected as the training set. This dataset contains 320 pairs of high-resolution noisy images and their corresponding clean images. The high-resolution images are cropped into image patches of 128×128 size, and the training samples are enhanced through rotation and flipping operations. In the model evaluation stage, 7 public datasets are used as the test set, including Set12 and BSD68 for grayscale images, CBSD68, Kodak24 and McMaster for color images, and the SIDD validation set and DND SRGB dataset for real images.
[0086] Step 2, construct an image denoising network based on a dual-branch residual sparse network with a series-parallel structure.
[0087] As Figure 2As shown in the figure, the image denoising network of the double-branch residual sparse network based on the series-parallel structure includes an upper branch network, a lower branch network, and a feature fusion module (RB).
[0088] 2.1) Construct the upper branch network:
[0089] As Figure 2 shown in the figure, the upper branch network is composed of five residual sparse modules connected in series. The first three residual sparse modules are connected through downsampling, and the last three residual sparse modules are connected through upsampling.
[0090] The formula of the upper branch network is as follows:
[0091]
[0092] Among them, is the output of each RSB module, O shallow represents the input of the upper branch network, RSB represents the residual sparse module, Down and Up respectively represent downsampling and upsampling, represents the output of the upper branch network of the deep feature extraction module.
[0093] As Figure 3 shown in the figure, the residual sparse module is composed of two standard convolution modules and two dilated convolution modules. The standard convolution module consists of a standard convolution layer with a convolution kernel size of 3×3 and a stride of 1, batch normalization, and a Relu activation function. The dilated convolution module consists of a dilated convolution layer with a convolution kernel size of 3×3, a stride of 1, and dilation rates of 2 and 3 respectively, batch normalization, and a Relu activation function. Connect the standard convolution module, the dilated convolution module with a dilation rate of 2, the standard convolution module, and the dilated convolution module with a dilation rate of 3 in sequence, and make a residual connection with the initial input of the module to form the residual sparse module. Its expression is as follows
[0094] O RSB = RSB(f rsb ) + f rsb = DBR3(CBR(DBR2(CBR(f rsb )))) + f rsb
[0095] Among them, f rsb represents the input of the RSB module; CBR represents the standard convolution module; DBR2 represents the dilated convolution module with a dilation rate of 2; DBR3 represents the dilated convolution module with a dilation rate of 3; the number of channels of the standard convolution and the dilated convolution is 64, and the convolution kernel is 3×3.
[0096] 2.2) Construct the lower branch network:
[0097] As Figure 2As shown, it is composed of five attention-guided residual sparse modules connected in series.
[0098] The formula of the lower branch network is as follows:
[0099]
[0100] Where represents the output of each ARSB module, O shallow represents the input of the upper branch network, ARSB represents the attention-guided residual sparse module, represents the output of the lower branch network of the deep feature extraction module.
[0101] As Figure 4 shown, the attention-guided residual sparse module is composed of two standard convolution modules, two dilated convolution modules, a channel attention mechanism, and a pixel attention mechanism. The standard convolution module consists of a standard convolution layer with a convolution kernel size of 3×3 and a stride of 1, batch normalization, and a Relu activation function. The dilated convolution module consists of a dilated convolution layer with a convolution kernel size of 3×3, a stride of 1, and dilation rates of 2 and 3 respectively, batch normalization, and a Relu activation function. First, the standard convolution module and the dilated convolution module with a dilation rate of 2 are connected and a residual connection is made to form an internal residual sparse module. Then, the internal residual sparse module, the channel attention mechanism, the pixel attention mechanism, the standard convolution module, and the dilated convolution module with a dilation rate of 3 are connected in sequence, and a residual connection is made with the initial input of the module to form the attention-guided residual sparse module. Its specific expression is as follows:
[0102] O ARSB = ARSB(f arsb ) + f arsb = DBR3(CBR(CAB(PAB(DBR2(CBR(f arsb )))))+ f arsb ) + f arsb
[0103] Where f arsb represents the input of the ARSB module, CAB represents the channel attention module, and PAB represents the pixel attention module.
[0104] 2.3) Construct the feature fusion module:
[0105] As Figure 5As shown in the figure, the feature fusion module consists of a standard convolution module, an attention-guided residual sparse module, two standard convolutions, and a Sigmoid activation function. The input of the feature fusion module is passed through the standard convolution module, the attention-guided residual sparse module, and the first standard convolution in sequence. After its output is concatenated with the normalized noise image, it is passed through the Sigmoid activation function and the second standard convolution in sequence. Then, the output of the first standard convolution is multiplied by the output of the second standard convolution, and finally, the residual is calculated with the original noise image to obtain the denoised image.
[0106] Its specific expression is as follows:
[0107]
[0108] Among them, f rb represents the input of the RB module; cat represents the concatenation operation; C represents the standard convolution; I n represents the input noise image; Sig represents the Sigmoid function.
[0109] Step 3, construct the loss function L of the dual-branch residual sparse denoising network based on the series-parallel structure: Use the MSE (Mean Squared Error) loss function as the loss function of the dual-branch residual sparse denoising network:
[0110]
[0111] Among them, and represent the denoised image of the model and the corresponding clean image respectively, and N is the total number of images;
[0112] Step 4, train the dual-branch residual sparse denoising network based on the series-parallel structure:
[0113] 4.1) Set the dual-branch residual sparse denoising network to the training mode;
[0114] 4.2) The noise images in the training set are propagated forward through the dual-branch residual sparse denoising network to calculate and output the fused denoised images;
[0115] 4.3) Calculate the loss between the model output and the original image according to the loss function;
[0116] 4.4) Calculate the gradient of the loss function according to the backpropagation method;
[0117] 4.5) Use the Adam optimizer to update the parameters of the model according to the gradient change of the loss function. The learning rate of the optimizer is set to 0.001, and the hyperparameters are β1 = 0.9 and β2 = 0.99;
[0118] 4.6) Repeat the process from 4.2) to 4.5) until the loss function reaches stability, and obtain the trained dual-branch residual sparse denoising network model.
[0119] Step 5, denoise the image using the trained denoising network:
[0120] 5.1) Set the feature interaction complementary learning denoising network to the test mode;
[0121] 5.2) Input the noisy images in the test set into the trained dual-branch residual sparse denoising network to obtain the denoised images.
[0122] The effects of the present invention will be further described below in conjunction with experiments.
[0123] 1. Experimental conditions:
[0124] The experiments of the present invention were carried out on a high-performance computing platform. The hardware configuration includes an AMD EPYC 8255C / 2.50GHz 12-core CPU, 128GB of memory, and an Nvidia GeForce GTX 2080Ti GPU. The software environment is based on the Ubuntu 20.04 operating system. The development and implementation use the PyTorch 1.11.0 framework and Python 3.8, and CUDA 11.3 and cuDNN 8.04 are used to accelerate the deep learning tasks. The combination of the hardware and software environments ensures that the model can efficiently process large-scale datasets and significantly improves the training speed and computing efficiency.
[0125] 2. Experimental content:
[0126] The training dataset of this experiment is divided into two parts: the synthetic noise image training dataset and the real noise image training dataset.
[0127] In terms of synthetic noise removal, the Waterloo Exploration Database (WED) is used to train the invention model. This database contains 4,744 natural images. To make full use of the computing resources, the receptive field size of the model is calculated, and the images in the training dataset are randomly cropped to a size of 128×128 for image denoising. For real noise elimination evaluation, the SIDD medium dataset is selected as the training set of the network model. This dataset contains 320 pairs of HR noisy images and their corresponding clean images. These HR images are arbitrarily cropped into image patches of 128×128 size, and rotation and flipping operations are used to augment the training samples.
[0128] In the model evaluation stage, to test the performance of the model in removing additive white Gaussian noise (AWGN), 5 public datasets are used as the test sets: Set12 and BSD68 for grayscale images, and CBSD68, Kodak24, and McMaster for color images. To test the performance of the model in real image denoising, the SIDD validation set and the DND SRGB dataset are selected as the test datasets.
[0129] The present invention uses PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) as the evaluation metrics for image denoising effect.
[0130] PSNR (Peak Signal-to-Noise Ratio) is a commonly used evaluation metric in the field of image denoising, which is used to quantify the quality of image restoration. PSNR is used to measure the difference between the image processed by noise and the original image, and the larger the value, the higher the restoration quality. The PSNR formula is defined as:
[0131]
[0132] where MAX I is the maximum possible pixel value in the image (for 8-bit images, MAX I = 255), and MSE is the Mean Squared Error, which is defined as the average of the sum of the squares of the pixel differences between the restored image and the original image. The formula is:
[0133]
[0134] where I(i,j) and K(i,j) are the pixel values of the original image and the restored image respectively, and m and n are the dimensions of the image.
[0135] SSIM (Structural Similarity Index) is a metric used to measure the similarity between two images, which pays particular attention to the structural information of the images rather than just the difference in pixel values. It is designed based on the principle of human visual perception and aims to evaluate the image quality by simulating the perception mechanism of the human eye. The value of SSIM is between 0 and 1. The closer the value is to 1, the more similar the two images are; the closer the value is to 0, the greater the difference between the two images.
[0136] Gaussian noise with σ of 15, 25, and 50 is added to the set12 grayscale image test dataset, and the existing mainstream methods and the method of the present invention are respectively used for image denoising. The PSNR and SSIM of the denoised images obtained by each method and the standard denoised image are shown in Table 1:
[0137] Table 1 Comparison of the set12 dataset metrics between the present invention and five existing mainstream methods
[0138]
[0139] Gaussian white noise with σ of 15, 25, and 50 was added to the CBSD68 color image test dataset. The existing mainstream methods and the method of the present invention were respectively used for image denoising. The PSNR of the denoised images obtained by each method and the standard denoised image is shown in Table 2:
[0140] Table 2 Comparison of the CBSD68 dataset metrics between the present invention and five existing mainstream methods
[0141]
[0142] On the SIDD and DND real image test datasets, the existing mainstream methods and the method of the present invention were respectively used for image denoising. The PSNR and SSIM of the denoised images obtained by each method and the standard denoised image are shown in Table 3:
[0143] Table 3 Comparison of the SIDD and DND dataset metrics between the present invention and five existing mainstream methods
[0144] Denoising method SIDD DND BM3D 33.40 / 0.879 37.38 / 0.929 TNRD 35.33 / 0.933 37.94 / 0.940 FFDNet 36.24 / 0.943 38.63 / 0.946 DnCNN 23.66 / 0.583 37.90 / 0.943 VDN 39.28 / 0.956 39.38 / 0.952 The present invention 39.51 / 0.957 39.66 / 0.957
[0145] In the table, the larger the peak signal-to-noise ratio PSNR and the structural similarity SSIM, the closer the denoised image is to the standard image of the dataset, and the better the denoising effect of the image.
[0146] As can be seen from Table 1, Table 2, and Table 3, the method proposed by the present invention is superior to other mainstream methods both in low-noise level scenarios and in high-noise level scenarios.
[0147] The comparison chart of the denoising effects of the present invention and several existing denoising methods in the high-noise scenario with σ = 15 is as Figure 6 shown, where (a) is the original clean image, (b) is the noisy image, (c) is the denoising effect diagram of the BM3D method, (d) is the denoising effect diagram of the DnCNN method, (e) is the denoising effect diagram of the ADNet method, and (f) is the denoising effect diagram of the method of the present invention.
[0148] From Figure 6 it can be seen that the present invention can retain more image texture features while removing noise, and obtain a higher-quality denoised image.
[0149] The complexity of the model of the present invention has been comprehensively evaluated. By comparing the number of parameters, the number of floating-point operations (flops), the average running time, and the memory usage, the lightweight and performance advantages of the model are demonstrated. To ensure the fairness of the comparison, all denoising models are implemented in the same device environment. As shown in Table 4, at a noise level of 15, images of different sizes (128×128, 256×256, and 512×512) are tested. The average running time is calculated based on 100 runs of each image denoising model. The experimental results show that compared with the DeamNet model with a similar number of parameters, the model of the present invention reduces the number of floating-point operations by 51.3% and shortens the running time by 62.2%. This indicates that the model proposed by the present invention has a significant advantage in computing efficiency and can complete the denoising task faster. In addition, the memory occupancy of the model of the present invention is reduced by 33.8% compared with DeamNet, which means that the model proposed by the present invention is more efficient in resource consumption and is suitable for deployment on devices with limited memory.
[0150] Although optimized in terms of model complexity, the model proposed by the present invention still exhibits superior peak signal-to-noise ratio (PSNR) performance in the experiment. This indicates that the model proposed by the present invention is not only lightweight but also performs excellently in terms of the quality of image denoising and detail retention, and can achieve efficient image denoising at a low computational cost, suitable for a variety of practical application scenarios.
[0151] Table 4 Flops (G), running time (ms), memory occupancy (M), and PSNR results of the present invention and existing mainstream methods
[0152]
[0153]
[0154] The specific embodiments described in the present invention are only used to illustrate the technical solutions of the present invention and do not constitute a limitation to the present invention. Although the present invention has been described in detail through specific embodiments, those skilled in the art should understand that various modifications, deformations, and equivalent replacements can be made to the present invention without departing from the basic idea and scope of the technical solution of the present invention. All such modifications, deformations, and replacements should be regarded as being within the protection scope of the claims of the present invention and should therefore be covered by the scope of the patent claims of the present invention.
Claims
1. A mobile device image denoising method based on a dual-branch residual sparse network, characterized in that It includes the following steps: S1. Data preparation and preprocessing: Images in multiple scenarios are captured by a mobile device, the images are data-annotated, enhanced, and a mobile device image dataset is constructed, and then the images in the dataset are preprocessed, and the preprocessing includes normalization, grayscale conversion, resolution adjustment, and data segmentation; S2. Construct a dual-branch residual sparse denoising network based on a series-parallel structure: The dual-branch residual sparse denoising network includes an upper-branch network, a lower-branch network, and a feature fusion module. Among them, the upper-branch network is formed by connecting five residual sparse modules in series through upsampling and downsampling, and is used to gradually extract image features and capture multi-scale information; the lower-branch network is formed by connecting five attention-guided residual sparse modules in series; the feature fusion module is used to integrate the feature information of the upper-branch network and the lower-branch network, perform final residual learning, and output the denoised image; S3. Design the loss function L of the dual-branch residual sparse denoising network as: Among them, and respectively represent the denoised image of the model and the corresponding clean image; N represents the total number of images; S4. The mobile device images obtained in S1 are divided into a training set and a test set through data segmentation. The training set images are input into the constructed dual-branch residual sparse denoising network, and the gradient of the loss function is calculated using the backpropagation method until the loss function converges, and a trained dual-branch residual sparse denoising network model is obtained; S5. The noisy images in the test set are input into the trained dual-branch residual sparse denoising network to obtain denoised images.
2. The mobile device image denoising method based on the dual-branch residual sparse network according to claim 1, wherein each of the residual sparse modules includes a standard convolution and a dilated convolution, and the input feature and the output feature are added through a residual connection to avoid the problem of gradient disappearance and improve the training stability; the attention-guided residual sparse module adds a channel attention and a pixel attention mechanism on the basis of the residual sparse module, so that the network focuses on the key regions of the image and improves the denoising effect; the residual sparse module and the attention-guided residual sparse module use a hybrid dilated convolution to increase the receptive field without increasing the amount of computation; the dilated convolution expands the coverage range by introducing gaps in the convolution kernel without increasing the parameters, and improves the receptive field and feature extraction ability of the model without increasing the amount of computation, realizing model lightweight.
3. The mobile device image denoising method based on a dual-branch residual sparse network according to claim 1, wherein The specific construction method of the upper-branch network in S2 includes: 1) A standard convolution module is composed of a standard convolution layer with a convolution kernel size of 3×3 and a stride of 1, batch normalization, and a ReLU activation function; 2) A dilated convolution module is composed of a dilated convolution layer with a convolution kernel size of 3×3, a stride of 1, and dilation rates of 2 and 3, batch normalization, and a ReLU activation function; 3) The standard convolution module, the dilated convolution module with a dilation rate of 2, the standard convolution module, and the dilated convolution module with a dilation rate of 3 are connected in sequence and connected to the input through a residual connection to form a residual sparse module; 4) Five residual sparse modules are connected in sequence. The first three residual sparse modules are connected through downsampling, and the last three residual sparse modules are connected through upsampling to form the upper-branch network; The functional representation of the upper-branch network is as follows: Among them, represents the output of each RSB module, where i ∈ 1, 2, 3, 4, 5; O shallow represents the input of the upper branch network; RSB represents the residual sparse module; Down and Up represent downsampling and upsampling respectively; represents the output of the upper branch network of the deep feature extraction module.
4. The method for denoising mobile device images based on a dual-branch residual sparse network according to claim 3, characterized in that, In the upper branch network, the residual sparse module expands the receptive field of the upper branch network by combining standard convolution and dilated convolution operations to capture multi-scale feature information. The specific content is as follows: First, extract image features through a standard convolution layer, and then use dilated convolution to increase the receptive field to obtain more extensive context information while ensuring computational efficiency; Apply residual connection between the input and output of the module to alleviate the vanishing gradient problem in deep networks, promote information flow, and accelerate the training and convergence of the network. Its functional representation is as follows: O RSB = RSB(f rsb ) + f rsb = DBR3(CBR(DBR2(CBR(f rsb )))) + f rsb Among them, f rsb represents the input of the RSB module; CBR represents the standard convolution module; DBR2 represents the dilated convolution module with a dilation rate of 2; DBR3 represents the dilated convolution module with a dilation rate of 3; the number of channels of both the standard convolution and the dilated convolution is 64, and the convolution kernel is 3×3.
5. The lower branch network of the mobile device image denoising method based on the double-branch residual sparse network according to claim 1, characterized in that The construction method of the lower branch network described in S2 specifically includes: 1) Connect a standard convolution module and a dilated convolution module with a dilation rate of 2 and perform a residual connection to form an internal residual sparse module; 2) Connect the internal residual sparse module, channel attention mechanism, pixel attention mechanism, standard convolution module, and dilated convolution module with a dilation rate of 3 in sequence, and perform a residual connection with the input to form an attention-guided residual sparse module; 3) Connect five attention-guided residual sparse modules in sequence to form the lower branch network; The functional representation of the lower branch network is as follows: Among them, represents the output of each ARSB module, where i ∈ 1, 2, 3, 4, 5; O shallow represents the input of the upper branch network; ARSB represents the attention-guided residual sparse module; represents the output of the lower branch network of the deep feature extraction module.
6. The method for denoising mobile device images based on a double-branch residual sparse network according to claim 5, wherein The attention-guided residual sparse module described in the lower branch network specifically includes the following content: First, connect a standard convolution module and a dilated convolution module with a dilation rate of 2 in sequence, and form an internal residual sparse module through a residual connection; use the standard convolution to extract local features, while the dilated convolution captures more extensive context information by increasing the receptive field; Next, connect the internal residual sparse module, channel attention mechanism, pixel attention mechanism, standard convolution module, and dilated convolution module with a dilation rate of 3 in sequence to form an attention-guided residual sparse module; in the attention-guided residual sparse module, the channel and pixel attention mechanisms adaptively assign weights to different features to enhance the network's attention to important features; Finally, add the output of the module to the initial input through a residual connection to ensure the efficient flow of information and strengthen the gradient propagation during the training process; the attention-guided residual sparse module combines multi-scale feature extraction, attention mechanism, and residual connection to improve the performance and stability of the lower branch network in complex tasks. Its functional representation is as follows: O ARSB = ARSB(f arsb ) + f arsb = DBR3(CBR(CAB(PAB(DBR2(CBR(f arsb )))))+ f arsb ) + f arsb Wherein, f arsb represents the input of the ARSB module; CAB represents the channel attention module; PAB represents the pixel attention module.
7. The method for denoising mobile device images based on a dual-branch residual sparse network according to claim 1, wherein, The feature fusion module described in S2 specifically includes the following content: Feature fusion input: Add the outputs of the upper branch network and the lower branch network as the input of the feature fusion module for feature fusion; Module composition: The feature fusion module consists of a standard convolution module, an attention-guided residual sparse module, two standard convolutions, and a Sigmoid activation function; Feature fusion process: 1) Pass the input of the feature fusion module through the standard convolution module, attention-guided residual sparse module, and the first standard convolution in sequence; 2) Concatenate the above output with the normalized noise image and pass it through the Sigmoid activation function and the second standard convolution in sequence; 3) Multiply the output of the first standard convolution by the output of the second standard convolution, and finally perform a residual with the original noise image to obtain the final denoised image; The functional representation of the feature fusion process is as follows: Among them, f rb represents the input of the RB module; cat represents the concatenation operation; C represents the standard convolution; I n represents the input noisy image; Sig represents the Sigmoid function.
8. The mobile device image denoising method based on a dual-branch residual sparse network according to claim 1, characterized in that In S4, the gradient of the loss function is calculated using the backpropagation method to train the model. The specific implementation steps are as follows: 1) Set the dual-branch residual sparse denoising network to the training mode; 2) Input the noisy images in the training set into the dual-branch residual sparse denoising network, perform forward propagation, and calculate and output the fused denoised images; 3) Calculate the loss between the model output and the original clean image according to the loss function; 4) Calculate the gradient of the loss function according to the backpropagation method; 5) Use the Adam optimizer to update the model parameters according to the gradient change of the loss function. The learning rate of the optimizer is set to 0.001, and the hyperparameters β1 = 0.9 and β2 = 0.99; 6) Repeat the above process until the loss function reaches stability, and obtain the trained dual-branch residual sparse denoising network model.
Citation Information
Patent Citations
SAR image denoising method based on multi-scale residual attention network
CN112233026A
Sparse angle CT artifact removal method
CN114596378A
Junk image denoising method based on multi-dimensional image information fusion
CN116543168A
Intensive LSTM residual network denoising method based on dynamic attention
CN116563144A
Self-supervised image denoising method based on three-stage feature extraction
CN118097159A
Cited By
Deep learning denoising model and method based on feature fusion and attention mechanism
CN120894557A
Lightweight denoising method based on convolutional neural network
CN121053033A
Geomagnetic data denoising method and system using dense residual shuffling attention network
CN121210849A