Image denoising method, device and equipment based on denoising diffusion probability model
The denoising diffusion probability model constructed through the U-shaped convolutional neural network and the dual attention mechanism directly uses image noise information to calculate the clean image, solving the problem of many iterative sampling steps in the existing technology, and achieving efficient and accurate image denoising effect, which is suitable for medical imaging and high-definition photography and other applications.
Patent Information
- Application Number
- CN202510545185.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
The existing diffusion model requires a large number of iterative sampling steps during image denoising, which leads to large consumption of computing resources and long processing time, making it difficult to efficiently generate high-quality images.
The denoising diffusion probability model built on U-type convolutional neural network is adopted, combined with the dual attention mechanism, and the clean image is calculated directly using image noise information to reduce the number of iterations and computing resource consumption.
Efficient image denoising is achieved in fewer iteration steps, significantly reducing the calculation amount and processing time, while improving the accuracy and efficiency of image denoising, and is suitable for fields such as medical imaging and high-definition photography.
Smart Images

Figure CN120471793A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an image denoising method, apparatus and device based on a denoising diffusion probability model. Background Art
[0002] Image denoising is a key research area in computer vision. Its core goal is to minimize or eliminate the effects of noise on an image, making the processed image as close to the original as possible. Essentially, it restores and reconstructs data, removing contamination. This technology has important applications in numerous fields, such as medical imaging diagnosis and security monitoring. Clear images provide a more reliable basis for subsequent analysis and decision-making.
[0003] Currently, image denoising methods primarily fall into three categories: filter-based, model-based, and learning-based. Filter-based methods utilize manually designed low-pass filters to remove noise. For example, the median filter, a commonly used nonlinear smoothing filter, effectively eliminates isolated noise points by replacing the value of a point in an image with the median of the values of all points in its neighborhood. Model-based methods model the distribution of natural images or noise, transforming the denoising task into a maximum a posteriori (MAP) optimization problem, relying on image priors to obtain a clear image. However, the models of such methods are often non-convex and involve multiple manually selected parameters, which limits performance to a certain extent. Learning-based methods, particularly those based on deep networks, have gradually become mainstream in recent years. These methods utilize deep learning models such as convolutional neural networks (CNNs) to extract image features and learn denoising mapping relationships through training.
[0004] While existing image denoising techniques have achieved some success, diffusion models still struggle to generate high-quality samples from noise. Existing diffusion models typically require hundreds or even thousands of iterative sampling steps, each of which involves complex computations, including interactions with model parameters, gradual noise reduction, and fine-tuning of data distribution. This not only consumes significant computational resources during backpropagation but also significantly prolongs the time required for training and generating new samples, necessitating innovative technologies to optimize and improve these efforts. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides an image denoising method, device and equipment based on a denoising diffusion probability model, which can output noise information by adopting a denoising diffusion probability model with a U-shaped convolutional neural network architecture, and efficiently use the noise information to directly calculate a clean image.
[0006] In order to achieve the above objectives, the technical solutions provided in the embodiments of the present application are as follows:
[0007] In the first aspect, the present application provides an image denoising method based on a denoising diffusion probability model, comprising: inputting a noisy image into a denoising diffusion probability model to obtain image noise output by the denoising diffusion probability model; wherein the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; based on the noisy image and the image noise, a clean image is calculated.
[0008] As an optional implementation provided in an embodiment of the present application, the U-shaped convolutional neural network includes dimensional layers, wherein each dimensional layer includes a residual block that introduces a dual attention mechanism.
[0009] As an optional implementation provided in an embodiment of the present application, the following steps are cyclically performed to train a denoising diffusion probability model: a clean sample image is randomly selected from a training data set, and a random time step is determined; within the random time step, the actual noise corresponding to the random time step is added to the clean sample image at each moment to convert the clean sample image into a pure noise image; the pure noise image is input into the denoising diffusion probability model to obtain predicted noise; the objective function between the predicted noise and the actual noise is calculated; and the model parameters of the denoising diffusion probability model are optimized based on the objective function until the objective function converges to obtain a trained denoising diffusion probability model.
[0010] As an optional implementation provided in an embodiment of the present application, the objective function is the square of the second norm of the difference between the predicted noise and the actual noise; optimizing the model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain the trained denoising diffusion probability model, including: calculating the square of the second norm of the difference between the predicted noise and the actual noise; calculating the gradient of the model parameters based on the square of the second norm of the difference; performing a gradient descent step to update the model parameters until the square of the second norm of the difference converges to obtain the trained denoising diffusion probability model.
[0011] As an optional implementation provided in an embodiment of the present application, a clean image is calculated based on a noisy image and image noise, including: determining a time step of the noisy image; if the time step of the noisy image is greater than time step 1, sampling noise from a standard normal distribution; based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, calculating an image of the previous time step of the noisy image; until the time step of the noisy image is equal to time step 1, a clean image is obtained.
[0012] As an optional implementation provided in an embodiment of the present application, based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, calculating the image of the noisy image at the previous time step includes obtaining the image of the noisy image at the previous time step according to the following calculation formula:
[0013]
[0014] Among them, x t-1 is the image of the previous time step of the noisy image, x t is a noisy image, α t , and σ t is a parameter related to the time step t, ∈ θ (x t ,t) is the image noise, σ t is the standard deviation parameter associated with time step t, and z is the noise sampled from a standard normal distribution.
[0015] In a second aspect, the present application provides an image denoising device based on a denoising diffusion probability model, the device comprising:
[0016] A noise prediction module is used to input the noisy image into a denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; wherein the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network;
[0017] The denoising module is used to calculate a clean image based on the noisy image and image noise.
[0018] In a third aspect, the present application provides an electronic device comprising: a processor, a memory, and a computer program stored on the memory and runnable on the processor, wherein when the computer program is executed by the processor, the image denoising method based on the denoising diffusion probability model as described in the first aspect or any one of its optional embodiments is implemented.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, comprising: a computer program stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the image denoising method based on the denoising diffusion probability model as described in the first aspect or any one of its optional embodiments.
[0020] In a fifth aspect, the present application provides a computer program product, comprising: the computer program product includes a computer program, and when the computer program is run on a computer, the computer implements the image denoising method based on the denoising diffusion probability model as described in the first aspect or any one of its optional embodiments.
[0021] The technical solution provided by the embodiments of the present application has the following advantages compared with the prior art:
[0022] The embodiments of the present application provide an image denoising method, apparatus and device based on a denoising diffusion probability model, wherein the method adopts a denoising diffusion probability model constructed based on a U-shaped convolutional neural network, so that the model can capture the characteristic information of the image at different scales, has high efficiency in the image denoising task, and can achieve a better denoising effect within fewer iterative steps, thereby reducing the consumption of computing resources and processing time; the image noise output by the denoising diffusion probability model contains the key information of the noise in the noisy image; the image noise and noisy image output directly based on the denoising diffusion probability model fully utilize this information, and directly restore the clean image through a reasonable calculation method, reducing unnecessary repeated calculations, avoiding a large number of iterative processes in the traditional diffusion model, and significantly reducing the amount of calculation and processing time. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] Figure 1 A flowchart of an image denoising method based on a denoising diffusion probability model provided in an embodiment of the present application;
[0026] Figure 2 A schematic diagram of the structure of the denoising diffusion probability model provided in an embodiment of the present application;
[0027] Figure 3 A schematic diagram of a pseudo-code algorithm for the denoising diffusion probability model training process provided in an embodiment of the present application;
[0028] Figure 4 A schematic diagram of a pseudo-code algorithm for the sampling process of a denoising diffusion probability model provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of the structure of an image denoising device based on a denoising diffusion probability model provided in an embodiment of the present application;
[0030] Figure 6 This is a schematic structural diagram of an electronic device described in an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the technical terms required to be used in the embodiments or the description of the prior art will be briefly introduced below.
[0032] The U-Net Convolutional Neural Network (UNET) is a convolutional neural network (CNN) architecture used for image segmentation. It features a symmetrical encoder-decoder structure. The encoder gradually reduces the spatial dimensions of the image while increasing the number of channels to capture deep features. The decoder performs the opposite operation, gradually restoring the spatial dimensions and reducing the number of channels, ultimately outputting a result of the same size as the input image. During this process, skip connections are used between the encoder and decoder to help the decoder better recover image details.
[0033] The Dual Attention Mechanism (DAM) is a technique used in deep learning to enhance a model's ability to focus on different information in the data. It combines spatial and channel-wise attention mechanisms, allowing the model to adaptively focus on and weight data across both spatial and channel dimensions. The spatial attention weight map and channel-wise attention weight vector are first calculated separately. These are then combined in some way, such as element-wise multiplication or addition, and then applied to the original feature map to generate a feature map enhanced with both spatial and channel-wise attention. This fully utilizes information in both spatial and channel-wise dimensions, enabling the model to more accurately capture key information in the data, improving both performance and expressiveness.
[0034] A time step is a division of time into a series of discrete, equally or unequally spaced steps in a time series or dynamic system. Each time step represents a specific time segment during which the state of the system or related variables is considered to be relatively stable or change according to a certain regularity.
[0035] In order to more clearly understand the above-mentioned objectives, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0036] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application can also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present application, not all of the embodiments.
[0037] In order to solve some or all of the technical problems existing in the related art, the embodiments of the present application provide an image denoising method, device and equipment based on a denoising diffusion probability model, wherein the method adopts a denoising diffusion probability model constructed based on a U-shaped convolutional neural network, so that the model can capture the characteristic information of the image at different scales, has high efficiency in the image denoising task, and can achieve better denoising effect within fewer iterative steps, thereby reducing the consumption of computing resources and processing time; the image noise output by the denoising diffusion probability model contains the key information of the noise in the noisy image; directly based on the image noise and noisy image output by the denoising diffusion probability model, make full use of this information, and directly restore the clean image through a reasonable calculation method, reducing unnecessary repeated calculations, avoiding a large number of iterative processes in the traditional diffusion model, and significantly reducing the amount of calculation and processing time.
[0038] An image denoising method based on a denoising diffusion probability model provided in an embodiment of the present application can be implemented by an image denoising device or electronic device based on a denoising diffusion probability model, and the electronic device includes but is not limited to a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The operating system of the electronic device may include Android, a mobile operating system (iOS) developed by Apple, an operating system (Windows) developed by Microsoft Corporation of the United States, etc., and the embodiment of the present application is not limited to this. The electronic device can be run alone to implement the present application, or it can be connected to a network and implement the present application through interactive operations with other computer devices in the network. Among them, the network in which the electronic device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0039] It should be noted that the protection scope of the image denoising method based on the denoising diffusion probability model described in the embodiment of the present application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, subtracting, or replacing steps in the existing technology based on the principles of the present application are included in the protection scope of the present application.
[0040] like Figure 1 As shown, Figure 1 This is a flow chart of an image denoising method based on a denoising diffusion probability model according to an embodiment of the present application. This method can be performed by an image denoising device based on a denoising diffusion probability model, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. The method mainly includes the following steps S101-S102:
[0041] S101 , inputting a noisy image into a denoising diffusion probability model to obtain image noise output by the denoising diffusion probability model.
[0042] S102: Calculate and obtain a clean image based on the noisy image and the image noise.
[0043] In the embodiments of this application, the Denoising Diffusion Probabilistic Model (DDPM) is a trained model for predicting noise. DDPM is built on the UNET neural network. Because the U-shaped convolutional neural network has powerful feature extraction capabilities, it can effectively extract semantic and detail information at different levels of the image under the encoder-decoder structure, thereby accurately separating the noise components from the noisy image.
[0044] In some embodiments, a dual attention mechanism is introduced into the residual block between each dimension layer (resolution level) of the UNET neural network of the present application. Figure 2 As shown, Figure 2 Schematic diagram of the structure of the denoising diffusion probability model provided in the embodiment of the present application. The dual attention mechanism includes a spatial attention module and a channel attention module, which are used to weight the spatial position importance and channel semantic importance of the feature map respectively. The denoising diffusion probability model enhances the recognition ability of noise features through the dual attention mechanism and predicts the image noise corresponding to the noisy image. By introducing the dual attention mechanism in the U-shaped network, the noise prediction accuracy and feature processing efficiency are significantly improved, the number of iterations required for the traditional diffusion model is effectively reduced, and the consumption of computing resources is reduced.
[0045] Each residual block includes: a group normalization component, an activation function component, a convolution component, a regularization component, and a skip connection component. The group normalization component is a variant of batch normalization used to control internal covariate shift. The activation function component can be a rectified linear unit (ReLU) activation function component, which allows the model to capture complex patterns and nonlinear relationships in the input data. The convolution component can be a 3×3 same-size convolution component. This convolution operation maintains the spatial size of the output feature map to the same as the input through appropriate padding. The regularization component can be a random dropout component, which prevents model overfitting by randomly dropping (i.e., setting to zero) some activation units in the network during training. The skip connection component directly passes the output of a previous layer to the next layer, making it easier for the network to learn the residual mapping between input and output during training. This helps to solve the vanishing gradient problem in deep networks and allows the model to retain information about primary features in deep layers.
[0046] Optionally, each dimensional layer of the UNET neural network includes two residual blocks. The skip connections in the residual blocks allow gradients to pass directly, making the network more convergent during training and alleviating the vanishing gradient problem. The two residual blocks can perform more complex nonlinear transformations on the input features, thereby learning richer feature representations and improving the segmentation performance of the UNET neural network.
[0047] Introducing a dual-attention mechanism within the residual block allows the network to adaptively select spatial regions and channel features that are more important for the current task. In UNET, different image regions and feature channels may contribute differently to the segmentation task. The dual-attention mechanism allows the network to focus on key feature information and suppress irrelevant information, thereby improving segmentation accuracy. Incorporating a dual-attention mechanism into the two residual blocks between dimensional layers promotes the interaction and fusion of features across different layers. The spatial attention mechanism helps the network capture dependencies between features at different spatial locations, while the channel attention mechanism facilitates information exchange between features across different channels, enabling the network to better utilize comprehensive feature information.
[0048] The following describes the training process of the denoising diffusion probability model. The following steps are performed in a loop:
[0049] Step 1. Randomly select clean sample images from the training dataset.
[0050] Step 2. Pick a random time step.
[0051] Optionally, a random time step is randomly selected from the noise plan. A noise plan is a pre-set plan for noise that specifies the type, intensity, distribution, and temporal behavior of the noise.
[0052] It can be understood as randomly selecting a time step from the time range covered by the pre-established noise plan.
[0053] Step 3. Within the random time step, the actual noise corresponding to the random time step is added to the clean sample image at each moment to convert the clean sample image into a pure noise image.
[0054] The actual noise corresponding to the time step is added to the clean sample image, and the forward diffusion process is simulated by the diffusion kernel to obtain the noisy sample image;
[0055] In some embodiments, at random time steps, a clean sample image is gradually converted into a pure noise image by adding Gaussian noise to the sample image at the previous moment, as shown in formula (1):
[0056]
[0057] Formula (1) is used to define the sample x at a given previous moment t-1 Under the condition of t The probability distribution of q(x t |x t-1 ) means that when x is known t-1 Under the condition of t The probability distribution of . Here x t and x t-1 Represents images (or data samples) at different time steps in the diffusion process, t represents the time step, t ranges from 0 to T, 0 corresponds to the original clean sample, and T corresponds to the completely noisy sample.
[0058] Represents a t is a normal distribution (Gaussian distribution) of random variables. is the mean, that is, the mean of the normal distribution is determined by the sample x at the previous moment t-1 Multiply by the coefficient Here β t is a parameter related to the time step t, which controls the variance of the noise added to the samples at each time step. is the covariance matrix, where is the identity matrix, which indicates that the normal distribution is independent in each dimension and has a variance of β t That is, in the t-1 Generate x t In the process, the noise added is consistent with the mean of 0 and the variance of β t Gaussian noise.
[0059] This formula describes a core step in the DDPM diffusion process: how, at each time step, clean samples are gradually converted into noisy samples by adding Gaussian noise to the previous sample. As t increases, the noise content in the sample gradually increases, eventually approaching pure noise. The process of adding noise multiple times can be viewed as the superposition of multiple normal distributions. Assuming that each added noise follows a normal distribution, after linear transformation and superposition, the final pure noise image also follows a new normal distribution.
[0060] Step 4. Input the pure noise image into the denoising diffusion probability model to obtain the predicted noise.
[0061] It should be noted that according to the current sample x t Predict the sample x at the previous moment t-1 , x t-1 It is not a completely determined value and contains random components. Because the diffusion process is essentially a process of gradually adding noise to the clean sample, from x t-1 to x tIt is a random transformation achieved by adding noise, so the reverse direction from x t-1 Push x t There is also randomness. To solve the above problem, this application uses the reparameterization technique. The core idea is to put x t-1 Specifically, in the probability distribution of the diffusion model, the original x t is based on x t-1 The random variable obtained by adding noise, reparameterization is to re-express this random process so that the noise part can be independently controlled. After the recursive formula is expanded by reparameterization, the input-output relationship of the network becomes clearer, and the properties of the normal distribution can be used for mathematical deduction and calculation, so that the network can be trained under the premise of meeting the probability distribution characteristics of the diffusion model, realizing the t Accurately predict x t-1 goal.
[0062] Step 5. Calculate the objective function between the predicted noise and the actual noise.
[0063] Optionally, the objective function is the square of the dichotomy of the difference between the predicted noise and the actual noise, or the mean square error between the predicted noise and the actual noise.
[0064] Step 6. Optimize the model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain the trained denoising diffusion probability model.
[0065] Optionally, the square of the second norm of the difference between the predicted noise and the actual noise difference is calculated, and then the gradient of the model parameters is calculated based on the square of the second norm of the difference; then a gradient descent step is performed to update the model parameters of the denoised diffusion probability model until the square of the second norm of the difference converges, thereby obtaining a trained denoised diffusion probability model.
[0066] For example, Figure 3 The figure shows the pseudocode algorithm for the denoising diffusion probability model (DDPM) training process. This algorithm aims to train a UNET with a dual-attention mechanism, enabling it to learn to predict noise from noisy images, which can then be used for subsequent tasks such as image denoising. The training process repeatedly samples, calculates gradients, and updates model parameters until convergence.
[0067] Repeat means starting a loop, which is used to continuously iterate the training process. x0~q(x0) means sampling the initial clean sample image x0 from the distribution q(x0), that is, the training dataset. x0 is a real clean image, such as a natural image, a medical image, etc. q(x0) represents the distribution of real data. t~Uniform({1,...,T}) means uniformly sampling a random time step t from the integer set {1,...,T}. In DDPM, the diffusion process has T time steps, and t is used to indicate the current time stage of the diffusion process. Different t corresponds to images with different noise levels. ∈~N(0,I): Sample noise ∈ from a standard normal distribution (mean 0, covariance matrix is the identity matrix I). This noise will be added to the image to simulate the noise addition step in the diffusion process.
[0068] This is the core step of training. First, represents the image after adding noise at random time step t, where is a parameter related to the random time step t, given by Calculated, β s is a predefined noise variance parameter for each time step. ∈ θ It is a neural network (such as UNET) controlled by parameter θ, whose input is a pure noise image and a random time step t, and the output is a prediction of the noise The prediction noise is calculated here The gradient of the mean-square error (MSE) with the true sampling noise ∈ A gradient descent step is performed to update the model parameters θ so that the model can accurately predict the noise from the noisy image.
[0069] Untilconverged means that the loop continues until the model converges. Convergence usually means that the loss function (here, the mean squared error) no longer decreases significantly and the model parameter updates tend to be stable.
[0070] It can be understood that the training process of the denoising diffusion probability model is as follows:
[0071]
[0072] In some embodiments, when executing step S102, the time step of the noisy image is first determined, and a comparison is made to see whether the time step of the noisy image is greater than time step 1. If so, noise is sampled from a standard normal distribution. Then, based on the noisy image, the time step of the noisy image, the image noise, and the sampling noise, the image of the previous time step of the noisy image is calculated until the time step equals time step 1, thereby obtaining a clean image. This achieves noise removal from the noisy image.
[0073] Optionally, the image of the previous time step of the noisy image is calculated according to the following formula (2):
[0074]
[0075] Among them, x t-1 is the image of the previous time step of the noisy image, x t is a noisy image, α t , and σ t is a parameter related to the time step t of the noisy image, ∈ θ (x t ,t) represents the image noise output by the denoising diffusion probability model, and z is the noise sampled from the standard normal distribution. t is a standard deviation parameter relative to the time step t, which controls the scale of the added noise z. β s is a predefined noise variance parameter for each time step.
[0076] like Figure 4 The pseudo code algorithm of the denoising diffusion probability model (DDPM) sampling process is shown in the figure, which is used to gradually generate a clean image from a noisy image. The algorithm describes the process of generating a clean image from a pure noise image x T The process of gradually removing noise through multiple iterations and finally obtaining a clean image x0 is based on the trained denoising diffusion probability model ∈ θ To achieve reverse diffusion. T ~N(0,I) represents the initialization of pure noise image x T , which is sampled from a standard normal distribution with mean 0 and covariance matrix I. This represents the final state of the diffusion process, which is a noisy state.
[0077] “for t=T,…,1do” is a loop structure that gradually decreases from the time step T to 1 and executes the reverse diffusion process.
[0078] “z~N(0,I)ift>1,elsez=0” means that if the current time step t is greater than 1, noise z is sampled from the standard normal distribution; when t=1, z is set to 0. This is because no additional noise needs to be added in the last step of back diffusion.
[0079] It is the core formula of back diffusion. The function of this formula is to calculate the value of the pure noise image x t , predicted noise ∈ θ (x t ,t) and the sampled noise z, calculate the image x of the previous time step t-1 , to gradually remove the noise.
[0080] returnx0 means that after the loop ends, the final clean image x0 is returned.
[0081] It can be understood that the sampling process of the denoising diffusion model is as follows:
[0082]
[0083] In summary, an embodiment of the present application provides an image denoising method based on a denoising diffusion probability model. The method first inputs a noisy image into a denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; wherein the denoising diffusion probability model is a trained model for predicting noise; the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network, wherein a dual attention mechanism is introduced in the residual block between each dimensional layer of the U-shaped convolutional neural network; then, based on the noisy image, the time step of the noisy image, and the image noise, a clean image is calculated.
[0084] This technical solution focuses on image denoising. Based on a denoising diffusion probability model and combined with a specific network construction method, it has significant technical effects in terms of image denoising accuracy, computational efficiency, and feature processing capabilities. The details are as follows:
[0085] (1) Improve the accuracy of noise prediction: The noisy image is input into a denoising diffusion probability model built and trained based on a U-shaped convolutional neural network, which can output image noise. Since the U-shaped convolutional neural network itself has a strong feature extraction capability, it can effectively extract semantic and detail information at different levels of the image under the encoder-decoder structure. At the same time, a dual attention mechanism is introduced in the residual block between each dimensional layer. The spatial attention mechanism enables the model to focus on the key areas where noise is located in the image, and the channel attention mechanism can enhance the sensitivity to the feature channels related to noise. The two work together to make the model's prediction of noise more accurate. Accurate noise prediction is a key prerequisite for achieving high-quality image denoising, which helps to more accurately separate the noise components from the noisy image, thereby obtaining a clean image that is closer to the original image.
[0086] (2) Optimizing the computational process and efficiency: Compared to existing diffusion models that require hundreds or even thousands of iterative sampling steps, this solution utilizes a trained denoising diffusion probability model to calculate a clean image based on the noisy image and the predicted image noise, reducing unnecessary iterations and complex calculations. By optimizing the network structure and mechanism, while ensuring the denoising effect, it reduces the computational resource consumption in backpropagation, shortens the processing time, and improves the overall efficiency of image denoising, making it more adaptable to the real-time and resource utilization requirements of practical applications.
[0087] (3) Enhanced feature processing and image restoration capabilities: The jump connection design of the U-shaped convolutional neural network enables the low-level detail features extracted by the encoder to be integrated with the features in the decoder's resolution restoration process. Combined with the residual block, it alleviates the gradient vanishing problem and ensures the network depth and feature extraction capabilities. The dual attention mechanism further enhances the processing capabilities of different spatial positions and channel features, enabling the model to better retain important information such as the image structure and texture when calculating the clean image, avoiding excessive blurring or loss of details during the denoising process, thereby generating clean images with higher quality and better visual effects. It has important application value in fields with high image quality requirements such as medical imaging and high-definition photography.
[0088] like Figure 5 As shown, Figure 5 A schematic diagram of the structure of an image denoising device based on a denoising diffusion probability model provided in an embodiment of the present application, the device comprising:
[0089] Noise prediction module 501, used to input the noisy image into the denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network;
[0090] The denoising module 502 is configured to calculate a clean image based on the noisy image and the image noise.
[0091] As an optional implementation provided in an embodiment of the present application, the U-shaped convolutional neural network includes dimensional layers, wherein each dimensional layer includes a residual block that introduces a dual attention mechanism.
[0092] As an optional implementation provided in an embodiment of the present application, the image denoising device based on the denoising diffusion probability model also includes a training module, which is used to cyclically execute the following steps to train and obtain the denoising diffusion probability model: randomly select a clean sample image from a training data set, and randomly select a random time step; within the random time step, successively add the actual noise corresponding to the random time step to the clean sample image at each moment to convert the clean sample image into a pure noise image; input the pure noise image into the denoising diffusion probability model to obtain predicted noise; calculate the objective function between the predicted noise and the actual noise; and optimize the model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain the trained denoising diffusion probability model.
[0093] As an optional implementation provided in an embodiment of the present application, the objective function is the square of the second norm of the difference between the predicted noise and the actual noise; the training module optimizes the model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain the trained denoising diffusion probability model, and is specifically used to: calculate the square of the second norm of the difference between the predicted noise and the actual noise; calculate the gradient of the model parameters based on the square of the second norm of the difference; and perform a gradient descent step to update the model parameters until the square of the second norm of the difference converges to obtain the trained denoising diffusion probability model.
[0094] As an optional implementation provided in an embodiment of the present application, the denoising module 502 is specifically used to: determine the time step of the noisy image, and if the time step of the noisy image is greater than time step 1, sample noise from a standard normal distribution; based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, calculate the image of the noisy image at the previous time step; until the time step of the noisy image is equal to time step 1, a clean image is obtained.
[0095] As an optional implementation provided in an embodiment of the present application, the denoising module 502 calculates the image of the noisy image at the previous time step based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution. Specifically, the denoising module 502 is configured to calculate the image of the noisy image at the previous time step according to the following calculation formula:
[0096]
[0097] Among them, x t-1 is the image of the previous time step of the noisy image, x t is a noisy image, α t , and σ t is a parameter related to the time step t, ∈ θ (x t ,t) is the image noise, σ t is the standard deviation parameter associated with time step t, and z is the noise sampled from a standard normal distribution.
[0098] Regarding the specific definition of the image denoising device based on the denoising diffusion probability model, please refer to the definition of the image denoising method based on the denoising diffusion probability model above, and will not be repeated here. The various modules in the above-mentioned image denoising device based on the denoising diffusion probability model can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of the above-mentioned modules.
[0099] In one embodiment, the present application provides an electronic device, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method for detecting a jam is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the electronic device housing, or an external keyboard, touchpad or mouse.
[0100] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0101] In one embodiment, the image denoising device based on the denoising diffusion probability model provided by the present application can be implemented in the form of a computer program. Figure 6 The memory of the electronic device can store various program modules constituting the image denoising device based on the denoising diffusion probability model, such as: Figure 5 The noise prediction module 501 and the denoising module 502 are shown. The computer program composed of various program modules enables the processor to execute the steps of the image denoising method based on the denoising diffusion probability model of various embodiments of the present application described in this specification.
[0102] For example, Figure 6 The electronic device shown can be Figure 5 The noise prediction module 501 based on the denoising diffusion probability model shown is executed by inputting the noisy image into the denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; wherein the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; the electronic device can perform calculations based on the noisy image and image noise through the denoising module 502 to obtain a clean image.
[0103] In one embodiment, the present application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0104] The noisy image is input into the denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; based on the noisy image and image noise, a clean image is calculated.
[0105] In one embodiment, when the computer program is executed, the computer program further implements the following steps: looping through the following steps to train and obtain a denoising diffusion probability model: randomly selecting a clean sample image from a training data set and randomly selecting a random time step; within the random time step, adding actual noise corresponding to the random time step to the clean sample image at each moment to convert the clean sample image into a pure noise image; inputting the pure noise image into the denoising diffusion probability model to obtain predicted noise; calculating an objective function between the predicted noise and the actual noise; and optimizing model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain a trained denoising diffusion probability model.
[0106] In one embodiment, when the computer program is executed, the computer program further implements the following steps: the objective function is the square of the second norm of the difference between the predicted noise and the actual noise; the model parameters of the denoised diffusion probability model are optimized based on the objective function until the objective function converges to obtain the trained denoised diffusion probability model, including: calculating the square of the second norm of the difference between the predicted noise and the actual noise; calculating the gradient of the model parameters based on the square of the second norm of the difference; and performing a gradient descent step to update the model parameters until the square of the second norm of the difference converges to obtain the trained denoised diffusion probability model.
[0107] In one embodiment, when the computer program is executed, the computer program further implements the following steps: calculating a clean image based on the noisy image and the image noise, including: determining a time step of the noisy image, and if the time step of the noisy image is greater than time step 1, sampling noise from a standard normal distribution; calculating an image of the previous time step of the noisy image based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution; and until the time step of the noisy image is equal to time step 1, obtaining a clean image.
[0108] In one embodiment, when the computer program is executed, the computer program further implements the following steps: calculating, based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, an image at a previous time step of the noisy image includes calculating the image at a previous time step of the noisy image according to the following calculation formula:
[0109]
[0110] Among them, x t-1 is the image of the previous time step of the noisy image, x t is a noisy image, α t , and σ t is a parameter related to the time step t, ∈ θ (x t ,t) is the image noise, σ t is the standard deviation parameter associated with time step t, and z is the noise sampled from a standard normal distribution.
[0111] When the processor in the electronic device provided by the present application executes a computer program, a denoising diffusion probability model constructed based on a U-shaped convolutional neural network is adopted, so that the model can capture the characteristic information of the image at different scales, has high efficiency in the image denoising task, and can achieve a better denoising effect within fewer iterative steps, thereby reducing the consumption of computing resources and processing time; the image noise output by the denoising diffusion probability model contains the key information of the noise in the noisy image; directly based on the image noise and noisy image output by the denoising diffusion probability model, this information is fully utilized, and a clean image is directly restored through a reasonable calculation method, which reduces unnecessary repeated calculations, avoids a large number of iterative processes in traditional diffusion models, and significantly reduces the amount of calculation and processing time.
[0112] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by the computer program, performs the following steps:
[0113] The noisy image is input into the denoising diffusion probability model to obtain the image noise output by the denoising diffusion probability model; the denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; based on the noisy image and image noise, a clean image is calculated.
[0114] In one embodiment, when the computer program is executed, the computer program further implements the following steps: looping through the following steps to train and obtain a denoising diffusion probability model: randomly selecting a clean sample image from a training data set and randomly selecting a random time step; within the random time step, adding actual noise corresponding to the random time step to the clean sample image at each moment to convert the clean sample image into a pure noise image; inputting the pure noise image into the denoising diffusion probability model to obtain predicted noise; calculating an objective function between the predicted noise and the actual noise; and optimizing model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain a trained denoising diffusion probability model.
[0115] In one embodiment, when the computer program is executed, the computer program further implements the following steps: the objective function is the square of the second norm of the difference between the predicted noise and the actual noise; the model parameters of the denoised diffusion probability model are optimized based on the objective function until the objective function converges to obtain the trained denoised diffusion probability model, including: calculating the square of the second norm of the difference between the predicted noise and the actual noise; calculating the gradient of the model parameters based on the square of the second norm of the difference; and performing a gradient descent step to update the model parameters until the square of the second norm of the difference converges to obtain the trained denoised diffusion probability model.
[0116] In one embodiment, when the computer program is executed, the computer program further implements the following steps: calculating a clean image based on the noisy image and the image noise, including: determining a time step of the noisy image; if the time step of the noisy image is greater than time step 1, sampling noise from a standard normal distribution; calculating an image of the previous time step of the noisy image based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution; until the time step of the noisy image is equal to time step 1, a clean image is obtained.
[0117] In one embodiment, when the computer program is executed, the computer program further implements the following steps: calculating, based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, an image at a previous time step of the noisy image includes calculating the image at a previous time step of the noisy image according to the following calculation formula:
[0118]
[0119] Among them, x t-1 is the image of the previous time step of the noisy image, x t is a noisy image, α t , and σ t is a parameter related to the time step t, ∈ θ (x t ,t) is the image noise, σ t is the standard deviation parameter associated with time step t, and z is the noise sampled from a standard normal distribution.
[0120] When the computer program in the computer-readable storage medium provided by the present application executes the computer program, a denoising diffusion probability model constructed based on a U-shaped convolutional neural network is adopted, so that the model can capture the characteristic information of the image at different scales, has high efficiency in the image denoising task, and can achieve a better denoising effect within fewer iterative steps, thereby reducing the consumption of computing resources and processing time; the image noise output by the denoising diffusion probability model contains the key information of the noise in the noisy image; directly based on the image noise and noisy image output by the denoising diffusion probability model, this information is fully utilized, and a clean image is directly restored through a reasonable calculation method, which reduces unnecessary repeated calculations, avoids a large number of iterative processes in traditional diffusion models, and significantly reduces the amount of calculation and processing time.
[0121] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a portion of code, and the module, program segment, or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0123] In this application, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0124] In this application, memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0125] In this application, computer-readable media includes permanent and non-permanent, removable and non-removable storage media. Storage media can be implemented by any method or technology to store information, and the information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0126] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device that includes the element.
[0127] The above are merely specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to these embodiments herein, but is intended to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. An image denoising method based on a denoising diffusion probability model, characterized in that: include: Inputting the noisy image into a denoising diffusion probability model to obtain image noise output by the denoising diffusion probability model; The denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; A clean image is calculated based on the noisy image and the image noise.
2. The method according to claim 1, characterized in that The U-shaped convolutional neural network includes dimensional layers, wherein each dimensional layer includes a residual block that introduces a dual attention mechanism.
3. The method according to claim 1, characterized in that The denoising diffusion probability model is obtained by looping through the following steps: Randomly select clean sample images from the training dataset and determine a random time step; In the random time step, adding actual noise corresponding to the random time step to the clean sample image at each moment, so as to convert the clean sample image into a pure noise image; Inputting the pure noise image into a denoising diffusion probability model to obtain predicted noise; Calculating an objective function between the predicted noise and the actual noise; The model parameters of the denoising diffusion probability model are optimized based on the objective function until the objective function converges to obtain a trained denoising diffusion probability model.
4. The method according to claim 3, characterized in that The objective function is the square of the two-norm of the difference between the predicted noise and the actual noise; Optimizing the model parameters of the denoising diffusion probability model based on the objective function until the objective function converges to obtain a trained denoising diffusion probability model includes: Calculating the square of the delta-norm of the difference between the predicted noise and the actual noise; Calculating the gradient of the model parameter based on the square of the second norm of the difference; A gradient descent step is performed to update the model parameters until the square of the second norm of the difference converges to obtain a trained denoised diffusion probability model.
5. The method according to claim 1, wherein The calculating a clean image based on the noisy image and the image noise includes: determining a time step of the noisy image; If the time step of the noisy image is greater than time step 1, sampling noise from a standard normal distribution; Obtaining an image at a previous time step of the noisy image by calculation based on the noisy image, the time step of the noisy image, the image noise, and noise sampled from the standard normal distribution; Until the time step of the noisy image is equal to time step 1, a clean image is obtained.
6. The method according to claim 5, characterized in that The calculating, based on the noisy image, the time step of the noisy image, the image noise, and the noise sampled from the standard normal distribution, to obtain the image of the previous time step of the noisy image includes obtaining the image of the previous time step of the noisy image according to the following calculation formula: Among them, x t-1 is the image of the previous time step of the noisy image, x t is the noisy image, α t , and σ t is a parameter related to the time step t, ∈ θ (x t ,t) is the image noise, σ t is the standard deviation parameter associated with time step t, and z is the noise sampled from the standard normal distribution.
7. An image denoising device based on a denoising diffusion probability model, characterized in that: include: A noise prediction module, configured to input a noisy image into a denoising diffusion probability model to obtain image noise output by the denoising diffusion probability model; The denoising diffusion probability model is constructed based on a U-shaped convolutional neural network; The denoising module is configured to calculate a clean image based on the noisy image and the image noise.
8. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the image denoising method based on the denoising diffusion probability model as claimed in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that include: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image denoising method based on the denoising diffusion probability model according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that include: The computer program product includes a computer program. When the computer program is run on a computer, the computer is enabled to implement the image denoising method based on a denoising diffusion probability model according to any one of claims 1 to 6.
Citation Information
Cited By
Image generation method, image synthesis method, computing device and electronic device
CN121329804A