Image defogging method and device based on nonlinear noise sequence and spatial feature enhancement
By introducing nonlinear noise sequences and spatial feature enhancement modules into the image defog removal model, the problems of uneven fog removal and failure to fully consider human perception in the prior art are solved, and a stronger defog removal effect and a more efficient training process are achieved.
Patent Information
- Application Number
- CN202510149103.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing image defog removal model deals with large changes in the depth of the scene and complex lighting conditions, it has uneven fog removal or residual fog removal, and fails to fully consider the physical properties of human perception, which limits its information recovery ability.
Using an image defogging method based on nonlinear noise sequence and spatial feature enhancement, a nonlinear variance sequence is designed to simulate human visual perception mechanism by improving the DDPM model, and a spatial feature enhancement module is added to the U-Net network to extract deep spatial information of the image.
It realizes a clear image closer to the human perceived quality, has a stronger defog removal effect, can process feature information at different scales, and improves the training efficiency and defog removal effect of the model.
Smart Images

Figure CN120070236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an image dehazing method and device based on non-linear noise sequence and spatial feature enhancement. Background Art
[0002] In haze weather, various particles suspended in the air medium scatter and absorb the light propagating in the atmosphere, which will cause degradation problems such as low contrast and blurred color projection, greatly increasing the difficulty of image analysis and understanding. Therefore, image dehazing, as a part of the field of image restoration and enhancement, has a positive impact on image vision applications.
[0003] Currently, the methods for improving the quality of degraded images can be roughly divided into three categories: physical model-based methods, model-free methods, and deep learning-based methods. First, the physical model-based methods construct a mathematical model based on the imaging mechanism of remote sensing scenes and use handcrafted priors to estimate unknown coefficients to obtain clear images. These commonly used priors include dark channel prior, underwater dark channel prior, red channel prior, minimum information loss prior, etc., but they are only effective and robust for specific scenes. Secondly, the model-free methods improve the visual quality by adjusting pixel values and histogram distributions. However, due to ignoring the imaging mechanism, there are problems of over-saturation and color distortion in remote sensing images with complex physical properties.
[0004] Finally, the deep learning-based methods use powerful feature extraction capabilities to predict haze density and restore clear images. Among these methods, the denoising diffusion probabilistic model, i.e., the DDPM model, has attracted extensive attention for various regressors due to its powerful generative ability. However, most existing diffusion dehazing models sometimes encounter problems such as uneven dehazing or residual dehazing. In particular, when the scene depth changes greatly and the lighting conditions are complex, the feature information captured by the model during training is incomplete, resulting in imperfect dehazing. At the same time, the deep learning-based methods usually use a simple linear variance to add noise to train the model, and do not consider the physical properties of human perception, limiting their information completion ability. Therefore, how to design a more effective method for enhancing the image quality in haze weather has become an urgent technical problem to be solved. Summary of the Invention
[0005] In order to solve the limitations of existing image enhancement methods in terms of detail enhancement and resilience to interference from irrelevant features, resulting in problems such as blurred details and low contrast, the primary object of the present invention is to provide an image dehazing method based on non-linear noise sequence and spatial feature enhancement that can produce clear images closer to human perception quality and has a more powerful dehazing effect.
[0006] To achieve the above object, the present invention adopts the following technical solutions: An image defogging method based on non-linear noise sequence and spatial feature enhancement, the method comprising the following steps in sequence:
[0007] (1) Collect foggy images, preprocess the collected foggy images, the preprocessed images form a data set, and the data set is divided into a training set, a validation set and a test set according to a ratio;
[0008] (2) Construct an image defogging model: Improve the DDPM model to obtain the improved DDPM model, that is, the image defogging model;
[0009] (3) Input the training set into the image defogging model for training to obtain the trained image defogging model;
[0010] (4) Input the foggy image to be processed into the trained image defogging model, and output the defogged image.
[0011] In step (1), the collected foggy images are from the haze data sets RSID and NID, and the preprocessing includes cropping, rotation and data augmentation.
[0012] Step (2) specifically refers to: Optimize the linear growth function of the DDPM model, design a non-linear variance sequence to simulate the perception mechanism of human vision, make the distribution of noise close to the perception quality of humans, and obtain the optimized noise schedule β t ; Add a spatial feature enhancement module to the middle layer of the backbone network of the DDPM model, that is, the U-Net network.
[0013] Step (3) specifically includes the following steps in sequence:
[0014] (3a) Divide the training process into three stages, namely the first third stage, the middle third stage and the last third stage, and the noise schedules of the three stages are β 1 、β 2 、β 3 , obtain the optimized noise schedule β t ;
[0015] (3b) According to the optimized noise schedule β t , gradually add noise to the original image sample x 0 at different rates, and obtain x 1 , x 2 , ···, x t , ···, x T through T iterations;
[0016] (3c) The image defogging model starts from x TStart progressive denoising. During each step of denoising, the image dehazing model learns various details of adding noise from the forward process and continuously updates. The original low-quality images to be enhanced in the training set and the noisy image \(x\) at the \(t\)-th step t are used as the guiding condition to input into the image dehazing model. The backbone network of the image dehazing model, i.e., the U-Net network, first downsamples the image;
[0017] (3d) The downsampled image enters the spatial feature enhancement module to fully extract the detailed information of the image and obtain the feature blocks;
[0018] (3e) Upsample the feature blocks to gradually restore the spatial dimensions of the original low-quality image and the noisy image \(x\) at the \(t\)-th step t and reconstruct the detailed information of the original low-quality image and the noisy image \(x\) at the \(t\)-th step t ;
[0019] (3f) Use the reconstructed detailed information of the original low-quality image and the noisy image \(x\) at the \(t\)-th step t to perform reverse denoising;
[0020] (3g) Continuously optimize the weights during the process of continuously training the image dehazing model until the loss reaches the minimum, the image dehazing model reaches fitting, and the trained weights are obtained.
[0021] Step (3a) specifically refers to: The optimized noise schedule \(\beta\) t is:
[0022] \(\beta\) t =\(\beta\) 1 +\(\beta\) 2 +\(\beta\) 3
[0023] The noise schedule \(\beta\) in the first one-third of the training stage 1 is:
[0024] \(\beta\) 1 =[t*[(a*\(\beta\) e -\(\beta\) s ) / (t sum / 3)]+\(\beta\) s
[0025] t\in(0,t sum / 3)
[0026] The noise schedule \(\beta\) in the middle one-third of the training stage 2 is:
[0027] \(\beta\) 2 =[(t - t sum / 3)·(b - a)·beta e / (t sum / 3)+a·beta e
[0028] t ∈ (t sum / 3, 2t sum )
[0029] The noise schedule β for the last one - third training stage 3 is as follows:
[0030] β 3 =[(t - 2t sum / 3)·(1 - b)·beta e / (t sum / 3)+b·beta e
[0031] t ∈ (2t sum / 3, t sum )
[0032] In the formula, t represents the noise - adding or denoising time at the t - th step in the forward and backward processes, a represents the noise - schedule coefficient for the first one - third training stage, b represents the noise - schedule coefficient for the middle one - third training stage, beta e represents the final value of the noise schedule, beta s represents the starting value of the noise schedule, and t sum represents the total noise - adding time.
[0033] Step (3b) specifically refers to: For the original image sample x 0 , gradually add Gaussian noise to the image by iterating the formula (1) T times to obtain x 1 , x 2 , ···, x t , ···, x T ;
[0034]
[0035] Among them, x t represents the noise - added image at the t - th step, and x t is only related to x t-1 ; x T represents the finally obtained image close to pure noise; N represents the normal distribution, and α t represents the variance of the noise added in each iteration process, 1 - α t ∈(0, I), where I represents the identity matrix;
[0036] Using the re - parameterization method, the closed - form representation of formula (1) is:
[0037]
[0038] Among them, β t is the optimized noise schedule, β t = 1 - α t ; z t-1 is the noise conforming to the standard normal distribution;
[0039] Derive the formula (3) for obtaining x 0 from x t from formula (2):
[0040]
[0041] In the formula, represents the product of β 1 , β 2 ,..., β t ; z 0 represents the standard Gaussian noise;
[0042] For formula (1), through parameter renormalization, generate the distribution of the t-th step q(x t |x 0 ):
[0043]
[0044] Among them, represents the product of α 1 , α 2 ,..., α t .
[0045] Step (3d) specifically includes the following steps in sequence:
[0046] (3d1) First, downsample the image composed of the original low-quality image and the noisy image x t at the t-th step to obtain the tensor Y. Divide the tensor Y into four channel blocks, and perform a depthwise separable convolution operation on the tensor Y. The depthwise separable convolution operation includes two steps: depth convolution and pointwise convolution. In depth convolution, perform a convolution operation on the tensor Y using n convolutional kernels matching the tensor Y. In pointwise convolution, perform a point convolution operation to increase the number of channels, and pointwise convolution is called dimension enhancement;
[0047] (3d2) Divide the tensor with the number of channels tripled after dimension enhancement into three groups for different random angular spatial displacement operations. The first group of tensors and the second group of tensors undergo spatial displacement feature transformation, while the third group of tensors remains unchanged;
[0048] (3d3) The split attention mechanism is adopted to synthesize three groups of tensors together to maintain the same dimension and number of channels as when input;
[0049] (3d4) The output vector after the split attention mechanism is subjected to a residual connection with tensor Y, and the output of the spatial feature enhancement module is formula (4), and tensor Y is represented by formula (5):
[0050]
[0051]
[0052] Among them, downsample represents downsampling in the U-Net network, DSConv represents the depthwise separable convolution operation, T i represents randomly performing up, down, left, and right spatial displacement operations on a tensor with 4n channels, Synthesize represents the split attention mechanism, and Resconnect represents the residual connection.
[0053] Step (3f) specifically refers to: using the image x' to be reverse-denoised t and the original low-quality image as the constraint conditions to generate the image x' t-1 , that is, the reverse denoising process is:
[0054]
[0055] Among them, I is the identity matrix, N is the normal distribution; p θ is the probability formula for reverse denoising; δ t is the standard deviation related to time and noise scheduling; represents the output of the spatial feature enhancement module; μ θ is the average function learned from the parameters θ of the image dehazing model:
[0056]
[0057] Among them, is the noise parameter, β t is the optimized noise scheduling, α t represents the variance of the noise added during each iteration;
[0058] The parameters θ of the image dehazing model are optimized based on the loss function L(θ):
[0059]
[0060] Among them, x t is the noisy image at the t-th step; represents x' t , and the mathematical expectation obtained from the constraint condition; ε represents the noise parameter obtained in the forward noise addition of the image defogging model.
[0061] Another object of the present invention is to provide an electronic device, including:
[0062] a processor; and
[0063] a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the image defogging method based on non-linear noise sequence and spatial feature enhancement as described above.
[0064] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the image defogging method based on non-linear noise sequence and spatial feature enhancement as described above.
[0065] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention proposes a non-linear variance sequence inspired by the perception mechanism of the human visual system, which uses a set variable slope to guide the model to focus on specific information in corresponding stages in different stages, such as the overall structure and texture, etc., which is beneficial for the model to understand image information more precisely and comprehensively, and pays more attention to the human perception ability in the process of forward diffusion and reverse denoising, so as to generate clearer images closer to the human perception quality; Second, the present invention proposes an optimized spatial feature enhancement module, which uses a modified spatial displacement network to capture enhanced spatial features, obtains deeper spatial information of the image, and performs residual connection with the original features to strengthen the transmission of features and the local association of image features, so that the feature learning of the image defogging model based on non-linear and spatial feature enhancement of the present invention is more efficient, can process feature information of different scales, combines multi-scale information to understand the image more deeply and comprehensively, speeds up the training speed of the model, improves the training efficiency of the model, and has a more powerful defogging effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of the method for training an image defogging model in the present invention;
[0067] Figure 2 is a visual comparison result diagram. DETAILED DESCRIPTION
[0068] An image defogging method based on non-linear noise sequence and spatial feature enhancement, the method includes the following steps in sequence:
[0069] (1) Collect foggy images, preprocess the collected foggy images, and the preprocessed images form a dataset. The dataset is divided into a training set, a validation set, and a test set according to a certain proportion;
[0070] (2) Build an image defogging model: Improve the DDPM model to obtain the improved DDPM model, which is the image defogging model;
[0071] (3) Input the training set into the image defogging model for training to obtain the trained image defogging model;
[0072] (4) Input the foggy image to be processed into the trained image defogging model, and output the defogged image.
[0073] In step (1), the collected foggy images come from the haze datasets RSID and NID, and the preprocessing includes cropping, rotation, and data augmentation.
[0074] Step (2) specifically refers to: Optimize the linear growth function of the DDPM model, design a non-linear variance sequence to simulate the perception mechanism of human vision, make the distribution of noise close to the perception quality of humans, and obtain the optimized noise schedule β t ; Add a spatial feature enhancement module to the middle layer of the backbone network of the DDPM model, that is, the U-Net network. The purpose is to extract the image spatial features again at the place where the image features in the middle layer are most obvious, so that the image defogging model can better learn the spatial detail information of the image during the noise addition process in the forward process and more effectively adjust the variance and mean estimated during the sampling process.
[0075] As Figure 1 shown, step (3) specifically includes the following steps in sequence:
[0076] (3a) Divide the training process into three stages, namely the first third stage, the middle third stage, and the last third stage. The noise schedules for the three stages are β 1 、β 2 、β 3 , and obtain the optimized noise schedule β t ;
[0077] (3b) According to the optimized noise schedule β t , gradually add noise to the original image sample x 0 at different rates, and obtain x 1 , x 2 , ···, x t , ···, x T through T iterations;
[0078] (3c) The image defogging model starts from x TStart progressive denoising. During each step of denoising, the image dehazing model learns and continuously updates various details of adding noise from the forward process, and uses the original low-quality images to be enhanced in the training set and the noisy image x at the t-th step t as the guiding condition to input into the image dehazing model. The backbone network of the image dehazing model, i.e., the U-Net network, first performs downsampling on the image;
[0079] (3d) The downsampled image enters the spatial feature enhancement module to fully extract the detailed information of the image and obtain the feature blocks;
[0080] (3e) Upsample the feature blocks to gradually restore the spatial dimensions of the original low-quality image and the noisy image x at the t-th step t and reconstruct the detailed information of the original low-quality image and the noisy image x at the t-th step t ;
[0081] (3f) Use the detailed information of the reconstructed original low-quality image and the noisy image x at the t-th step t to perform reverse denoising;
[0082] (3g) Continuously optimize the weights during the process of continuously training the image dehazing model until the loss reaches the minimum, the image dehazing model reaches fitting, and the trained weights are obtained.
[0083] Step (3a) specifically refers to: the optimized noise schedule β t is:
[0084] β t =β 1 +β 2 +β 3
[0085] When the image is denoised or noise is added in the first third stage, the growth function increases rapidly, that is, the noise in the initial stage is rapidly enhanced. The rapid growth of the noise in the initial stage enables the model to quickly blur the image, providing more information for the model to perform reverse denoising in subsequent steps, so as to better restore the clear image. The noise schedule β in the first third training stage 1 is:
[0086] β 1 =[t*[(a*beta e -beta s ) / (t sum / 3)]·+beta s
[0087] t∈(0,tsum / 3)
[0088] When denoising or adding noise to the image in the middle one-third stage, the growth function starts to slow down the growth rate, that is, the noise growth slightly decreases. Controlling the noise addition speed helps the model gradually adjust the image details, restore the intermediate frequency information, and provide a good transition for the subsequent reverse denoising process. The noise schedule β in the middle one-third training stage 2 is as follows:
[0089] β 2 = [(t - t sum / 3)·(b - a)·beta e / (t sum / 3)+a·beta e
[0090] t ∈ (t sum / 3, 2t sum / 3)
[0091] Finally, when denoising or adding noise to the image in the last one-third stage, the change rate of the growth function reaches the lowest, that is, the noise changes the slowest. The slow change of noise allows the image defogging model to spend more time learning the image that is closest to the noise to the greatest extent to solve the problems in the inverse process of the image defogging model, and allows the model to focus on restoring the high-frequency details such as edges and texture information that significantly improve the contrast and clarity of the defogged image. The noise schedule β in the last one-third training stage 3 is as follows:
[0092] β 3 = [(t - 2t sum / 3)·(1 - b)·beta e / (t sum / 3)+b·beta e
[0093] t ∈ (2t sum / 3, t sum )
[0094] In the formula, t represents the moment of adding noise or denoising at the t-th step in the forward and backward processes, a represents the noise schedule coefficient in the first one-third training stage, b represents the noise schedule coefficient in the middle one-third training stage, beta e represents the final value of the noise schedule, beta s represents the starting value of the noise schedule, and t sum represents the total noise addition time.
[0095] Step (3b) specifically refers to: for the original image sample x 0 , gradually add Gaussian noise to the image by iterating formula (1) T times to obtain x 1 , x2 , ···, x t , ···, x T ;
[0096]
[0097] where x t represents the noisy image at the t-th step, and x t is only related to x t-1 ; x T represents the finally obtained image close to pure noise; N represents the normal distribution, and α t represents the variance of the added noise in each iteration process, and 1 - α t ∈(0, I), where I represents the identity matrix;
[0098] Using the reparameterization method, the closed-form representation of formula (1) is:
[0099]
[0100] where β t is the optimized noise schedule, and β t = 1 - α t ; z t-1 is the noise conforming to the standard normal distribution;
[0101] Derive formula (3) for obtaining x 0 from x t from formula (2):
[0102]
[0103] In the formula, represents the product of β 1 , β 2 ,..., β t ; z 0 represents the standard Gaussian noise;
[0104] For formula (1), through parameter renormalization, the distribution of q(x t |x 0 ) at the t-th step is generated:
[0105]
[0106] where represents the product of α 1 , α 2 ,..., α t .
[0107] Step (3d) specifically includes the following steps in sequence:
[0108] (3d1) First, downsample the image composed of the original low-quality image and the noisy image x at the t-th step t to obtain the tensor Y. Divide the tensor Y into four channel blocks, and perform a depthwise separable convolution operation on the tensor Y. The depthwise separable convolution operation includes two steps: depth convolution and pointwise convolution. In depth convolution, perform a convolution operation on the tensor Y using n convolution kernels matching the tensor Y. In pointwise convolution, perform a point convolution operation to increase the number of channels, and pointwise convolution is called dimension enhancement;
[0109] (3d2) Divide the tensor with the number of channels tripled after dimension enhancement into three groups for different random angular spatial displacement operations. The first group of tensors and the second group of tensors undergo spatial displacement feature transformation, while the third group of tensors remains unchanged;
[0110] (3d3) Adopt a split attention mechanism to combine the three groups of tensors together to maintain the same dimensions and number of channels as at the input;
[0111] (3d4) Perform a residual connection between the output vector after the split attention mechanism and the tensor Y. The output of the spatial feature enhancement module is formula (4), and the tensor Y is represented by formula (5):
[0112]
[0113]
[0114] where, downsample represents downsampling in the U-Net network, DSConv represents the depthwise separable convolution operation, T i represents randomly performing up, down, left, and right spatial displacement operations on a tensor with 4n channels, Synthesize represents the split attention mechanism, and Resconnect represents the residual connection.
[0115] Step (3f) specifically refers to: using the image x' to be reverse denoised t and the original low-quality image as constraint conditions to generate the image x' t-1 , that is, the reverse denoising process is:
[0116]
[0117] θ is the probability formula for reverse denoising; δ t
[0118] is the standard deviation related to time and noise scheduling; represents the output of the spatial feature enhancement module;
[0119] μ θ is the average function learned from the parameters θ of the image dehazing model:
[0120]
[0121] where is the noise parameter, β t is the optimized noise schedule, α t represents the variance of the added noise in each iteration;
[0122] The parameters θ of the image dehazing model are optimized based on the loss function L(θ):
[0123]
[0124] where x t is the noisy image at the t-th step; represents x' t , and is the mathematical expectation obtained from the constraint conditions; ε represents the noise parameter already obtained in the forward noise addition of the image dehazing model.
[0125] In Figure 2 , the original hazy image is shown as (a) in Figure 2 , the hazy image processed by the prior art GT is shown as (b) in Figure 2 , the hazy image processed by the prior art FDP is shown as (c) in Figure 2 , the hazy image processed by the prior art GDCP is shown as (d) in Figure 2 , the hazy image processed by the prior art DAUMR is shown as (e) in Figure 2 , the hazy image processed by the prior art DUBCDP is shown as (f) in Figure 2 , the hazy image processed by the prior art AODNet is shown as (g) in Figure 2 , the hazy image processed by the prior art DchazeNet is shown as (h) in Figure 2 , the hazy image processed by the prior art SCANet is shown as (i) in Figure 2 , the hazy image processed by the prior art EMPFNet is shown as (j) in Figure 2 , the hazy image processed by the prior art RSHazeNet is shown as (k) in Figure 2 , the hazy image processed by the method proposed in the present invention is shown as (l) in Figure 2 . It can be seen that the present invention has a more powerful dehazing effect.
[0126] In summary, the present invention proposes a non-linear variance sequence inspired by the human visual system's perception mechanism, which uses a variable slope to guide the model to focus on specific information at different stages, such as the overall structure and texture, etc.; the present invention pays more attention to human perception ability in forward diffusion, thereby generating clearer images closer to human perception quality; the present invention proposes an optimized spatial feature enhancement module, which uses a modified spatial displacement network-based method to capture enhanced spatial features and integrate them with the original features to strengthen feature transmission and local correlation of image features, thereby preventing gradient disappearance and guiding the denoising process; the present invention can activate vanishing gradients and enhance the feature completion ability.
Claims
1. An image defogging method based on nonlinear noise sequence and spatial feature enhancement, characterized in that: The method comprises the following steps in order: (1) Collect foggy images, preprocess the collected foggy images, form a data set after preprocessing, and divide the data set into a training set, a validation set, and a test set in proportion; (2) Constructing an image defogging model: improving the DDPM model to obtain an improved DDPM model, i.e., an image defogging model; (3) Inputting the training set into the image defogging model for training to obtain a trained image defogging model; (4) The foggy image to be processed is input into the trained image defogging model, and the defogged image is output.
2. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 1, characterized in that: In step (1), the collected foggy images are from the haze datasets RSID and NID, and the preprocessing includes cropping, rotation and data enhancement.
3. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 1, characterized in that: Step (2) specifically refers to: optimizing the linear growth function of the DDPM model, designing a nonlinear variance sequence to simulate the perception mechanism of human vision, making the distribution of noise close to the human perception quality, and obtaining the optimized noise scheduling β t ; A spatial feature enhancement module is added to the middle layer of the U-Net network, which is the backbone network of the DDPM model.
4. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 1, characterized in that: Step (3) specifically includes the following steps in order: (3a) The training process is divided into three stages, namely the first third stage, the middle third stage and the last third stage. The noise schedules of the three stages are β1, β2 and β3 respectively. The optimized noise schedule β t ; (3b) According to the optimized noise scheduling β t , gradually add noise to the original image sample x0 at different rates, and obtain x1, x2, ···, x through T iterations t ,···,x T ; (3c) Image dehazing model from x T The image defogging model starts to denoise step by step. During each step of denoising, the image defogging model learns the details of the noise from the forward process and continuously updates it, and the original low-quality image to be enhanced in the training set is and the noisy image x at step t t The image dehazing model is input as a guiding condition, and the backbone network of the image dehazing model, namely the U-Net network, first downsamples the image; (3d) The downsampled image enters the spatial feature enhancement module to fully extract the image detail information and obtain the feature block; (3e) Upsample the feature blocks and gradually restore the original low-quality image and the noisy image x at step t t spatial dimensions, reconstructing the original low-quality image and the noisy image x at step t t Detailed information; (3f) Using the reconstructed original low-quality image and the noisy image x at step t t Detailed information is collected and reverse denoised; (3g) In the process of continuously training the image dehazing model, the weights are continuously optimized, the loss is minimized, the image dehazing model is fitted, and the trained weights are obtained.
5. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 4, characterized in that: Step (3a) specifically refers to: optimizing the noise scheduling β t for: β t =β1+β2+β3 The noise schedule β1 for the first third of the training phase is: β1=[t*[(a*beta e -beta s ) / (t sum / 3)]+beta s ]t∈(0,t sum / 3) The noise schedule β2 in the middle third of the training phase is: β2=[(t-t sum / 3)·(b-a)·beta e ] / (t sum / 3)+a·beta e t∈(t sum / 3,2t sum / 3) The noise schedule β3 in the last third of the training phase is: β3=[(t-2t sum / 3)·(1-b)·beta e ] / (t sum / 3)+b·beta e t∈(2t sum / 3,t sum ) Where t represents the t-th step of noise addition or denoising in the forward and backward processes, a represents the noise scheduling coefficient in the first one-third of the training stage, b represents the noise scheduling coefficient in the middle one-third of the training stage, and beta e represents the final value of the noise schedule, beta s represents the starting value of noise scheduling, t sum Indicates the total noise addition time.
6. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 4, characterized in that: Step (3b) specifically means: for the original image sample x0, gradually add Gaussian noise to the image by iterating formula (1) T times to obtain x1, x2, ···, x t ,···,x T ; Among them, x t represents the noisy image at step t, x t Only with x t-1 Related; x T represents the final image close to pure noise; N represents normal distribution, α t Indicates the variance of the noise added during each iteration, 1-α t ∈(0,I), I represents the identity matrix; Using the reparameterization method, the closed form of formula (1) is expressed as: Among them, β t For the optimized noise scheduling, β t =1-α t ;z t-1 is the noise that conforms to the standard normal distribution; From formula (2), we can deduce that from x0 we get x t Formula (3): In the formula, denote β1, β2, ..., β t The cumulative multiplication of; z0 represents standard Gaussian noise; For formula (1), the parameters are renormalized and the t-th step q(x t |x0) distribution: in, represents α1, α2, ..., α t The cumulative multiplication of .
7. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 4, characterized in that: Step (3d) specifically comprises the following steps in order: (3d1) First, the original low-quality image and the noisy image x at step t t The composed image is downsampled to obtain tensor Y, which is divided into four channel blocks. A depth-wise separable convolution operation is performed on tensor Y. The depth-wise separable convolution operation includes two steps: depth-wise convolution and point-wise convolution. In depth-wise convolution, a convolution operation is performed on tensor Y using n convolution kernels matching tensor Y. In point-wise convolution, a point convolution operation is performed to increase the number of channels. Point-wise convolution is called dimensionality enhancement. (3d2) The tensors whose channels are tripled after dimension enhancement are divided into three groups for different random angle spatial displacement operations. The first group of tensors and the second group of tensors undergo spatial displacement feature transformation, while the third group of tensors remains unchanged. (3d3) A split attention mechanism is used to synthesize the three sets of tensors together to maintain the same dimension and number of channels as the input; (3d4) The output vector after the split attention mechanism is connected to the tensor Y by residual connection. The output of the spatial feature enhancement module is formula (4), and the tensor Y is expressed by formula (5): Among them, downsample represents the downsampling in the U-Net network, DSConv represents the depth-separable convolution operation, and T i represents randomly performing up, down, left, and right spatial shift operations on a tensor with 4n channels, Synthesize represents the split attention mechanism, and Resconnect represents the residual connection.
8. The image defogging method based on nonlinear noise sequence and spatial feature enhancement according to claim 4, characterized in that: Step (3f) specifically refers to: taking the image x' to be reverse denoised t Compared with the original low quality image Generate image x' for the constraints t-1 , that is, the reverse denoising process is: Where I is the unit matrix, N is the normal distribution; p θ is the probability formula for reverse denoising; δ t is the standard deviation associated with the time and noise schedule; Represents the output of the spatial feature enhancement module; μ θ is the average function learned from the parameters θ of the image dehazing model: in, is the noise parameter, β t For the optimized noise scheduling, α t Represents the variance of the noise added during each iteration; The parameters θ of the image dehazing model are optimized based on the loss function L(θ): Among them, x t is the noisy image at step t; Represents x' t , as well as is the mathematical expectation obtained by the constraint conditions; ε represents the noise parameter obtained in the forward denoising of the image dehazing model.
9. An electronic device, comprising: processor; as well as A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the image defogging method based on nonlinear noise sequence and spatial feature enhancement as described in any one of claims 1 to 8.
10. A computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the processor is enabled to execute the image defogging method based on nonlinear noise sequence and spatial feature enhancement as described in any one of claims 1 to 8.
Citation Information
Cited By
Image restoration method and device based on residual denoising diffusion model
CN121563847A