Generalized image restoration apparatus based on contour prior guided mamba diffusion

CN122391018BActive Publication Date: 2026-09-08INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610848912.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-08
Estimated Expiration
2046-06-12

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请实施例提供了一种基于轮廓先验引导的Mamba扩散的通用图像恢复模型,以解决现有技术中图像恢复时使用的先验信息大多自图像退化信息中提取得到,导致图像恢复性能首受限的问题

Benefits of technology

[0026] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment extracts contour prior information from the self-degrading image using CPGMD and encodes it to obtain contour prior features; it performs feature interaction between the contour prior features and input features to obtain spatially enhanced interactive features, and then performs channel attention enhancement on the spatially enhanced interactive features and contour prior features to obtain channel-enhanced interactive features; next, it uses the channel-enhanced interactive features to generate prediction noise for each time step in DDPM, and finally uses DDPM based on this prediction noise to recover a clean image from the degraded image. This image restoration model has high robustness and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391018B_ABST
    Figure CN122391018B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, and provides a general image restoration model based on contour prior guided Mamba diffusion. The model extracts contour prior information from a self-degraded image through CPGMD, and encodes the contour prior information to obtain contour prior features; the contour prior features and input features are subjected to feature interaction to obtain spatially reinforced interaction features, and the spatially reinforced interaction features and the contour prior features are subjected to channel attention enhancement to obtain channel-reinforced interaction features; next, the channel-reinforced interaction features are used to generate prediction noise of each time step in a DDPM; finally, the DDPM is used to restore a clean image from the degraded image based on the prediction noise. The image restoration model has high robustness and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a general image restoration model based on contour prior-guided Mamba diffusion. Background Technology

[0002] Image restoration aims to recover high-quality images from degraded inputs. In real-world scenarios, images are frequently affected by various degradation factors, including visual degradation such as low light, underwater attenuation, and noise, as well as semantic degradation such as adversarial perturbations. These degradation factors not only reduce the perceptual quality of images but also impair the robustness of advanced visual tasks such as classification, detection, and segmentation. Therefore, developing restoration methods with strong generalization capabilities is crucial for improving the reliability of computer vision systems.

[0003] Traditional image restoration models rely on hand-designed physical priors, which are interpretable and effective in specific image degradation scenarios. However, they often lack robustness and generalization ability when faced with the complex and diverse image degradation in the real world. With the advent of deep learning, data-driven methods have become the mainstream paradigm for image restoration. Convolutional Neural Network (CNN)-based methods significantly improve restoration quality by learning a direct mapping from degraded to clean images end-to-end. However, their limited receptive field restricts the modeling of long-range dependencies. Generative Adversarial Networks (GAN)-based methods enhance perceptual realism through adversarial training. However, their training process is inherently unstable and prone to unpredictable artifacts.

[0004] In recent years, Denoising Diffusion Probabilistic Models (DDPM) have become one of the most popular generative frameworks due to their stable training process and high-fidelity output. In image restoration, DDPM-based models restore images through a progressive denoising process conditioned on degraded inputs and often incorporate external prior information to improve restoration quality. However, in most existing diffusion-based methods, whether based on visual or textual priors, this prior information is extracted directly from the degraded input. Since degradation artifacts are not explicitly decoupled, this prior information inevitably retains degradation-related information, reducing the reliability of guidance and limiting the restoration performance of degraded images. This problem is exacerbated in cases of severe degradation, as damaged information significantly contaminates the guidance prior. Therefore, extracting degradation-independent prior information is crucial for ensuring reliable guidance and robust image restoration. Summary of the Invention

[0005] In view of this, embodiments of this application provide a general image restoration model based on contour prior-guided Mamba diffusion to solve the problem that the prior information used in image restoration in the prior art is mostly extracted from image degradation information, which leads to the limitation of image restoration performance.

[0006] A first aspect of this application provides a general image restoration model based on contour prior-guided Mamba diffusion, including a general image restoration framework based on contour prior-guided Mamba diffusion (CPGMD) and DDPM.

[0007] CPGMD includes:

[0008] The contour prior extraction module is configured to extract contour prior information from self-degenerate images;

[0009] The contour prior encoding module is configured to encode the extracted contour prior information to obtain contour prior features; the contour prior features are high-dimensional features and the distribution density is greater than a preset density threshold.

[0010] The contour prior spatial perception module is configured to perform feature interaction between contour prior features and input features to obtain spatially enhanced interactive features; wherein, the input features include the input features of the degraded image and the input features of the image after adding noise at the current time step, and the image after adding noise at the current time step is determined by the DDPM model;

[0011] The contour prior channel perception module is configured to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain the channel-enhanced interactive features.

[0012] The noise prediction module is configured to generate predicted noise for each time step in DDPM based at least on the channel-enhanced interaction features.

[0013] DDPM is configured to recover a clean image from a degraded image based at least on predicted noise.

[0014] A second aspect of this application provides a general image restoration method based on contour-prior-guided Mamba diffusion, comprising:

[0015] Obtain the degraded image;

[0016] Use CPGMD to generate the prediction noise for each time step in DDPM;

[0017] Using DDPM, a clean image can be recovered from a degraded image based at least on the predicted noise;

[0018] CPGMD generates the prediction noise for each time step in DDPM in the following manner:

[0019] Contour prior information is extracted from self-degenerate images using the contour prior extraction module;

[0020] The extracted contour prior information is encoded using the contour prior encoding module to obtain contour prior features; the contour prior features are high-dimensional features and their distribution density is greater than a preset density threshold.

[0021] The contour prior spatial perception module is used to perform feature interaction between the contour prior features and the input features to obtain spatially enhanced interactive features; among them, the input features include the input features of the degraded image and the input features of the image after adding noise at the current time step, and the image after adding noise at the current time step is determined by the DDPM model;

[0022] The contour prior channel perception module is used to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain the channel-enhanced interactive features.

[0023] The noise prediction module is used to generate predicted noise for each time step in DDPM based at least on the channel-enhanced interaction features.

[0024] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0025] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0026] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment extracts contour prior information from the self-degrading image using CPGMD and encodes it to obtain contour prior features; it performs feature interaction between the contour prior features and input features to obtain spatially enhanced interactive features, and then performs channel attention enhancement on the spatially enhanced interactive features and contour prior features to obtain channel-enhanced interactive features; next, it uses the channel-enhanced interactive features to generate prediction noise for each time step in DDPM, and finally uses DDPM based on this prediction noise to recover a clean image from the degraded image. This image restoration model has high robustness and generalization ability. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of the structure of a general image restoration model based on contour prior-guided Mamba diffusion according to an embodiment of this application.

[0029] Figure 2 This is a schematic diagram of the structure of the CPSAM module provided in the embodiments of this application.

[0030] Figure 3 This is a schematic diagram of the CPCAM module provided in the embodiments of this application.

[0031] Figure 4 This is a comparison chart of the restoration performance of the model provided in this application on the underwater image enhancement dataset LSUI with the restoration performance of other models.

[0032] Figure 5 This is a comparison chart of the restoration performance of the model provided in this application on the low-light image enhancement dataset LOLv2_syn with the restoration performance of other models.

[0033] Figure 6 This is a comparison chart of the restoration effect of the model provided by the embodiments of this application on the image denoising dataset BSD68 with the restoration effect of other models.

[0034] Figure 7 This is a schematic flowchart of a general image restoration method based on contour prior-guided Mamba diffusion provided in an embodiment of this application.

[0035] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0036] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0037] The following will describe in detail, with reference to the accompanying drawings, a general image restoration model and method based on contour-guided Mamba diffusion according to embodiments of this application.

[0038] As mentioned above, in image restoration, DDPM-based models recover images through a progressive denoising process conditioned on degraded inputs and typically incorporate external prior information to improve restoration quality. However, in most existing diffusion-based methods, whether based on visual or textual priors, this prior information is extracted directly from the degraded input. Since degradation artifacts are not explicitly decoupled, this prior information inevitably retains degradation-related details, reducing the reliability of the guidance and limiting the restoration performance of degraded images. This problem is exacerbated in cases of severe degradation, as damaged information significantly contaminates the guidance prior.

[0039] Therefore, embodiments of this application consider extracting prior information unrelated to degradation for image restoration.

[0040] Among various prior information sources, structural priors demonstrate particular effectiveness in mitigating degradation and preserving fine details, as evidenced by the Low-light image enhancement via structure modeling and guidance (SMG) method and the Semantic and structure priors for diffusion-based realistic image restoration (SSP-IR) method. SMG employs an encoder-decoder based on analyzing and improving the image quality of StyleGAN (StyleGAN) to extract structural priors from degraded images, while SSP-IR uses pixel-level processors to achieve the same goal. However, both methods rely on implicit extraction, making it difficult to distinguish between structure-related information and degradation artifacts.

[0041] Building upon this foundation, this application proposes a contour-prior-guided Mamba diffusion model for general image restoration. Considering that various degradation factors such as underwater distortion, low-light conditions, noise, and adversarial perturbations significantly alter the color distribution, brightness, and fine texture of an image, while having minimal impact on overall structural features such as object boundaries and scene geometry, CPGMD extracts contour information, independent of specific degradation types, from the degraded image as a prior and integrates it into the diffusion noise prediction process using a specially designed spatial and channel-aware module. This prior information provides a stable guiding signal, helping to maintain content consistency and fine details.

[0042] Furthermore, this application provides a general image restoration model based on contour prior-guided Mamba diffusion. Contour prior information is extracted from the self-degraded image using CPGMD and encoded to obtain contour prior features. Feature interaction is performed between the contour prior features and the input features to obtain spatially enhanced interactive features. Then, channel attention enhancement is applied to the spatially enhanced interactive features and the contour prior features to obtain channel-enhanced interactive features. Next, the channel-enhanced interactive features are used to generate prediction noise for each time step in DDPM. Finally, DDPM is used to recover a clean image from the degraded image based on this prediction noise. This image restoration model exhibits high robustness and generalization ability.

[0043] Figure 1 This is a schematic diagram of the structure of a general image restoration model based on contour prior-guided Mamba diffusion provided in an embodiment of this application. Figure 1 As shown, the model includes CPGMD and DDPM, with DDPM in the upper dashed box and CPGMD in the lower dashed box.

[0044] In some embodiments of this application, CPGMD includes a contour prior extraction module, a contour prior encoding module, a contour prior spatial perception module, a contour prior channel perception module, and a noise prediction module.

[0045] The contour prior extraction module is configured to extract contour prior information from the self-degenerate image.

[0046] The contour prior encoding module is configured to perform feature encoding on the extracted contour prior information to obtain contour prior features; these contour prior features are high-dimensional features with a distribution density greater than a preset density threshold. Here, high-dimensional features refer to feature representations composed of multi-channel feature vectors.

[0047] The contour prior spatial awareness module is configured to perform feature interaction between contour prior features and input features to obtain spatially enhanced interactive features. The input features include input features of the degraded image and input features of the noisy image at the current time step, with the noisy image at the current time step determined by the DDPM model.

[0048] The contour prior channel perception module is configured to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain the channel-enhanced interactive features.

[0049] The noise prediction module is configured to generate predicted noise for each time step in DDPM based at least on the channel-enhanced interaction features.

[0050] Furthermore, DDPM is configured to recover a clean image from a degraded image based at least on predicted noise.

[0051] In some embodiments of this application, extracting contour prior information from a degraded image may include: segmenting the degraded image using an image segmentation model to obtain a segmented image of the degraded image; converting the segmented image into a grayscale image and determining the horizontal and vertical gradient components of the grayscale image; calculating the edge intensity value of each pixel in the degraded image based on the horizontal and vertical gradient components; determining pixels with edge intensity values ​​higher than a preset intensity threshold and pixels connected to strong edge pixels as edge pixels; and determining the binary contour map composed of each edge pixel as contour prior information.

[0052] Strong edge pixels can be determined as follows: calculate the gradient value of each pixel, and determine the pixels whose gradient value is greater than a preset gradient threshold as strong edge pixels.

[0053] In other words, in the CPGMD provided in the embodiments of this application, the input degraded image... A contour prior extraction module based on a segmentation model was constructed. This module requires no additional training and can extract structurally stable contour information from degraded images under zero-shot conditions, which can be used for structural constraints and artifact suppression in the subsequent diffusion-based restoration process.

[0054] Specifically, the contour prior extraction module first uses an image segmentation model (such as the SAM model) to process the degraded image. Perform segmentation processing to obtain its segmented image. The image segmentation model can be any segmentation model (SAM) or any other image segmentation model; there are no restrictions here.

[0055] The segmented image can then be converted to a grayscale image. and calculate respectively Gradient components in the horizontal and vertical directions and This is to reflect local intensity changes in the image.

[0056] Next, we can further calculate the edge intensity value of each pixel based on the above gradient components. ;in, Let be the coordinates of any pixel.

[0057] Finally, a dual-threshold hysteresis strategy can be used to filter edge responses: when the edge intensity value of a pixel is higher than a preset emphasis threshold, or when a pixel is connected to a strong edge pixel, it can be identified as an edge pixel, thereby obtaining prior contour information represented by a binarized contour map. .

[0058] The process of extracting prior information about the contour is as follows: Figure 1 The leftmost module is shown in the dotted box at the bottom center.

[0059] In some embodiments of this application, feature encoding is performed on the extracted contour prior information to obtain contour prior features, which may include: constructing an N-layer convolutional structure; and using the N-layer convolutional structure to perform feature encoding on the contour prior information to obtain contour prior features. Here, N is a positive integer greater than or equal to 3 and less than or equal to 6.

[0060] While the extracted contour prior information can effectively suppress degradation noise and provide a reliable structural reference, it is inherently sparse and has limited semantic information. To enhance its usability in the recovery model, this application further sets up a contour prior encoding module, using an N-layer convolutional structure. right Feature encoding is performed to obtain high-dimensional semantic representation and densely distributed contour prior features: ;in For contour prior features, This represents the contour prior encoding operator.

[0061] Through the above steps, the embodiments of this application can automatically extract stable structural contour priors from degraded images and convert them into semantic feature representations suitable for diffusion models, thereby effectively improving the ability to preserve structure and reducing artifact propagation during image restoration.

[0062] The process of feature encoding the extracted contour prior information is as follows: Figure 1 The middle module is shown in the dotted box at the bottom center.

[0063] In some embodiments of this application, feature interaction is performed on contour prior features and input features to obtain spatially enhanced interactive features. This may include: mapping the input features using the MambaScan module to obtain spatial features; performing linear projection on the spatial features to obtain a query vector; performing average pooling compression on the contour prior features to obtain pooled compressed features; generating key vectors and value vectors from the pooled compressed features through linear mapping; and using a scaled dot product attention mechanism to perform feature interaction on the query vector, key vector, and value vector to obtain spatially enhanced interactive features.

[0064] In other words, to effectively integrate degradation-independent contour priors into spatial feature modeling, this application constructs a Contour Prior Spatial-Aware Module (CPSAM). This CPSAM module, through a feature interaction mechanism guided by contour priors, enables the network to focus on salient contour regions in the spatial domain, avoiding the propagation of degradation artifacts and thus improving the ability to recover fine structures.

[0065] Specifically, the inputs to CPSAM include: spatial features output from the Mamba scan structure. and contour prior features .

[0066] in, Input features can be processed using the MambaScan module. Spatial features are obtained by performing spatial feature mapping. The MambaScan module is the Scan module in the Mamba visual model.

[0067] Furthermore, this spatial feature can be... Perform linear projection to obtain the query vector. This process can be achieved using a learnable projection matrix. This is used to establish subsequent attentional interactions. .

[0068] Next, the contour prior features can be used to construct the key and value. Considering that directly involving the contour prior features in attention calculation would result in a large computational load, this embodiment of the application processes the contour prior features before constructing the key and value. Perform adaptive average pooling compression.

[0069] Specifically, the kernel size can be... Adaptive pooling is performed to obtain compressed features for the segmented regions. Subsequently, the pooled features are used to generate key vectors through linear mapping. AND value vector Its calculation form can be described as:

[0070] ;

[0071] ;

[0072] in, and These are the height and width of the feature image, respectively. and All are learnable parameters. This is the set of pixels in each spatial region after the pooling operation. and The first line, number Column spatial location index, For the prior feature map of the contour in spatial location The feature vector at that location. By adjusting the pooling kernel size... By comparing performance, the optimal compression scale can be obtained.

[0073] After obtaining the query ,key Sum Subsequently, the CPSAM module can employ a scaled dot product attention mechanism for feature interaction. Its output features are obtained using the following formula: ;in, For spatially enhanced interactive features, It is a linear transformation function. For normalized exponential functions, The dimension of the key vector is used to scale the dot product result for stable training.

[0074] Through the aforementioned attention mechanism, CPSAM can leverage contour priors to enhance spatially salient regions, improve the guidance capability of fine structural information, and minimize the interference of degenerate features in the recovery process. The specific structure of the CPSAM module is as follows: Figure 2 As shown.

[0075] In some embodiments of this application, channel attention enhancement is performed on spatially enhanced interaction features and contour prior features to obtain channel-enhanced interaction features. This may include: inputting the spatially enhanced interaction features and the current time step into a multilayer perceptron to obtain perceptual features; performing global average pooling on the contour prior features in the spatial dimension to obtain channel descriptor vectors; performing a linear transformation on the channel descriptor vectors and using the SiLU activation function to generate channel attention weights based on the linear transformation result; the value range of the channel attention weights is [0,1]; and performing channel-wise weighting on the perceptual features and channel attention weights to obtain channel-enhanced interaction features.

[0076] In other words, this application also constructs a Contour Prior Channel-Aware Module (CPCAM) to guide feature selection along the channel dimension, enabling the model to prioritize enhancing feature channels carrying contour-related semantic information, thereby further improving the ability to preserve contour details. This module uses contour prior features... As a channel-level semantic guide, it achieves selective channel enhancement through channel statistics and adaptive weighting mechanisms.

[0077] First, we can examine the prior features of the contour. Global average pooling is performed along the spatial dimension to obtain a compact channel descriptor vector. Its calculation form is as follows: This vector represents the importance of each channel in the contour prior.

[0078] Next, the channel attention weights can be obtained. To convert channel statistics into weighted coefficients, this embodiment of the application uses the channel descriptor vector... A linear transformation is performed, and channel attention weights in the range [0, 1] are generated using the SiLU activation function. ,in The SiLU activation function is used. The attention weights for this channel are... Used to characterize the importance of each channel to the contour representation.

[0079] Then the channel attention weights mentioned above can be... With feature output from the Multilayer Perceptron (MLP) module Channel-by-channel weighting is performed to obtain the output characteristics of the CPCAM module. .

[0080] Finally, to map the weighted features to a unified feature space, a linear projection layer can be used to modify the output features of the CPCAM module. The transformation is then performed to obtain the final output of the module. Through the above steps, CPCAM can effectively capture the semantic importance of contour priors in the channel dimension, achieve channel enhancement based on contour information, and thus further improve the model's ability to preserve structural contours.

[0081] The specific structure of the CPCAM module is as follows: Figure 3 As shown.

[0082] In some embodiments of this application, the time step in DDPM The predicted noise is obtained through the formula Determined; among them, For time step Predictive noise, Indicates CPGMD, For time step The image after adding noise, For degraded images, For contour prior information, This represents the contour prior encoding operator.

[0083] ;in, For time step Add to clean image Noise in , For time step The percentage of the original image signal retained. This represents a predefined noise scheduling scheme, i.e., a predefined time step. The noise added. In one example, It can be gradually increased from 0.0001 to 0.02; It is a multiplication operator.

[0084] The process of sequentially acquiring spatially enhanced interaction features, channel-enhanced interaction features, and determining prediction noise is as follows: Figure 1 The module on the right side of the dotted frame in the lower middle is shown.

[0085] In some embodiments of this application, recovering a clean image from a degraded image based at least on predicted noise may include: sampling Gaussian noise from a Gaussian distribution. As the initial value for progressive denoising; in response to determining Based on the formula Determine each time step Noisy image ;in, For time step The percentage of the original image signal retained. , Follows a Gaussian distribution. A positive integer greater than 2; in response to determination Based on the formula Determine a clean image .

[0086] The technical solution provided in this application extracts contour prior information from the self-degrading image using CPGMD and encodes it to obtain contour prior features. Feature interaction is performed between the contour prior features and the input features to obtain spatially enhanced interactive features. Then, channel attention enhancement is applied to the spatially enhanced interactive features and the contour prior features to obtain channel-enhanced interactive features. Next, the channel-enhanced interactive features are used to generate prediction noise for each time step in DDPM. Finally, DDPM is used to recover a clean image from the degraded image based on this prediction noise. This image restoration model has high robustness and generalization ability.

[0087] To verify the image restoration effect of the model provided in this application embodiment, the following experiment was designed: The model provided in this application embodiment was tested on multiple datasets, including the underwater image enhancement dataset LSUI, the low-light image enhancement dataset LOLv2_syn, and the image denoising dataset BSD68. The comparison results are as follows: Figures 4 to 6 As shown.

[0088] refer to Figure 4The first column is the input image. Columns 2-8 are the comparison methods: FUnIE (Fully Convolutional Conditional Generative Adversarial Network Real-Time Underwater Image Enhancement Model), WaterNet (Underwater Image Enhancement Network), Ucolor (Medium Transmission Guided Multi-Color Space Embedding Method), Ushape (U-Shaped Transformer Underwater Image Enhancement Method), WWPE (Weighted Wavelet Visual Perception Fusion Underwater Image Enhancement), Diffwater (Conditional Denoising Diffusion Model Underwater Image Enhancement Method), and DM-water (Transformer Diffusion Model Underwater Image Enhancement Method). Column 9 is the method of this application (CPGMD), and column 10 is the real image. In the first row, the target object appears brown instead of white. Similar color distortion also appears in the turtle (second row), fish (third row), and aquatic plants (fourth row). Similar quality degradation also exists in rows four through seven.

[0089] refer to Figure 5 The first column is the input image. Columns 2-7 represent the comparison methods: KinD (Light Up Dark Network), MIRNet (Multi-Scale Information Residual Image Enhancement), Restormer (Transformer Image Restoration Network), SNR (Signal-to-Noise Ratio Aware Low-Light Enhancement), Retinexformer (Retinex Transformer Low-Light Image Enhancement), and CIDNet (Color Space Low-Light Image Enhancement). Column 8 is the method proposed in this application (CPGMD), and column 9 is the real image. The comparison methods typically suffer from color distortion, insufficient or excessive enhancement, and loss of naturalness, especially in sky and mountain areas. In contrast, CPGMD maintains natural brightness and color fidelity across various scenes.

[0090] refer to Figure 6 The first line is the input noisy image; lines 2-4 compare the methods: FDnCNN (Flexible Feedforward Denoising Convolutional Neural Network), FFDNet (Fast Flexible Image Denoising Network), and IRCNN (Deep CNN Denoising Prior Image Restoration); line 5 is the method of this application (CPGMD); and line 6 is the real image. Although all methods have similar noise variance... Both 15 and 25 can reduce noise to some extent, but their performance degrades significantly when the noise variance is 50. In contrast, even under severe noise conditions, CPGMD can robustly generate denoised images that are visually closest to the real image (GT).

[0091] The general image restoration model based on contour prior-guided Mamba diffusion provided in this application utilizes an image segmentation model to extract contour priors that are independent of image degradation, providing clear and reliable guidance to achieve robust restoration for various degrees of degradation. The designed CPSAM module uses the contour prior-guided model to focus on spatially salient regions, improving the restoration of local details. The designed CPCAM module uses the contour prior to generate channel-level guidance, enhancing the semantic consistency and feature representation of the restored image.

[0092] The CPGMD provided in this application achieves advanced performance and demonstrates strong robustness and generalization ability.

[0093] Figure 7 This is a schematic flowchart of a general image restoration method based on contour-prior-guided Mamba diffusion provided in an embodiment of this application. Figure 7 As shown, the method includes the following steps:

[0094] In step S701, the degraded image is acquired.

[0095] In step S702, the predicted noise for each time step in the denoising diffusion probability model is generated using a general image restoration framework based on contour prior-guided Mamba diffusion.

[0096] In step S703, a clean image is recovered from the degraded image using a denoising diffusion probability model based at least on the predicted noise.

[0097] CPGMD generates the prediction noise for each time step in DDPM in the following manner:

[0098] Contour prior information is extracted from self-degenerate images using the contour prior extraction module;

[0099] The extracted contour prior information is encoded using the contour prior encoding module to obtain contour prior features; the contour prior features are high-dimensional features and their distribution density is greater than a preset density threshold.

[0100] The contour prior spatial perception module is used to perform feature interaction between the contour prior features and the input features to obtain spatially enhanced interactive features; among them, the input features include the input features of the degraded image and the input features of the image after adding noise at the current time step, and the image after adding noise at the current time step is determined by the DDPM model;

[0101] The contour prior channel perception module is used to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain the channel-enhanced interactive features.

[0102] The noise prediction module is used to generate predicted noise for each time step in DDPM based at least on the channel-enhanced interaction features.

[0103] In other words, the CPGMD method provided in this application extracts contour information, which is independent of the specific degradation type, from the degraded image as prior information, and integrates it into the diffusion noise prediction process using a specially designed spatial and channel-aware module. This prior information provides a stable guiding signal, which helps to maintain the consistency and fine details of the content.

[0104] Specifically, CPGMD leverages the powerful zero-shot segmentation capabilities of SAM to extract stable contour-based prior information from various degraded inputs without retraining or using labeled data. To effectively integrate this prior information into the diffusion framework, this embodiment designs two complementary perceptual modules: a spatial perceptual module that guides the model to focus on contour-salient regions; and a channel perceptual module that adaptively enhances contour-related feature channels. These two modules work together to enable CPGMD to utilize both the spatial and semantic aspects of contour information, thereby improving the fidelity of image restoration in various degraded scenarios.

[0105] The technical solution provided in this application extracts contour prior information from the CPGMD self-degrading image and encodes it into contour prior features. Feature interaction is performed between the contour prior features and the input features to obtain spatially enhanced interactive features. Channel attention enhancement is then applied to the spatially enhanced interactive features and the contour prior features to obtain channel-enhanced interactive features. Next, the channel-enhanced interactive features are used to generate prediction noise for each time step in DDPM. Finally, DDPM is used to recover a clean image from the degraded image based on this prediction noise. This image restoration method exhibits high robustness and generalization ability.

[0106] In some implementations, extracting contour prior information from a degraded image includes: segmenting the degraded image using an image segmentation model to obtain a segmented image of the degraded image; converting the segmented image into a grayscale image and determining the horizontal and vertical gradient components of the grayscale image; calculating the edge intensity value of each pixel in the degraded image based on the horizontal and vertical gradient components; identifying pixels with edge intensity values ​​higher than a preset intensity threshold and pixels connected to strong edge pixels as edge pixels; and determining the binary contour map composed of each edge pixel as contour prior information.

[0107] In some implementations, feature encoding is performed on the extracted contour prior information to obtain contour prior features, including: constructing an N-layer convolutional structure; N is a positive integer greater than or equal to 3 and less than or equal to 6; and using the N-layer convolutional structure to perform feature encoding on the contour prior information to obtain contour prior features.

[0108] In some implementations, feature interaction is performed on the contour prior features and input features to obtain spatially enhanced interactive features, including: mapping the input features using the MambaScan module to obtain spatial features; performing linear projection on the spatial features to obtain a query vector; performing average pooling compression on the contour prior features to obtain pooled compressed features; generating key vectors and value vectors from the pooled compressed features through linear mapping; and using a scaled dot product attention mechanism to perform feature interaction on the query vector, key vector, and value vector to obtain spatially enhanced interactive features.

[0109] In some implementations, channel attention enhancement is applied to the spatially enhanced interaction features and contour prior features to obtain channel-enhanced interaction features. This includes: inputting the spatially enhanced interaction features and the current time step into a multilayer perceptron to obtain perceptual features; performing global average pooling on the contour prior features in the spatial dimension to obtain channel descriptor vectors; performing a linear transformation on the channel descriptor vectors and using the SiLU activation function to generate channel attention weights based on the linear transformation result; the value range of the channel attention weights is [0,1]; and performing channel-wise weighting on the perceptual features and channel attention weights to obtain channel-enhanced interaction features.

[0110] In some implementations, time steps in DDPM The predicted noise is obtained through the formula Determined; among them, For time step Predictive noise, Indicates CPGMD, For time step The image after adding noise, For degraded images, For contour prior information, Represents the contour prior encoding operator; ;in, For time step Add to clean image Noise in , For time step The percentage of the original image signal retained. This represents a predefined noise scheduling scheme. It is a multiplication operator.

[0111] In some implementations, recovering a clean image from a degraded image based at least on predicted noise includes: sampling Gaussian noise from a Gaussian distribution. As the initial value for progressive denoising; in response to determining Based on the formula Determine each time step Noisy image ;in, For time step The percentage of the original image signal retained. , , Follows a Gaussian distribution. A positive integer greater than 2; in response to determination Based on the formula Determine a clean image .

[0112] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0113] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0114] Figure 8 This is a schematic diagram of the electronic device provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 8 of this embodiment includes a processor 801, a memory 802, and a computer program 803 stored in the memory 802 and executable on the processor 801. When the processor 801 executes the computer program 803, it implements the steps in the various method embodiments described above. Alternatively, when the processor 801 executes the computer program 803, it implements the functions of each module / unit in the various device embodiments described above.

[0115] Electronic device 8 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 8 may include, but is not limited to, processor 801 and memory 802. Those skilled in the art will understand that... Figure 8 This is merely an example of electronic device 8 and does not constitute a limitation on electronic device 8. It may include more or fewer components than shown, or different components.

[0116] The processor 801 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0117] The memory 802 can be an internal storage unit of the electronic device 8, such as a hard disk or RAM of the electronic device 8. The memory 802 can also be an external storage device of the electronic device 8, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 8. The memory 802 can also include both internal and external storage units of the electronic device 8. The memory 802 is used to store computer programs and other programs and data required by the electronic device.

[0118] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0119] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0120] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A universal image restoration device based on contour-prior-guided Mamba diffusion, characterized in that, It includes the general image restoration module CPGMD based on contour prior-guided Mamba diffusion and the denoising diffusion probability module DDPM; CPGMD includes: The contour prior extraction module is configured to extract contour prior information from self-degenerate images; The contour prior encoding module is configured to perform feature encoding on the extracted contour prior information to obtain contour prior features; the contour prior features are high-dimensional features and the distribution density is greater than a preset density threshold. The contour prior spatial perception module is configured to perform feature interaction on the contour prior features and input features to obtain spatially enhanced interactive features; wherein, the input features include the input features of the degraded image and the input features of the image after adding noise at the current time step, and the image after adding noise at the current time step is determined by DDPM; The contour prior channel perception module is configured to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain channel-enhanced interactive features. The noise prediction module is configured to generate predicted noise for each time step in the DDPM based at least on the channel-enhanced interaction features. DDPM is configured to recover a clean image from a degraded image based at least on the predicted noise; The interaction between the prior contour features and the input features yields spatially enhanced interactive features, including: The MambaScan module is used to map the input features to obtain spatial features; The spatial features are linearly projected to obtain the query vector; The prior contour features are compressed by average pooling to obtain the compressed pooled features. The pooled compressed features are used to generate key vectors and value vectors respectively through linear mapping; The query vector, the key vector, and the value vector are used to perform feature interaction using a scaled dot product attention mechanism to obtain the spatially enhanced interaction features. Channel attention enhancement is applied to the spatially enhanced interaction features and the contour prior features to obtain channel-enhanced interaction features, including: The spatially enhanced interaction features and the current time step are input into the multilayer perceptron to obtain the perception features; Global average pooling is performed on the contour prior features in the spatial dimension to obtain the channel descriptor vector; The channel descriptor vector is linearly transformed, and the channel attention weights are generated based on the linear transformation result using the SiLU activation function; the value range of the channel attention weights is [0,1]. The perceptual features and the channel attention weights are weighted channel by channel to obtain the channel-enhanced interaction features; Time step in DDPM The predicted noise is obtained through the formula Determined; among them, For time step Predictive noise, Indicates CPGMD, For time step The image after adding noise, For degraded images, For contour prior information, Represents the contour prior encoding operator; ;in, For time step Add to clean image Noise in , For time step The percentage of the original image signal retained. This represents a predefined noise scheduling scheme. It is a multiplication operator; Recovering a clean image from a degraded image based at least on the predicted noise includes: Sample a Gaussian noise from a Gaussian distribution. As the initial value for progressive denoising; Response to determination Based on the formula Determine each time step Noisy image ;in, For time step The percentage of the original image signal retained. , Follows a Gaussian distribution. It is a positive integer greater than 2; Response to determination Based on the formula Determine a clean image .

2. The universal image restoration device based on contour prior-guided Mamba diffusion according to claim 1, characterized in that, Extracting contour prior information from self-degrading images, including: The degraded image is segmented using an image segmentation model to obtain the segmented image of the degraded image; The segmented image is converted into a grayscale image, and the horizontal and vertical gradient components of the grayscale image are determined. The edge intensity value of each pixel in the degraded image is calculated based on the horizontal gradient component and the vertical gradient component. Pixels with edge intensity values ​​higher than a preset intensity threshold and pixels connected to strong edge pixels are identified as edge pixels; The binarized contour map composed of each edge pixel is determined as the contour prior information.

3. The universal image restoration device based on contour prior-guided Mamba diffusion according to claim 1, characterized in that, The extracted contour prior information is feature-encoded to obtain contour prior features, including: Construct an N-layer convolutional structure; N is a positive integer greater than or equal to 3 and less than or equal to 6; The contour prior information is feature-encoded using the N-layer convolutional structure to obtain the contour prior features.

4. A general image restoration method based on contour prior-guided Mamba diffusion, characterized in that, include: Obtain the degraded image; The general image restoration module CPGMD, guided by contour priors and Mamba diffusion, is used to generate the predicted noise at each time step in the denoising diffusion probability module DDPM. A clean image can be recovered from the degraded image using DDPM based at least on the predicted noise; The CPGMD generates the prediction noise for each time step in DDPM in the following manner: Contour prior information is extracted from self-degenerate images using the contour prior extraction module; The extracted contour prior information is feature-encoded using a contour prior coding module to obtain contour prior features; the contour prior features are high-dimensional features and their distribution density is greater than a preset density threshold. The contour prior spatial perception module is used to perform feature interaction between the contour prior features and the input features to obtain spatially enhanced interactive features; wherein, the input features include the input features of the degraded image and the input features of the image after adding noise at the current time step, and the image after adding noise at the current time step is determined by DDPM; The contour prior channel perception module is used to perform channel attention enhancement on the spatially enhanced interactive features and the contour prior features to obtain the channel-enhanced interactive features. The noise prediction module generates predicted noise for each time step in DDPM based at least on the channel-enhanced interaction features. The interaction between the prior contour features and the input features yields spatially enhanced interactive features, including: The MambaScan module is used to map the input features to obtain spatial features; The spatial features are linearly projected to obtain the query vector; The prior contour features are compressed by average pooling to obtain the compressed pooled features. The pooled compressed features are used to generate key vectors and value vectors respectively through linear mapping; The query vector, the key vector, and the value vector are used to perform feature interaction using a scaled dot product attention mechanism to obtain the spatially enhanced interaction features. Channel attention enhancement is applied to the spatially enhanced interaction features and the contour prior features to obtain channel-enhanced interaction features, including: The spatially enhanced interaction features and the current time step are input into the multilayer perceptron to obtain the perception features; Global average pooling is performed on the contour prior features in the spatial dimension to obtain the channel descriptor vector; The channel descriptor vector is linearly transformed, and the channel attention weights are generated based on the linear transformation result using the SiLU activation function; the value range of the channel attention weights is [0,1]. The perceptual features and the channel attention weights are weighted channel by channel to obtain the channel-enhanced interaction features; Time step in DDPM The predicted noise is obtained through the formula Determined; among them, For time step Predictive noise, Indicates CPGMD, For time step The image after adding noise, For degraded images, For contour prior information, Represents the contour prior encoding operator; ;in, For time step Add to clean image Noise in , For time step The percentage of the original image signal retained. This represents a predefined noise scheduling scheme. It is a multiplication operator; Recovering a clean image from a degraded image based at least on the predicted noise includes: Sample a Gaussian noise from a Gaussian distribution. As the initial value for progressive denoising; Response to determination Based on the formula Determine each time step Noisy image ;in, For time step The percentage of the original image signal retained. , Follows a Gaussian distribution. It is a positive integer greater than 2; Response to determination Based on the formula Determine a clean image .

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the general image restoration method based on contour prior-guided Mamba diffusion as described in claim 4.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the general image restoration method based on contour prior-guided Mamba diffusion as described in claim 4.

Citation Information

Patent Citations

  • Super-division method and device for using priori knowledge in diffusion model without training, and storage medium

    CN118570064A

  • Multi-feature enhancement fusion Mama image denoising method

    CN120707868A