Method for non-paired image defogging by learning implicit neurodegeneration characterization

By learning the NeDR-Dehaze method of implicit neural degenerate representation, combined with KAN-CID and implicit dense residual module, the balance problem between feature representation and global consistency modeling in complex haze scenes in existing technologies is solved, and a more efficient image dehazing effect is achieved.

CN120725899APending Publication Date: 2025-09-30CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818240.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing image dehazing methods have difficulty striking a balance between fine feature representation of uneven haze distribution and global consistency modeling when dealing with complex scenes, resulting in poor dehazing effects and insufficient generalization ability.

Method used

The NeDR-Dehaze method, which learns implicit neural degradation representation, enhances feature representation by combining channel-independent and channel-dependent mechanisms through the KAN-CID module, and optimizes haze degradation features using the implicit dense residual module and dense residual enhancement module to restore image details.

Benefits of technology

It significantly improves the performance and robustness of image dehazing, can more effectively restore clear and detail-rich images, adapt to complex haze scenes, reduce high-frequency redundant information, and enhance the ability to restore details in structural areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725899A_ABST
    Figure CN120725899A_ABST
Patent Text Reader

Abstract

The invention discloses a non-paired image defogging method for learning implicit neurodegeneration representation, and relates to the technical field of computer vision. In order to more effectively process spatial feature changes and enhance fine-grained feature representation and global consistency modeling capability of a model, the invention provides a channel-independent and channel-related mechanism (KAN-CID), the architecture is based on a Kolmogorov-Arnold network (KAN) framework, channel-independent and channel-dependent mechanisms are integrated, and in particular, the framework has the advantages that the framework is simplified, the framework is simplified, and the framework is simplified. The KAN is based on the Commogov-Arnod representation theorem, a high-dimensional function can be effectively decomposed into a plurality of one-dimensional functions, meanwhile, the shape of an activation function is adaptively adjusted, the expressive force of the network is greatly improved, the KAN-CID module provided by the invention gives consideration to local sensing and global modeling, the structure details of the image are reserved, and the image quality is improved. And the context representation between the channels is enhanced, so that more accurate feature modeling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for learning implicit neural degenerate representations for unpaired image defogging. Background Art

[0002] Image dehazing is an important task in the field of computer vision, which aims to restore clear and detailed visual content from images affected by haze.

[0003] This technology is crucial for improving the performance of downstream applications such as object detection, for example, PED-YOLO, RT-DETR, and autonomous driving (VLP). The formation of haze images is usually modeled using the atmospheric scattering model (ASM):

[0004] I(x)=J(x)t(x)+A(1-t(x))

[0005] Where I(x) is the foggy image, J(x) is the clear image, t(x) is the transmittance, and A is the global atmospheric light. It is usually defined as t(x) = e -βd(x) β represents the atmospheric scattering coefficient, and d(x) is the scene depth;

[0006] Existing image dehazing methods often include the following:

[0007] Image dehazing based on prior

[0008] Traditional image dehazing methods primarily rely on prior knowledge to estimate haze thickness or transmittance maps, thereby restoring a clear image. A classic example is the dark channel prior (DCP), which estimates the transmittance map based on the statistical observation that local regions in haze-free images typically contain low-intensity pixels. This approach has achieved remarkable dehazing performance. Subsequently, methods such as color attenuation prior (CAP) and non-local dehazing have been proposed. These methods construct priors based on image brightness, saturation, or color distribution, thereby improving the adaptability and robustness of the algorithms. While prior-based methods are physically interpretable, their heavy reliance on predefined assumptions makes them prone to artifacts and color distortion when the input image deviates from these assumptions, limiting their effectiveness, especially in extreme weather conditions or when haze distribution is highly heterogeneous.

[0009] Learning-based image dehazing

[0010] In recent years, deep learning has achieved breakthroughs in image restoration, making learning-based dehazing methods increasingly mainstream. Early methods, such as DehazeNet and AOD-Net, leveraged convolutional neural networks to automatically learn the mapping from input foggy images to transmission maps or clear images, significantly improving dehazing performance. These methods typically require a large number of synthetic grayscale images paired with corresponding clear images as supervision.

[0011] To overcome the reliance on synthetic paired data, subsequent research has gradually shifted to unsupervised or weakly supervised learning strategies. For example, Cycle-Dehaze employs the CycleGAN framework to achieve image-to-image dehazing without the need for paired training data. Furthermore, methods such as FFA-Net and GridDehazeNet further enhance the model's generalization capabilities by introducing attention mechanisms and multi-scale feature fusion modules.

[0012] Furthermore, to improve the practicality of dehazing models in real-world scenarios, recent research has increasingly focused on evaluating them using real-world hazy image datasets. SGDN addresses the limitations of RGB-based representations by incorporating YCbCr structural information as guidance, leveraging dual-domain complementarity to better recover subtle edge and color relationships in real images. Furthermore, the RW2AH dataset, with its rich geographical and climatic diversity, was introduced, providing a more reliable foundation for training and evaluating dehazing models in real-world applications. ODCR proposed an orthogonal separation contrast regularization method, which achieves more effective dehazing without paired data by orthogonally separating content and haze information and applying a contrast constraint. Lan et al. proposed Diff-Dehazer, a diffusion-based unpaired image dehazing framework. This framework combines physical priors with the generative power of diffusion models to achieve superior dehazing performance on real-world images. Existing methods often over-rely on explicit feature extraction and physical models for complex scenes, struggling to strike a balance between capturing fine-grained features of non-uniform haze distributions and modeling global consistency, thus limiting their dehazing effectiveness and generalization in real-world scenarios.

[0013] Implicit neural representation

[0014] Implicit neural representations have also been widely used in other fields. Chen et al. integrated INR into a multi-scale Transformer architecture in the image deraining task to better explore multi-scale information and model complex rain patterns, thereby improving the robustness of the model in complex scenes. Nam et al. proposed a framework using neural image representations (NIRs) to effectively merge multiple inputs into a single standard view without selecting one of the images as a reference frame. Yang et al. proposed a collaborative low-light image enhancement method NeRCo based on implicit neural representation. It robustly restores perceptually friendly results in an unsupervised manner. However, local redundant features are encountered when processing images, which affects the modeling of high-frequency details and the ability to recover details in structural areas. Inspired by the above fields, the work of the present invention further optimizes the implicit representation to better promote the removal of haze.

[0015] In summary, haze is extremely unevenly distributed in images and has complex multi-scale characteristics, such as different sizes, shapes, angles, depths, and densities. These characteristics pose severe challenges to convolutional neural networks (CNNs). Due to the fixed receptive field, existing CNN architectures have difficulty effectively capturing spatially varying features and non-local structural information, especially in images with uneven haze distribution. This limitation hinders their ability to effectively represent fine-grained features and achieve consistent global modeling, thereby affecting overall performance. Although haze images often exhibit similar visual degradation patterns (such as typical haze-induced distortion), current methods rely heavily on traditional feature representations, which are sensitive to input variations and cannot fully model the implicit latent functions behind these common degradations. This limits their performance in complex scenes.

[0016] When dealing with complex scenes, existing methods often have difficulty striking a balance between fine feature representation of uneven haze distribution and global consistency modeling;

[0017] Therefore, learning the potential correlations between features from spatially varying haze is of great research significance for restoring clearer images, and a new solution to the above problem is needed. Summary of the Invention

[0018] The purpose of the present invention is to provide a method for learning implicit neural degradation representation for unpaired image dehazing, namely, an unsupervised dehazing method named NeDR-Dehaze. This method focuses on implicit neural degradation representation and aims to deeply explore the nonlinear degradation characteristics of haze and its implicit representation, so as to more efficiently restore clear and detailed image content to solve the technical problems raised in the background technology.

[0019] To achieve the above objectives, the present invention provides the following technical solution: a method for learning implicit neural degenerate representation for unpaired image dehazing, comprising at least the following steps:

[0020] S1: The dehazing network architecture includes at least a KAN-CID module, a multi-level feature fusion module, an implicit dense residual module and a dense residual enhancement module. The implicit dense residual module is an IDRM module, and the dense residual enhancement module is a DREM module.

[0021] S2: After multi-scale feature extraction, the feature representation is further enhanced through the KAN-CID module, and the DREM module is used to enhance the expressive power of IDRM in detailed modeling;

[0022] S3: gradually restore image details under the multi-level feature fusion module;

[0023] S4: An implicit dense residual module is embedded in the image reconstruction process, and the haze degradation features are modeled as a continuous function using implicit neural representation to enhance the robustness and adaptability of the model. In addition, redundant information in high-frequency features is removed to enhance the recovery of details in structural areas.

[0024] Preferably, the KAN-CID module integrates channel-independent and channel-dependent mechanisms. The KAN-CID module jointly models the spatial and channel dimensions. The channel-independent branch extracts spatial details through spatial convolution, while the channel-independent branch uses a learnable kernel function to capture inter-channel dependencies, thereby achieving seamless fusion of local features and global context information.

[0025] The KAN-CID module consists of two branches: a channel-independent branch and a channel-dependent branch. The channel-independent branch is a CI branch that is independent of the channel, and the channel-dependent branch is a CD branch that is related to the channel.

[0026] The application process of the KAN-CID module is as follows:

[0027] First, each channel is processed independently to deeply extract the information within the channel, and then the correlation between channels is modeled and integrated through cross-channel information fusion.

[0028] Preferably, each channel of the channel-independent branch independently performs depth-separable convolution to extract spatial structural information, and by avoiding inter-channel interference, can more accurately preserve local details such as edges and textures. Given an input feature map F∈R C×H×W , the output of the channel independent branch is: F CI =DWConv 7×7 (F);

[0029] The output of the CI branch is first flattened and then fed into the CD branch.

[0030] Preferably, the KAN-CID module is an innovative network architecture based on KAN, which uses Kolmogorov-Arnold Networks with kernels of different sizes, so that it can effectively model complex and multi-level semantic dependencies across channels, which is conducive to nonlinear cross-channel modeling. The KAN process includes at least the following steps:

[0031]

[0032] Where N=3 is a hyperparameter, I represents the input feature vector, Φ i Represents the i-th KAN-Layer layer, each i-th KAN-Layer layer Φ i , with n-dimensional input and n-dimensional output, expressed as follows:

[0033] Φ={φ q,p},p=1,2,...,n in ,q=1,2,...,n out

[0034] The Kolmogorov-Arnold representation theorem in matrix form is expressed as:

[0035]

[0036] Φ out =[Φ1(·),…,Φ 2n (·),Φ 2n+1 (·)]

[0037] where Φ includes the nin×nout learnable activation function φ; q,p Represents the learnable parameters, and the calculation results are expressed in the form of a matrix:

[0038]

[0039] Among them, Φ i (I i ) represents the feature I input to the KAN-Layer layer i Output;

[0040] The KAN-CID module enhances the representation and global consistency of fine-grained features by combining channel-independent and channel-dependent mechanisms. Specifically, the channel-independent branch extracts spatial details through spatial convolution, while the channel-dependent branch uses a learnable kernel function to capture inter-channel dependencies. This achieves a seamless fusion of local features and global contextual information, significantly improving the network's dehazing performance and image restoration quality.

[0041] Preferably, the implicit dense residual module optimizes the implicit neural representation of the haze image in the form of dense residual enhancement, aiming to effectively restore image quality by learning the implicit representation of haze degradation characteristics. By utilizing the parameterization function of the neural network, the image content is represented as a continuous function, thereby flexibly capturing the intricate structure and changes in the image. In addition, by performing a differential operation on the original input image, it can effectively highlight the key image content obscured by haze.

[0042] The application process of the implicit dense residual module includes at least the following steps:

[0043] Given an input foggy image I hazy ∈R H×W×3 , where H×W represents the spatial resolution of the image, and the foggy image is converted into a feature map E'∈R H×W×C ;

[0044] At the same time, the corresponding two-dimensional coordinates are recorded in X'∈R H×W×2 middle;

[0045] Then, E' and X' are fused and the fused information is decoded through a multi-layer perceptron to generate the final neural representation;

[0046] I IDRM [i,j]=F MLP (E'[i,j],X'[i,j])

[0047] Among them, I IDRM [i,j] represents the processed pixel value at position (i,j), E'[i,j] is the feature vector of that position, and X'[i,j] is the corresponding coordinate vector;

[0048] In order to eliminate redundant features and capture common haze degradation characteristics, they are integrated into the dense residual enhancement module and the multilayer perceptron is trained by learning the following mapping function:

[0049] f θ :R 2 →R 3 ,f(x,y)=(r,g,b)

[0050] Among them, f θ:R 2 →R 3 represents the mapping defined by the parameters θ, which projects the two-dimensional image coordinates into a three-dimensional color space, where (r, g, b) represents the values ​​of the red, green, and blue channels respectively;

[0051] In order to enhance the network's ability to capture spatial information, a spatial encoding mechanism is introduced. Specifically, the image coordinate X is mapped to a high-dimensional space through the encoding function γ, and then used as one of the inputs of the network. This process is expressed by the following formula:

[0052] X'=γ(X)

[0053] γ(x)=[sin(2 0 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)]

[0054] Where X represents the image coordinate, X' represents the encoding coordinate, and L = 4 is a hyperparameter that determines the encoding dimension.

[0055] Preferably, the dense residual enhancement module is used to enhance the expressiveness of the IDRM module in terms of detailed modeling;

[0056] The IDRM structure can effectively learn the continuous distribution of the overall image structure, but it will inevitably encounter some local redundant features during the processing. To address this problem, DREM is embedded as an auxiliary module during the training process to focus on removing redundant information in high-frequency features and enhance the ability to recover details in structural areas.

[0057] The DREM module integrates residual learning and dense connection mechanism to construct a two-layer residual dense block (RDB) structure;

[0058] Through the initial convolutional layer and residual path, the feature response is strengthened, and then the feature hierarchy is gradually enriched through dense connection and local fusion. The final output features are improved through convolution and differential operation with the original input image, which effectively highlights the key image content under haze interference, helping IDRM to fit the implicit representation of the target image more accurately.

[0059] Compared with the prior art, the present invention has the following beneficial effects:

[0060] 1. To enhance fine-grained feature representation and global consistency modeling under non-uniform haze distribution, this paper proposes a novel KAN-CID module, which integrates channel-independent and channel-dependent mechanisms in the Kolmogorov-Arnold network architecture.

[0061] 2. The present invention designs an implicit dense residual module to learn haze degradation characteristics and capture spatial details through continuous function mapping, thereby achieving more effective implicit haze modeling and stronger degradation processing capabilities;

[0062] 3. We conducted extensive experiments on several publicly available synthetic and real-world haze datasets and achieved impressive performance, demonstrating the robustness and effectiveness of our approach. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0064] Figure 1 Schematic diagram of the method architecture of the present invention;

[0065] Figure 2 Schematic diagram of visual comparison between the input image and IDRM result of the present invention;

[0066] Figure 3 This is a visual comparison chart of the haze removal effect of the present invention on the SOTS-Outdoor dataset samples;

[0067] Figure 4 This is a visual comparison chart of the defogging of the present invention on the SOTS-Indoor dataset sample;

[0068] Figure 5 This is a visual comparison chart of the defogging of the present invention on the SOTS-Indoor dataset sample;

[0069] Figure 6 This is a visual comparison diagram of the defogging of the present invention on the HSTS dataset sample;

[0070] Figure 7 This is a visual comparison diagram of the dehazing performance of the present invention on the I-HAZE dataset sample;

[0071] Figure 8 This is a visual comparison chart of the dehazing method of the present invention on the PhoneHazy dataset sample;

[0072] Figure 9 This is a visual comparison chart of the dehazing method of the present invention on the NH-Haze dataset samples. DETAILED DESCRIPTION

[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0074] See also Figure 1 , Figure 1 The architecture of the method embodying the present invention consists of a dehazing network. (a) shows the channel-independent and channel-dependent mechanisms of the KAN network designed based on the present invention, while (b) shows the implicit neural representation. The feature extraction component represents the extracted unscaled feature map; A, β, t, and d are intermediate parameters of the atmospheric scattering model; the ChnMapper module converts feature maps with different channel counts into feature maps with a specific expanded channel count.

[0075] A method for learning implicit neural degradation representation for unpaired image dehazing includes at least the following steps:

[0076] S1: The dehazing network architecture includes at least a KAN-CID module, a multi-level feature fusion module, an implicit dense residual module, and a dense residual enhancement module. The implicit dense residual module is the IDRM module, and the dense residual enhancement module is the DREM module.

[0077] S2: After multi-scale feature extraction, the feature representation is further enhanced through the KAN-CID module, and the DREM module is used to enhance the expressive power of IDRM in detailed modeling;

[0078] S3: gradually restore image details under the multi-level feature fusion module;

[0079] S4: An implicit dense residual module is embedded in the image reconstruction process, and the haze degradation characteristics are modeled as a continuous function using implicit neural representation to enhance the robustness and adaptability of the model. In addition, redundant information in high-frequency features is removed to enhance the recovery of structural area details.

[0080] In recent years, the Kolmogorov-Arnold network, as an emerging neural network architecture, has attracted increasing attention from researchers. Inspired by the Kolmogorov-Arnold (KAN) representation theorem, KAN aims to decompose high-dimensional functions into multiple one-dimensional functions, thereby enhancing the network's expressive power. Unlike traditional multi-layer perceptrons (MLPs), KANs use learnable activation functions on edges (weights) rather than fixed activation functions on nodes (neurons). This flexible architecture offers potential advantages for image processing tasks. The theorem states that there exists a representation of the following form:

[0081]

[0082] where x=(x1,x2,...,x n ) is represented by a single variable continuous function h and a series of continuous two variable functions xi and gq, formed by composition, for any continuous function f(x) defined in n-dimensional real space;

[0083] This paper aims to enhance the dehazing network's ability to model complex haze distributions by improving fine-grained feature representation and global consistency. We propose a feature enhancement module, KAN-CID, that integrates both channel-independent and channel-dependent mechanisms. This module jointly models the spatial and channel dimensions: the channel-independent branch extracts spatial details through spatial convolution, while the channel-independent branch employs a learnable kernel function to capture inter-channel dependencies. This design enables seamless fusion of local features with global contextual information.

[0084] In real-world image dehazing tasks, the uneven distribution of haze makes it challenging for convolutional neural networks to capture global semantics by relying solely on local feature extraction, often resulting in blurred details and color distortion. The KAN-CID module addresses this issue through its dual-branch architecture, significantly improving the network's dehazing performance and image restoration quality.

[0085] The KAN-CID module integrates channel-independent and channel-dependent mechanisms. It jointly models the spatial and channel dimensions. The channel-independent branch extracts spatial details through spatial convolution, while the channel-independent branch uses a learnable kernel function to capture inter-channel dependencies, achieving seamless fusion of local features and global contextual information.

[0086] like Figure 1 As shown in (a), the KAN-CID module consists of two branches: the channel-independent branch and the channel-dependent branch. The channel-independent branch is the CI branch that is independent of the channel, and the channel-dependent branch is the CD branch that is related to the channel.

[0087] The application process of the KAN-CID module is as follows:

[0088] First, each channel is processed independently to deeply extract the information within the channel, and then the correlation between channels is modeled and integrated through cross-channel information fusion.

[0089] Each channel of the channel-independent branch performs depth-separated convolution independently to extract spatial structure information. By avoiding inter-channel interference, local details such as edges and textures can be more accurately preserved. Given an input feature map F∈R C ×H×W , the output of the channel independent branch is: F CI =DWConv 7×7 (F);

[0090] The output of the CI branch is first flattened and then fed into the CD branch.

[0091] The KAN-CID module is an innovative network architecture based on KAN. It uses Kolmogorov-Arnold Networks with kernels of different sizes, which can effectively model complex and multi-level semantic dependencies across channels and facilitate nonlinear cross-channel modeling. The KAN process includes at least the following steps:

[0092]

[0093] Where N=3 is a hyperparameter, I represents the input feature vector, Φ i Represents the i-th KAN-Layer layer, each i-th KAN-Layer layer Φ i , with n-dimensional input and n-dimensional output, expressed as follows:

[0094] Φ={φ q,p},p=1,2,...,n in ,q=1,2,...,n out

[0095] The Kolmogorov-Arnold representation theorem in matrix form is expressed as:

[0096]

[0097] Φ out =[Φ1(·),…,Φ 2n (·),Φ 2n+1 (·)]

[0098] where Φ includes the nin×nout learnable activation function φ; q,p Represents the learnable parameters, and the calculation results are expressed in the form of a matrix:

[0099]

[0100] Among them, Φ i (I i ) represents the feature I input to the KAN-Layer layer i Output;

[0101] The KAN-CID module enhances the representation and global consistency of fine-grained features by combining channel-independent and channel-dependent mechanisms. Specifically, the channel-independent branch extracts spatial details through spatial convolution, while the channel-dependent branch uses a learnable kernel function to capture inter-channel dependencies. This achieves a seamless fusion of local features and global contextual information, significantly improving the network's dehazing performance and image restoration quality.

[0102] The implicit dense residual module optimizes the implicit neural representation of haze images in the form of dense residual enhancement. It aims to effectively restore image quality by learning implicit representations of haze degradation characteristics. It leverages the parameterization capabilities of neural networks to represent image content as a continuous function, thereby flexibly capturing the intricate structure and changes in the image. In addition, by performing a differential operation on the original input image, it can effectively highlight key image content obscured by haze.

[0103] like Figure 1 As shown in (b), the application process of the implicit dense residual module includes at least the following steps:

[0104] Given an input foggy image I hazy ∈R H×W×3 , where H×W represents the spatial resolution of the image, and the foggy image is converted into a feature map E'∈R H×W×C ;

[0105] At the same time, the corresponding two-dimensional coordinates are recorded in X'∈R H×W×2 middle;

[0106] Then, E' and X' are fused and the fused information is decoded through a multi-layer perceptron to generate the final neural representation;

[0107] I IDRM [i,j]=F MLP (E'[i,j],X'[i,j])

[0108] Among them, I IDRM [i,j] represents the processed pixel value at position (i,j), E'[i,j] is the feature vector of that position, and X'[i,j] is the corresponding coordinate vector;

[0109] In order to eliminate redundant features and capture common haze degradation characteristics, they are integrated into the dense residual enhancement module and the multilayer perceptron is trained by learning the following mapping function:

[0110] f θ :R 2 →R 3 ,f(x,y)=(r,g,b)

[0111] Among them, fθ :R 2 →R 3 represents the mapping defined by the parameters θ, which projects the two-dimensional image coordinates into a three-dimensional color space, where (r, g, b) represents the values ​​of the red, green, and blue channels respectively;

[0112] In order to enhance the network's ability to capture spatial information, a spatial encoding mechanism is introduced. Specifically, the image coordinate X is mapped to a high-dimensional space through the encoding function γ, and then used as one of the inputs of the network. This process is expressed by the following formula:

[0113] X'=γ(X)

[0114] γ(x)=[sin(2 0 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)]

[0115] Where X represents the image coordinates, X' represents the encoding coordinates, and L = 4 is a hyperparameter that determines the encoding dimension;

[0116] Redundant features in high-frequency features are effectively removed through the dense residual enhancement module. Compared with traditional explicit representation methods, IDRM has excellent flexibility in capturing complex structures and subtle changes in images, making it particularly suitable for dealing with diverse nonlinear degradations caused by haze. Specifically, the addition of local aggregation and feature expansion mechanisms greatly enhances the network's ability to model fine-grained local features and texture information. In addition, the use of position encoding further enriches the network's spatial perception capabilities, enabling it to have a deeper understanding of scene geometry. It is worth noting that IDRM adopts an end-to-end training strategy, allowing the model to learn the best feature representation directly from the original input image without relying on hand-crafted priors or physical models. This not only simplifies the dehazing pipeline, but also improves performance and efficiency. As Figure 2 As shown in the figure, it can be clearly seen that IDRM has a significant effect on reducing the high-intensity pixel values ​​caused by haze. This method can effectively reduce the interference of haze on the image and reconstruct the underlying haze-free image, thereby significantly improving the image clarity and detail expression. IDRM shows a significant advantage in image dehazing, effectively suppressing high-intensity white fog and restoring haze-free images.

[0117] The dense residual enhancement module is used to enhance the expressiveness of the IDRM module in detailed modeling;

[0118] The IDRM structure can effectively learn the continuous distribution of the overall image structure, but it will inevitably encounter some local redundant features during the processing. To address this problem, DREM is embedded as an auxiliary module during the training process, focusing on removing redundant information in high-frequency features and enhancing the ability to recover details in structural areas.

[0119] like Figure 1 As shown in (b), the DREM module integrates residual learning and dense connection mechanism to construct a two-layer residual dense block (RDB) structure;

[0120] Through the initial convolutional layer and residual path, the feature response is strengthened, and then the feature hierarchy is gradually enriched through dense connection and local fusion. The final output features are refined through convolution and differential operation with the original input image, effectively highlighting the key image content under the interference of haze, helping IDRM to more accurately fit the implicit representation of the target image;

[0121] This design improves the network's adaptability to complex haze distribution without relying on traditional explicit estimation. It also makes up for the structural expression ability of IDRM in modeling high-frequency details, providing stronger restoration capabilities for image dehazing tasks.

[0122] In order to evaluate the performance of the proposed method on dehazing networks, a series of qualitative and quantitative comparative experiments were conducted.

[0123] Experimental setup

[0124] Training and test data. To train the model, the present invention uses the RESIDE dataset. The indoor training set (ITS) of RESIDE contains 13,990 synthetic haze images and 1,399 clean images, while the outdoor training set (OTS) contains 17,500 synthetic haze images and 500 clean images. To be more fair, the present invention also trains the baseline on the outdoor training set (OTS). In addition, the present invention also evaluates the proposed method on synthetic datasets and real-world datasets. The test dataset includes a synthetic dataset containing 500 indoor images and 500 outdoor images, and a real-world dataset containing 35 images, such as the I-HAZE dataset, the PhoneHazy dataset, containing 40 images; the HSTS dataset, containing 10 images; the NH-Haze dataset, containing 5 images; and the O-HAZE dataset, containing 45 images.

[0125] Evaluation metrics. For quantitative comparison, we selected three metrics: PSNR, SSIM

[30] , and LPIPS. These metrics are commonly used to evaluate the effectiveness of dehazing networks. Higher PSNR and SSIM values ​​indicate better visual quality, while lower LPIPS values ​​correspond to better visual quality.

[0126] Experimental details. During the experiment, the present invention carefully set up the training process of the proposed network model. Specifically, the present invention used the Adam optimizer with hyperparameters = 0.9, = 0.999, and set the number of samples per batch to 4. In addition, the learning rate was set to 0.0001. The model was trained on an NVIDIA 3090 GPU using the PyTorch framework. Through experimental verification, the present invention found that after 1560000 iterations, the network model achieved optimal performance.

[0127] Performance Evaluation

[0128] To comprehensively evaluate the defogging model of the present invention, a comparative study with the most advanced defogging methods was conducted. This comparison included both qualitative and quantitative analysis.

[0129] Qualitative analysis. Figure 3-9 Qualitative evaluation results of different dehazing methods are presented on multiple datasets, including SOTS-Indoor, SOTS-Outdoor, HSTS, O-HAZE, NH-Haze, PhoneHazy, and I-HAZE. As can be seen, the proposed method achieves higher PSNR and SSIM scores in most scenarios. In particular, on complex real-world datasets with realistic fog degradation (such as PhoneHazy, O-HAZE, and I-HAZE), the proposed method demonstrates stronger generalization and detail restoration performance.

[0130] Further analysis Figure 3 From the visual results in Figure 2, it can be seen that the proposed method outperforms other competing methods in restoring image structure, texture details, and color naturalness. Deep learning-based methods (e.g., RIDCP, SGDN, and D4) often perform poorly in dehazing performance. In contrast, the proposed method can accurately restore the structure of distant objects, remove residual fog, and avoid over-enhancement while maintaining overall image clarity. Figure 4 As shown in Figure 3, USID-NET and RefineDNet methods often have significant deviations in color restoration, which makes the method of the present invention more advantageous in terms of visual consistency and practical applicability.

[0131] Quantitative analysis. In order to comprehensively evaluate the performance of the dehazing method proposed in the present invention, the present invention performed quantitative analysis on multiple publicly available datasets. For example, the present invention compared the method of the present invention with the most advanced dehazing methods using standard evaluation metrics such as peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learning-perceptual image patch similarity (LPIPS). The experimental results in Tables 1-2 show that the method of the present invention achieved excellent performance on all test datasets. Specifically, on the SOTS-Indoor dataset, the method of the present invention achieved a PSNR of 25.73dB and an SSIM of 0.937, which is better than the existing methods. On the SOTS-Outdoor dataset, the method of the present invention achieved a PSNR of 26.98dB and an SSIM of 0.963, which shows that the method of the present invention has a strong dehazing ability in outdoor scenes. The I-HAZE dataset, which consists of real fog and its corresponding fog-free indoor images, also achieved a PSNR of 16.47dB and an SSIM of 0.783, respectively, which are better than the existing methods. More experimental data are shown in detail in the corresponding tables. The results show that the method of the present invention has achieved satisfactory results in practical applications.

[0132] Table 1: Quantitative comparison with dehazing methods on SOTS-Outdoor and HSTS datasets

[0133]

[0134] Table 2: Quantitative comparison with dehazing methods on PhoneHazy and NH-Haze datasets

[0135]

[0136] As shown in Table 3, the proposed dehazing model, which does not require paired training data, demonstrates excellent efficiency, significantly reduces the number of parameters, and uses minimal computational effort. Although lightweight models such as SANet, FFANet, and DEANet also perform competitively in terms of parameter count and computational cost, they rely on supervised training and are inferior to the proposed method in terms of dehazing performance. Experimental results further confirm that the proposed method outperforms existing methods in both efficiency and dehazing quality, while also exhibiting greater stability.

[0137] Table 3 Comparison of the efficiency of different dehazing methods

[0138]

[0139] Ablation experiments

[0140] In order to evaluate the effectiveness of the method of the present invention in the image decontamination task, the present invention conducted ablation experiments on different modules on the SOTS-Outdoor test set. As shown in Table 4, specifically, Variant 2 (V2): KAN-CID greatly enhances the model's ability to learn the nonlinear degradation relationship between grayscale images and clean images, thereby making the restored image closer to the true distribution. In order to achieve an effective balance between degradation normalization and content fidelity, Variant 3 (V3) further shows that the strategy of the present invention focuses on implicit neural representation and irregular degradation features. This improves the expressive power and generalization performance of the enhanced model.

[0141] Table 4: Ablation experiments on indoor and outdoor datasets

[0142]

[0143] × and √ decibels indicate that the corresponding modules are not used and are used.

[0144] To investigate the impact of the number of training iterations on model performance, we conducted ablation studies with different numbers of iterations. As shown in Table 5, both PSNR and SSIM show an overall upward trend with increasing number of iterations, indicating that the reconstruction quality is gradually improving. The model achieves optimal performance when N = 156, with a PSNR of 26.98 and an SSIM of 0.963. Notably, a slight decrease in performance is observed when N = 160, which may be due to overfitting or training saturation. These findings highlight the importance of carefully selecting the number of training iterations to ensure a balance between convergence and generalization.

[0145] More ablation study results are shown in Table 3.

[0146] Table 5: Ablation study of different training iterations (in 10,000)

[0147] N=40 N=70 N=100 N=130 N=156 N=160 PSNR 25.63 26.39 25.90 26.47 26.98 26.27 SSIM 0.953 0.960 0.960 0.959 0.963 0.962

[0148] In summary:

[0149] In order to more effectively handle spatial feature changes and enhance the model's fine-grained feature representation and global consistency modeling capabilities, the present invention proposes a channel-independent and channel-dependent mechanism (KAN-CID). This architecture is based on the Kolmogorov-Arnold network (KAN) framework and integrates channel-independent and channel-dependent mechanisms. Specifically, KAN is based on the Kolmogorov-Arnold representation theorem and can effectively decompose high-dimensional functions into multiple one-dimensional functions while adaptively adjusting the shape of the activation function, which greatly improves the network's expressiveness. The KAN-CID module proposed in this invention takes into account both local perception and global modeling, retaining the structural details of the image while enhancing the contextual representation between channels, thereby achieving more accurate feature modeling.

[0150] Furthermore, to more effectively learn the underlying correlations between features from spatially varying haze and enhance the model's adaptive representation, the present invention designs an implicit neural representation within the feature fusion process. Building on this, the present invention further designs a dense residual enhancement module (DREM). This module not only optimizes the processing of feature information but also significantly eliminates redundant information, significantly enhancing the model's ability to capture haze degradation characteristics and spatially varying details. This improvement significantly enhances the model's robustness in complex situations.

[0151] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for learning implicit neural degenerate representation for unpaired image dehazing, characterized by: At least the following steps are included: S1: The dehazing network architecture includes at least a KAN-CID module, a multi-level feature fusion module, an implicit dense residual module and a dense residual enhancement module. The implicit dense residual module is an IDRM module, and the dense residual enhancement module is a DREM module. S2: After multi-scale feature extraction, the feature representation is further enhanced through the KAN-CID module, and the DREM module is used to enhance the expressive power of IDRM in detailed modeling; S3: gradually restore image details under the multi-level feature fusion module; S4: An implicit dense residual module is embedded in the image reconstruction process, and the haze degradation features are modeled as a continuous function using implicit neural representation to enhance the robustness and adaptability of the model. In addition, redundant information in high-frequency features is removed to enhance the recovery of details in structural areas.

2. The method for learning implicit neural degenerate representation for unpaired image dehazing according to claim 1, characterized in that: The KAN-CID module integrates channel-independent and channel-dependent mechanisms. The KAN-CID module jointly models the spatial and channel dimensions. The channel-independent branch extracts spatial details through spatial convolution, while the channel-independent branch uses a learnable kernel function to capture inter-channel dependencies, achieving seamless fusion of local features and global context information. The KAN-CID module consists of two branches: a channel-independent branch and a channel-dependent branch. The channel-independent branch is a CI branch that is independent of the channel, and the channel-dependent branch is a CD branch that is related to the channel. The application process of the KAN-CID module is as follows: First, each channel is processed independently to deeply extract the information within the channel, and then the correlation between channels is modeled and integrated through cross-channel information fusion.

3. The method for learning implicit neural degenerate representation for unpaired image dehazing according to claim 2, characterized in that: Each channel of the channel-independent branch independently performs depth-separable convolution to extract spatial structure information and more accurately retain local details by avoiding inter-channel interference. Given an input feature map F∈R C×H×W , the output of the channel independent branch is: F CI =DWConv 7×7 (F); The output of the CI branch is first flattened and then fed into the CD branch.

4. The method for learning implicit neural degenerate representation for unpaired image dehazing according to claim 2, characterized in that: The KAN-CID module is an innovative network architecture based on KAN. It adopts a Kolmogorov-Arnold network with kernels of different sizes, which can effectively model complex and multi-level semantic dependencies across channels and is conducive to nonlinear cross-channel modeling. The KAN process includes at least the following steps: Where N=3 is a hyperparameter, I represents the input feature vector, Φ i Represents the i-th KAN-Layer layer, each i-th KAN-Layer layer Φ i , with n-dimensional input and n-dimensional output, expressed as follows: Φ={φ q,p },p=1,2,...,n in ,q=1,2,...,n out The Kolmogorov-Arnold representation theorem in matrix form is expressed as: F out =[Φ1(·),…,Φ 2n (·),Φ 2n+1 (·)] where Φ includes the nin×nout learnable activation function φ; q,p Represents the learnable parameters, and the calculation results are expressed in the form of a matrix: Among them, Φ i (I i ) represents the feature I input to the KAN-Layer layer i Output; The KAN-CID module enhances the representation and global consistency of fine-grained features by combining channel-independent and channel-dependent mechanisms. Specifically, the channel-independent branch extracts spatial details through spatial convolution, while the channel-dependent branch uses a learnable kernel function to capture inter-channel dependencies. This achieves a seamless fusion of local features and global contextual information, thereby significantly improving the network's dehazing performance and image restoration quality.

5. The method for learning implicit neural degenerate representation for unpaired image dehazing according to claim 1, characterized in that: The implicit dense residual module optimizes the implicit neural representation of haze images in the form of dense residual enhancement. It aims to effectively restore image quality by learning implicit representations of haze degradation characteristics. It uses the parameterization capabilities of neural networks to represent image content as a continuous function, thereby flexibly capturing the intricate structure and changes in the image. In addition, by performing a differential operation on the original input image, it can effectively highlight key image content obscured by haze. The application process of the implicit dense residual module includes at least the following steps: Given an input foggy image I hazy ∈R H×W×3 , where H×W represents the spatial resolution of the image, and the foggy image is converted into a feature map E'∈R H×W×C ; At the same time, the corresponding two-dimensional coordinates are recorded in X'∈R H×W×2 middle; Then, E' and X' are fused and the fused information is decoded through a multi-layer perceptron to generate the final neural representation; I IDRM [i,j]=F MLP (E'[i,j],X'[i,j]) Among them, I IDRM [i,j] represents the processed pixel value at position (i,j), E'[i,j] is the feature vector of that position, and X'[i,j] is the corresponding coordinate vector; In order to eliminate redundant features and capture common haze degradation characteristics, they are integrated into the dense residual enhancement module and the multilayer perceptron is trained by learning the following mapping function: f θ :R 2 →R 3 ,f(x,y)=(r,g,b) Among them, f θ :R 2 →R 3 represents the mapping defined by the parameters θ, which projects the two-dimensional image coordinates into a three-dimensional color space, where (r, g, b) represents the values ​​of the red, green, and blue channels respectively; In order to enhance the network's ability to capture spatial information, a spatial encoding mechanism is introduced. Specifically, the image coordinate X is mapped to a high-dimensional space through the encoding function γ, and then used as one of the inputs of the network. This process is expressed by the following formula: X'=γ(X) γ(x)=[sin(2 0 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)] Where X represents the image coordinate, X' represents the encoding coordinate, and L = 4 is a hyperparameter that determines the encoding dimension.

6. The method for learning implicit neural degenerate representation for unpaired image dehazing according to claim 1, characterized in that: The dense residual enhancement module is used to enhance the expressiveness of the IDRM module in terms of detailed modeling; The IDRM structure can effectively learn the continuous distribution of the overall image structure, but it will inevitably encounter some local redundant features during the processing. To address this problem, DREM is embedded as an auxiliary module during the training process to focus on removing redundant information in high-frequency features and enhance the ability to recover details in structural areas. The DREM module integrates residual learning and dense connection mechanism to construct a double-layer residual dense block structure; Through the initial convolutional layer and residual path, the feature response is strengthened, and then the feature hierarchy is gradually enriched through dense connection and local fusion. The final output features are improved through convolution and differential operation with the original input image, which effectively highlights the key image content under haze interference, helping IDRM to fit the implicit representation of the target image more accurately.