Wavelet transform and prior-based diffusion road image enhancement method
By combining wavelet transform, prior maps, and diffusion models, the problem of enhancing highway images in complex environments was solved, achieving efficient image enhancement while preserving the natural colors and texture details of the highway images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2025-03-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing highway image enhancement methods suffer from insufficient generalization ability and unstable enhancement effects when processing images in complex environments, especially in cases of lighting changes and occlusion, where they are difficult to effectively preserve details and structure.
By combining wavelet transform, prior maps, and diffusion models, a model is constructed using a training dataset and a Unet neural network to enhance highway images. Low-frequency and high-frequency information is captured using transmittance and texture prior maps for image enhancement.
It improves image processing speed and efficiency, preserves the natural colors and texture details of highway scenes, and provides a satisfying visual experience.
Smart Images

Figure CN120410864B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and image enhancement, specifically to a diffusion highway image enhancement method based on wavelet transform and priors. Background Technology
[0002] In recent years, with the development of intelligent transportation systems, the acquisition and processing of highway images has received increasing attention. Highway images play a crucial role in applications such as traffic monitoring, accident detection, and vehicle recognition. However, due to environmental factors, changes in lighting, and limitations of imaging equipment, these images often suffer from problems such as blurriness, low contrast, and color distortion. These challenges make traditional image processing methods insufficient to meet practical needs.
[0003] Wavelet transform, as an effective image processing technique, can analyze local features of images at different scales and has good time-frequency localization capabilities, making it suitable for processing images with multi-scale and multi-frequency characteristics. Therefore, it is widely used in image enhancement tasks to improve image sharpness and contrast.
[0004] Meanwhile, the introduction of prior knowledge can further improve image enhancement. By establishing physically meaningful prior models, it is possible to better adapt to the characteristics of highway images, especially image degradation in complex environments. Combining wavelet transform with the diffusion method of prior knowledge can effectively address the diversity and uncertainty in highway images, thereby improving image quality and meeting the high standards required for traffic management and monitoring.
[0005] In summary, the diffusion-based highway image enhancement method based on wavelet transform and prior knowledge will provide effective technical support for improving the clarity, contrast, and color reproduction of highway images, and promote the further development of intelligent transportation systems.
[0006] In recent years, road surface image enhancement methods can be divided into the following three types:
[0007] 1. Road surface image enhancement based on image processing methods;
[0008] 2. Road surface image restoration based on degradation model;
[0009] 3. Road surface image enhancement based on deep learning.
[0010] Image processing-based road image enhancement methods aim to improve the visual quality of road images by adjusting individual pixel values. Typically, these methods employ two input channels generated from a degraded image and adjust the image in terms of detail, structure, and illumination levels based on different sparse representation capabilities of the square norm of the image's spatial information, thereby enhancing image detail and contrast. While these methods have made some progress in improving road image quality, several limitations remain. First, image processing-based methods primarily focus on adjusting individual pixels, making it difficult to effectively capture and connect the global information of the image. In the context of highways, global information is crucial for reproducing the appearance of realistic scenes and objects. Due to the limitations of this local operation, the methods may lead to over- or under-enhancement of the image, especially in the presence of complex lighting and scattering conditions. Second, road images are subject to strong scattering and absorption of light, resulting in loss or blurring of details and structure. Image processing-based methods often struggle to effectively handle this complex optical distortion, thus posing a challenge in restoring image details and structure.
[0011] Road surface image enhancement methods based on degradation models aim to model the specific degradation process of highway images, thereby achieving inverse solving of undegraded road surface images. First, the complex physical processes of light propagation and scattering in the road environment make accurate modeling of image degradation difficult. Phenomena such as illumination changes, weather conditions, and vehicle types affect the quality of highway images in many ways, and these effects vary under different road surface conditions. Therefore, when modeling image degradation based on specific models, it is often difficult to fully consider the complexity of the highway environment, thus limiting the generalization ability of the methods. Second, due to the diversity of highway image acquisition conditions, such as different illumination, weather, and traffic densities, methods based on specific models exhibit poor robustness in different scenarios. This can lead to significant performance fluctuations in practical applications, making it difficult to achieve stable enhancement results in various road surface environments.
[0012] Deep learning-based highway image enhancement methods aim to improve the quality of road images by leveraging the powerful learning capabilities of neural networks to learn complex image features and patterns from large amounts of road image data. Generally, these methods use architectures such as convolutional neural networks (CNNs) to learn image mappings end-to-end for effective highway image enhancement. However, despite significant progress in highway image processing, deep learning still faces several challenges and limitations. First, deep learning methods typically require large amounts of labeled highway image data for training, and obtaining large-scale, richly labeled highway image datasets is a challenging task due to the diversity and complexity of highway scenes. Therefore, the generalization ability of deep learning models to various weather, lighting, and traffic conditions may be limited. Second, factors such as lighting variations, reflections, and occlusions in the highway environment have complex and non-linear effects on image quality, making it difficult for deep learning models to understand and simulate these physical processes. In summary, existing highway image enhancement models face three major challenges:
[0013] First, due to the complexity of the highway environment, the generalization ability of the method is limited;
[0014] Second, the images enhanced by current methods still have shortcomings in color and texture details;
[0015] Third, factors such as changes in lighting, interference, and occlusion in the highway environment have a complex and nonlinear impact on image quality, which makes it difficult for deep learning models to understand and simulate these physical processes. Summary of the Invention
[0016] The technical problem to be solved by the present invention is to provide a diffusion highway image enhancement method based on wavelet transform and prior art, so as to solve the problems existing in the prior art.
[0017] The technical solution adopted by the present invention to solve the above-mentioned technical problems includes the following steps:
[0018] A diffusion highway image enhancement method based on wavelet transform and prior knowledge includes the following steps:
[0019] Step A: Select training data and construct the training set;
[0020] Step B: Build the model based on the parameters defined in the configuration file;
[0021] Step C: Train the model using the training set, calculate the loss between the model's output value and the true value according to the defined loss function, backpropagate the calculated loss, and update and optimize the model's parameters; stop training when the prediction performance reaches the preset value, and save the model for testing.
[0022] Step D: Construct the model based on the model parameters saved during the training phase;
[0023] Step E: Test the model on the test set; calculate the quantitative indicators by comparing the test results with real images, and use the quantitative indicators to measure the performance of the model.
[0024] Furthermore, in step A, the training data includes a road surface image and its corresponding ground reality image, as well as corresponding transmission prior maps and texture prior maps. The transmission prior map is determined based on the red channel degradation region of the road surface image, and the texture prior map is calculated by an edge detection operator.
[0025] The training set consists of six parts: ① highway images ② Ground image x0; ③ Transmission prior map m1 corresponding to the highway image; ④ Texture prior map m2 corresponding to the highway image; ⑤ Transmission prior map m3 corresponding to the ground image; ⑥ Texture prior map m4 corresponding to the ground image.
[0026] Furthermore, in step B, the parameters of the configuration file include: input dimension, output dimension, number of channels in the intermediate layer, and number of residual block counts.
[0027] Furthermore, step C includes the following steps:
[0028] Step C1: Randomly select a time t from 1,...,T for training. Here, T is a pre-set parameter, set to 1000 in this method. Two random noise values ∈ [0,1] that follow a standard normal distribution N(0,1) are used. t and ∈ h , used for noise addition processing in the forward process of the diffusion model;
[0029] Step C2: Obtain the highway image After wavelet transform, the image is divided into low-frequency sub-bands of the highway image. and high-frequency subband The ground image x0 is divided into a low-frequency sub-band x1 and a high-frequency sub-band x2 by wavelet transform.
[0030] Step C3: Add noise to the low-frequency sub-band x1 of the ground image using the following formula:
[0031]
[0032] Obtain the image x with added noise t Each time t corresponds to an α t α here t These are constants defined in the model. From α tThe factorial to α1 is a constant;
[0033] Step C4: Extract the low-frequency subband of the highway image. The transmission prior art image m1 corresponding to the highway image and the image x with added noise. t Concatenated together, and input into the neural network ∈ θ1 In, here ∈ θ1 This refers to the UET neural network, and the formula is as follows:
[0034]
[0035] The output of the neural network is the predicted noise and the added noise ∈ t Calculate the loss function and update the network parameters;
[0036] Step C5: Add noise to the high-frequency subband x2 of the ground image using the following formula:
[0037]
[0038] Obtain the image x with added noise h Each time t corresponds to an α t α here t These are constants defined in the model. From α t The factorial to α1 is also a constant;
[0039] Step C6: Extract the high-frequency subband of the highway image. The transmission prior art image m1 corresponding to the highway image and the image x with added noise. h Concatenated together, and input into the neural network ∈ θ2 In, here ∈ θ2 This refers to the Unet neural network formula, as follows:
[0040]
[0041] The output of the neural network is the predicted noise and the added noise ∈ h Calculate the loss function and update the network parameters;
[0042] Step C7: Randomly generate noise map L T H T These two noise maps follow a standard normal distribution N(0,1). The time step i is defined from T to 1. Here, T is a pre-set hyperparameter, which is set to 1000 in this method. So i takes values from 1000 to 1. When i loops to 1, the two variables z1 and z2 are defined as 0; otherwise, z1 and z2 are defined as two random noises, both following a standard normal distribution N(0,1).
[0043] Step C8: For each time step i, L i-1 It is obtained through the following formula:
[0044]
[0045] H i-1 It is obtained through the following formula:
[0046]
[0047] Where, α i It is a constant defined in the model, and there is a corresponding α for each time step i. i L i It is L i-1 The image from the previous moment, H i It is H i-1 The image from the previous moment, From α i factorial to α1, ∈ θ1 ,∈ θ2 It is a neural network in the feedforward process, σ i It is variance. It is also a constant;
[0048] Step C9: Repeat the denoising operation using the previous image L. i Get L i-1 Then through L i-1 Get L i-2 Continue until L0 is obtained, and H0 is obtained in the same way. Return the denoising results L0 and H0. Use L0 and the low-frequency sub-band x1 of the ground image as the loss function to update the network parameters, and use H0 and the high-frequency sub-band x2 of the ground image as the loss function to update the network parameters.
[0049] Step C10: Save the parameter file of the trained model during the training process for testing.
[0050] Furthermore, step E includes the following steps:
[0051] Step E1: Randomly generate noise map L T H T These two noise maps follow a standard normal distribution N(0,1). The time step i is defined from T to 1. Here, T is a pre-set hyperparameter, which is set to 1000 in this method. So i takes values from 1000 to 1. When i loops to 1, the two variables z1 and z2 are defined as 0; otherwise, z1 and z2 are defined as two random noises, both following a standard normal distribution N(0,1).
[0052] Step E2, for each time step i, Li-1 It can be obtained through the following formula:
[0053]
[0054] H i-1 It can be obtained through the following formula:
[0055]
[0056] Where, α i It is a constant defined in the model, and there is a corresponding α for each time step i. i L i It is L i-1 The image from the previous moment, H i It is H i-1 The image from the previous moment, From α i factorial to α1, ∈ θ1 ,∈ θ2 It is a neural network in the feedforward process, σ i It is variance. It is also a constant;
[0057] Step E3: Repeat the denoising operation using the previous image L. i Get L i-1 Then through L i-1 Get L i-2 Continue until L0 is obtained, and the same applies to H0, and return the noise reduction results L0 and H0;
[0058] Step E4: Input L0 and H0 into the inverse wavelet module for inverse wavelet processing to obtain the enhanced image.
[0059] Compared with the prior art, the present invention has the following technical effects:
[0060] This invention organically combines highway image enhancement methods with wavelet transform, prior maps, and diffusion models. By introducing wavelet transform, the computational burden is effectively reduced when processing road images, thereby improving processing speed and efficiency. The characteristics of wavelet transform enable us to capture frequency information in the image more accurately while reducing computational costs, laying the foundation for subsequent processing steps.
[0061] Secondly, this invention fully considers the importance of prior images in highway scenes. By combining prior images, we successfully incorporated rich low-frequency information, making the image processing more consistent with the characteristics of real road surface environments. This step effectively alleviates the low-frequency color cast problem, providing more natural color reproduction for the enhanced image.
[0062] Most importantly, this invention cleverly incorporates a diffusion model, which, through its superior ability to generate high-frequency information, successfully preserves important texture details in highway scenes. The diffusion model's generative capabilities make the enhanced image more perceptually realistic, providing users with satisfactory visual results. This innovative design gives this invention a unique advantage in solving the problem of high-frequency texture loss, significantly improving the visual results of highway image enhancement.
[0063] In summary, this invention utilizes the generative capability of the diffusion model to produce enhanced results with satisfactory perceptual fidelity; based on wavelet transform, it greatly accelerates inference and reduces the use of computational resources without sacrificing information; and based on prior maps, it makes full use of the abundant low- and high-frequency information in the prior maps, enabling the model to effectively solve the problems of low-frequency color cast and loss of high-frequency texture details. Attached Figure Description
[0064] Figure 1 This is a flowchart illustrating the operation of the model in this invention. Detailed Implementation
[0065] All features disclosed in this specification, or steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps. To enable those skilled in the art to better understand this invention, it will be further described in detail below with reference to the accompanying drawings and the following embodiments.
[0066] Step A: Select training data, process the data, and construct the training set and test set.
[0067] The training data includes highway images and their corresponding ground reality maps, as well as corresponding transmission prior maps and texture prior maps. The transmission prior map is determined based on the red channel degradation region of the highway image, while the texture prior map is calculated by the edge detection operator. Therefore, our training set consists of six parts: ① Highway images ② Ground truth image x0; ③ Transmission prior image m1 corresponding to the highway image; ④ Texture prior image m2 corresponding to the highway image; ⑤ Transmission prior image m3 corresponding to the ground truth image; ⑥ Texture prior image m4 corresponding to the ground truth image. The training and testing sets are divided in a 10:1 ratio for subsequent training and testing operations, respectively.
[0068] Step B: Use the configuration file to create the model. The parameters of the configuration file include: input dimension, output dimension, number of channels in the intermediate layers, and number of residual buffer locks. Based on these parameters, an overall framework for the model can be created for training.
[0069] Step C: Train the model using the training set. The training process includes the following steps:
[0070] Step 1: Randomly select a time t from 1,...,T for training. Here, T is a pre-set parameter, set to 1000 in this method. Two random noise values ∈ [0,1] that follow a standard normal distribution N(0,1) are used. t and ∈ h , used for noise addition processing in the forward process of the diffusion model;
[0071] Step 2: Obtain the highway image After wavelet transform, the image is divided into low-frequency sub-bands of the highway image. and high-frequency subband The ground image x0 is divided into a low-frequency sub-band x1 and a high-frequency sub-band x2 by wavelet transform.
[0072] Step 3: Add noise to the low-frequency sub-band x1 of the ground image using the following formula:
[0073]
[0074] Obtain the image x with added noise t Each time t corresponds to an α t α here t These are constants defined in the model. From α t The factorial to α1 is a constant;
[0075] Step 4: Extract the low-frequency subband from the highway image. The transmission prior art image m1 corresponding to the highway image and the image x with added noise. t Concatenated together, and input into the neural network ∈ θ1 In, here ∈ θ1 This refers to the UET neural network, and the formula is as follows:
[0076]
[0077] The output of the neural network is the predicted noise and the added noise ∈ t Calculate the loss function and update the network parameters;
[0078] Step 5: Add noise to the high-frequency subband x2 of the ground image using the following formula:
[0079]
[0080] Obtain the image x with added noise h Each time t corresponds to an α t α here t These are constants defined in the model. From αt The factorial to α1 is also a constant;
[0081] Step 6: Extract the high-frequency subband of the highway image. The transmission prior art image m1 corresponding to the highway image and the image x with added noise. h Concatenated together, and input into the neural network ∈ θ2 In, here ∈ θ2 This refers to the Unet neural network formula, as follows:
[0082]
[0083] The output of the neural network is the predicted noise and the added noise ∈ h Calculate the loss function and update the network parameters;
[0084] Step 7: Randomly generate noise map L T H T These two noise maps follow a standard normal distribution N(0,1). The time step i is defined from T to 1. Here, T is a pre-set hyperparameter, which is set to 1000 in this method. So i takes values from 1000 to 1. When i loops to 1, the two variables z1 and z2 are defined as 0; otherwise, z1 and z2 are defined as two random noises, both following a standard normal distribution N(0,1).
[0085] Step 8: For each time step i, L i-1 It can be obtained through the following formula:
[0086]
[0087] H i-1 It can be obtained through the following formula:
[0088]
[0089] Where, α i It is a constant defined in the model, and there is a corresponding α for each time step i. i L i It is L i-1 The image from the previous moment, H i It is H i-1 The image from the previous moment, From α i factorial to α1, ∈ θ1 ,∈ θ2 It is a neural network in the feedforward process, σ i It is variance. It is also a constant;
[0090] Step 9: Repeat the denoising operation using the previous image L. i Get L i-1 Then through L i-1 Get L i-2 Continue until L0 is obtained, and the same applies to H0. Return the denoising results L0 and H0. Use L0 and the low-frequency sub-band x1 of the ground image as the loss function to update the network parameters, and use H0 and the high-frequency sub-band x2 of the ground image as the loss function to update the network parameters.
[0091] Step 10: Save the parameter file of the trained model during the training process for testing;
[0092] In the model testing phase, the overall framework of the model is constructed based on the parameters saved during training for testing purposes.
[0093] Step D: Construct the model based on the model parameters saved during the training phase;
[0094] Step E: Test the model using the test set. The testing process includes the following steps:
[0095] Step 1: Randomly generate noise map L T ,z T These two noise maps follow a standard normal distribution N(0,1). The time step i is defined from T to 1. Here, T is a pre-set hyperparameter, which is set to 1000 in this method. So i takes values from 1000 to 1. When i loops to 1, the two variables z1 and z2 are defined as 0; otherwise, z1 and z2 are defined as two random noises, both following a standard normal distribution N(0,1).
[0096] Step 2: For each time step i, L i-1 It can be obtained through the following formula:
[0097]
[0098] H i-1 It can be obtained through the following formula:
[0099]
[0100] Where, α i It is a constant defined in the model, and there is a corresponding α for each time step i. i L i It is L i-1 The image from the previous moment, H i It is H i-1 The image from the previous moment, From α i factorial to α1, ∈θ1 ,∈ θ2 It is a neural network in the feedforward process, σ i It is variance. It is also a constant;
[0101] Step 3: Repeat the denoising operation using the previous image L. i Get L i-1 Then through L i-1 Get L i-2 Continue until L0 is obtained, and the same applies to H0, and return the noise reduction results L0 and H0;
[0102] Step 4: Input L0 and H0 into the inverse wavelet module for inverse wavelet processing to obtain the enhanced image.
[0103] Although the present invention has been described herein with reference to embodiments thereof, the above embodiments are merely preferred embodiments of the present invention, and the implementation of the present invention is not limited to the above embodiments. It should be understood that those skilled in the art can design many other modifications and implementations, which will fall within the scope and spirit of the principles disclosed in this application.
Claims
1. A diffusion highway image enhancement method based on wavelet transform and prior knowledge, characterized in that, Specifically, the steps include the following: Step A: Select training data and construct the training set; Step B: Build the model based on the parameters defined in the configuration file; Step C: Train the model using the training set, calculate the loss between the model's output value and the true value according to the defined loss function, backpropagate the calculated loss, and update and optimize the model's parameters; stop training when the prediction performance reaches the preset value, and save the model for testing. Step C includes the following steps: Step C1: Randomly select a time t from 1,...,T for training, T=1000, and randomly add two noise sources that follow a standard normal distribution N(0, 1). and , used for noise addition processing in the forward process of the diffusion model; Step C2: Obtain the highway image After wavelet transform, the image is divided into low-frequency sub-bands of the highway image. and high-frequency subband ; ground image After wavelet transform, it is divided into low-frequency sub-bands of the ground image. and high-frequency subband ; Step C3: Low-frequency sub-band of the ground image Perform a noise-adding operation to obtain a noisy image. ; Step C4: Extract the low-frequency subband of the highway image. The transmission prior art image m1 corresponding to the highway image and the image with added noise. The pieces are concatenated together and then fed into the neural network. inside, Refers to the Unet neural network; The output of the neural network is the predicted noise and the added noise. Calculate the loss function and update the network parameters; Step C5: High-frequency subband of the ground image Perform a noise-adding operation to obtain a noisy image. ; Step C6: Extract the high-frequency subband of the highway image. The transmission prior art image m1 corresponding to the highway image and the image with added noise. The pieces are concatenated together and then fed into the neural network. inside, Refers to the Unet neural network; The output of the neural network is the predicted noise and the added noise. Calculate the loss function and update the network parameters; Step C7: Randomly generate noise maps , These two noise graphs follow a standard normal distribution N(0, 1). The time step i is defined as ranging from T to 1, where T=1000, and i ranges from 1000 to 1. When i cycles to 1, the time step i is... , Both variables are defined as 0; otherwise, , Defined as two random noises, both following a standard normal distribution N(0, 1); Step C8: For each time step i, It is obtained through the following formula: It is obtained through the following formula: in, It is a constant, and each time step i has a corresponding one. , yes The image from the previous moment, yes From arrive factorial, , It is a neural network in the feedforward process. It is variance. , is a constant; Step C9: Repeat the denoising operation using the previous image. get Then through get until you get ,get Similarly, return the noise reduction result. , ,Will Low-frequency subband of ground images Perform loss function, update network parameters, and High-frequency subband of ground images Calculate the loss function and update the network parameters; Step C10: Save the parameter file of the trained model during the training process for testing; Step D: Construct the model based on the model parameters saved during the training phase; Step E: Test the model on the test set; calculate the quantitative indicators by comparing the test results with real images, and use the quantitative indicators to measure the performance of the model.
2. The diffusion highway image enhancement method based on wavelet transform and prior as described in claim 1, characterized in that, In step A, the training data includes road surface images and their corresponding ground reality images, as well as corresponding transmission prior maps and texture prior maps; wherein, the transmission prior map is determined based on the red channel degradation region of the road surface image, and the texture prior map is calculated by the edge detection operator; The training set consists of six parts: ① highway images ② Ground images ; ③ Transmission prior map m1 corresponding to the highway image; ④ Texture prior map m2 corresponding to the highway image; ⑤ Transmission prior map m3 corresponding to the ground image; ⑥ Texture prior map m4 corresponding to the ground image.
3. The diffusion highway image enhancement method based on wavelet transform and prior as described in claim 1, characterized in that, In step B, the parameters of the configuration file include: input dimension, output dimension, number of channels in the intermediate layer, and number of residual blocks.
4. The diffusion highway image enhancement method based on wavelet transform and prior as described in claim 1, characterized in that, Step E includes the following steps: Step E1: Randomly generate noise maps , These two noise graphs follow a standard normal distribution N(0, 1). The time step i is defined as ranging from T to 1, where T=1000, and i ranges from 1000 to 1. When i cycles to 1, the time step i is... , Both variables are defined as 0; otherwise, , Defined as two random noises, both following a standard normal distribution N(0, 1); Step E2: For each time step i, It is obtained through the following formula: It is obtained through the following formula: in, It is a constant, and each time step i has a corresponding one. , yes The image from the previous moment, yes From arrive factorial, , It is a neural network in the feedforward process. It is variance. , is a constant; Step E3: Repeat the denoising operation using the previous image. get Then through get until you get ,get Similarly, return the noise reduction result. , ; Step E4, then , The image is input into the inverse wavelet module for inverse wavelet processing, resulting in the enhanced image.