Unsupervised image quality enhancement method, device, equipment, medium and product
Through the unsupervised image quality enhancement method, the image content and details are restored by using the generator in the feature extraction and image enhancement network, which solves the problem of quality degradation caused by bad weather and noise during the image acquisition process, and improves image clarity and autonomous driving recognition performance.
Patent Information
- Application Number
- CN202510452766.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
The image acquisition process in the real world is affected by factors such as bad weather conditions and sensor noise, resulting in reduced image visibility and reduced quality, making it difficult for the existing technology to effectively restore image content and details.
Using an unsupervised image quality enhancement method, the unsupervised feature enhancement module in the feature extractor and image enhancement network, including the first generator and the second generator, restore image content and details through the trained generator, and build a dual learning architecture to enhance image quality.
Effectively remove image artifacts, restore image content and details, improve image clarity, and improve visual recognition performance of autonomous driving systems.
Smart Images

Figure CN120387942A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image enhancement, and in particular, to an unsupervised image quality enhancement method, device, equipment, medium and product. Background Art
[0002] Due to the improvement of the reliability of vehicle environment perception, significant progress has been made in autonomous driving technology in recent years. On-vehicle RGB (RGB color mode) imaging cameras have become one of the most widely used sensors for ensuring vehicle navigation safety due to their high cost performance and installation convenience. However, the image acquisition process in the real world will inevitably be affected by various factors, such as bad weather conditions, sensor noise, and motion blur, resulting in reduced image visibility and degraded image quality. Summary of the Invention
[0003] The purpose of the present application is to provide an unsupervised image quality enhancement method, device, equipment, medium and product, which can restore the content of the image and improve the quality and quality of the image.
[0004] To achieve the above purpose, the present application provides the following solutions:
[0005] In a first aspect, the present application provides an unsupervised image quality enhancement method, including:
[0006] Obtaining a video frame image of the original vehicle environment;
[0007] Using a feature extractor to extract features from the video frame image to obtain initial image features;
[0008] Constructing an image enhancement network and training an unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a fourth discriminator; the unsupervised feature enhancement module includes a first generator and a second generator;
[0009] Inputting the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; wherein, the trained first generator is used to correct the content of the video frame image, and the trained second generator is used to restore the detailed content of the video frame image;
[0010] Determining a video frame image with enhanced features based on the enhanced image features.
[0011] In a second aspect, the present application provides an unsupervised image quality enhancement device, including:
[0012] An acquisition module, configured to acquire a video frame image of the original vehicle environment;
[0013] A feature extraction module, configured to extract features from the video frame image by using a feature extractor to obtain initial image features;
[0014] A network construction module, configured to construct an image enhancement network and train an unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a second discriminator group; the unsupervised feature enhancement module includes a first generator and a second generator;
[0015] An enhancement module, configured to input the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; wherein, the trained first generator is used to correct the content of the video frame image, and the trained second generator is used to restore the detailed content of the video frame image;
[0016] A fusion module, configured to determine a video frame image with enhanced features based on the enhanced image features.
[0017] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned unsupervised image quality enhancement method.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned unsupervised image quality enhancement method is implemented.
[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned unsupervised image quality enhancement method is implemented.
[0020] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0021] The present application provides an unsupervised image quality enhancement method, device, equipment, medium, and product. The first generator in the trained unsupervised feature enhancement module is used to remove artifacts from the video frame image of the original vehicle environment, so that the content of the video frame image can be restored. The second generator is used to restore the detailed content of the video frame image, making the image content clearer, thereby improving the image quality. At the same time, the unsupervised feature enhancement module consists of two generators, forming a plug-and-play dual learning architecture, and the unsupervised feature enhancement module can be applied to the vehicle-mounted system to enhance the quality of the driving video frame images captured by the camera in the vehicle-mounted system, improve the clarity of the images, and further improve the visual recognition performance of the vehicle-mounted system in the real world, effectively enhancing the recognition performance of the autonomous driving system. Brief Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is an application environment diagram of an unsupervised image quality enhancement method in an embodiment of the present application;
[0024] Figure 2 It is a schematic flowchart of an unsupervised image quality enhancement method provided in an embodiment of the present application;
[0025] Figure 3 It is a schematic diagram of the VGG16 network structure provided in an embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of the training process of an unsupervised feature enhancement module provided in another embodiment of the present application;
[0027] Figure 5 It is a schematic diagram for comparing the ablation study results of inserting UFEM in different network layers provided in an embodiment of the present application;
[0028] Figure 6 It is a schematic diagram of the performance results of testing different UFEM variants provided in an embodiment of the present application;
[0029] Figure 7 It is the result of the data sensitivity experiment on UFEMs obtained from different training sets provided in another embodiment of the present application. Among them, (a) is a schematic diagram of the result of the data sensitivity experiment using UFEMs trained by selecting different numbers of images in the Fog3 dataset, and (b) is a schematic diagram of the result of the data sensitivity experiment using UFEMs trained by selecting different numbers of images in the ExDARK and RTTS datasets.
[0030] Figure 8 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed Description of the Embodiments
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0032] To make the objectives, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] The unsupervised image quality enhancement method provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, placed in the cloud or on other servers. The terminal 102 can send the video frame images of the original vehicle environment to the server 104. After receiving the video frame images of the original vehicle environment, for the video frame images of the original vehicle environment, the server 104 uses a feature extractor to extract features from the video frame images to obtain initial image features; constructs an image enhancement network, and trains the unsupervised feature enhancement module in the image enhancement network; inputs the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; determines the video frame images with enhanced features based on the enhanced image features. The server 104 can feedback the obtained video frame images with enhanced features to the terminal 102. In addition, in some embodiments, the unsupervised image quality enhancement method can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly process the video frame images of the original vehicle environment, or the server 104 can obtain the video frame images of the original vehicle environment from the data storage system and process the video frame images of the original vehicle environment.
[0034] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, and tablet computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0035] In an exemplary embodiment, as Figure 2 shown in the figure, an unsupervised image quality enhancement method is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in the figure as an example for description, it includes the following steps 201 to 205. Among them:
[0036] Step 201, obtain the video frame images of the original vehicle environment.
[0037] Specifically, the video frame image is an image obtained after processing the video captured by the capturing device, which can be a vehicle-mounted camera or other capturing devices.
[0038] Step 202: Use a feature extractor to extract features from the video frame image to obtain initial image features.
[0039] Step 203: Construct an image enhancement network and train the unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a second discriminator group; the unsupervised feature enhancement module includes a first generator and a second generator.
[0040] Step 204: Input the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; among them, the trained first generator is used to correct the content of the video frame image, and the trained second generator is used to restore the detailed content of the video frame image.
[0041] Step 205: Determine the video frame image with enhanced features based on the enhanced image features.
[0042] Implementing the above steps 201 to 205, the content of the video frame image can be restored by using the first generator and the second generator in the trained unsupervised feature enhancement module, improving the quality and quality of the image.
[0043] In an exemplary embodiment, the training process of the unsupervised feature enhancement module includes steps 301 - 306:
[0044] Step 301: Construct a training set; the training set includes a plurality of training samples, and the training samples include clear images and degraded images.
[0045] In this embodiment, one large synthetic dataset and seven real-world degradation datasets are selected for training. Among them, three datasets, namely ImageNet-C, Haze-20, and ExDARK-Exclusively Dark Image Dataset, are used for image classification tasks; two datasets, namely RTTS-Realworld TaskDriven Testing Set and DAWN Fog-Detection in AdverseWeather Nature Fog, are used for object detection tasks; and three datasets, namely ACDC Nightime-Adverse Conditions Dataset with Correspondences, Dark_Zurich, and Nighttime Driving, are used for semantic segmentation tasks.
[0046] Step 302: Use a feature extractor to extract features from the clear image and the degraded image respectively to obtain clear features and degraded features.
[0047] In this embodiment, the VGG16 (Visual Geometry Group 16-layer network) network is adopted. As Figure 3 shown, the VGG16 is divided into a shallow pre-trained layer (SPL) and a deep pre-trained layer (DPL), and the "Conv3_2" layer in the SPL is used as the feature extractor.
[0048] Step 303: Calculate a first loss function based on the clear features, the degraded features, the first generator, the third generator, and the first discriminator group; the first loss function includes a first adversarial loss, a cycle consistency loss, and an identity preservation loss.
[0049] As Figure 4 shown, the training process of the first generator G D2C is completed in the first stage, which can restore the spatial content and remove artifacts, so that the trained first generator G D2C corrects the content of the image. For the convenience of expressing the input and output of the first generator G D2C and the third generator G C2D , Figure 4 among the multiple first generators G D2C in are all the same first generator GD2C , multiple third generators G C2D are all the same third generator G C2D , and the specific process is as follows:
[0050] Input the degraded feature X D ( Figure 4 the DF in it) into the first generator G D2C to generate a pseudo-enhanced feature Fake_CF; input the pseudo-enhanced feature Fake_CF into the third generator G C2D to generate a double pseudo-degraded feature; input the degraded feature DF into the third generator G C2D to generate a degradation-maintaining feature; input the pseudo-enhanced feature( Figure 4 the Fake_CF in it) and the clear feature( Figure 4 the CF in it) into the first discriminator group to obtain a first discrimination result; calculate a first adversarial loss according to the first discrimination result; calculate a cycle consistency loss according to the degraded feature and the double pseudo-degraded feature; calculate an identity preservation loss according to the degraded feature and the degradation-maintaining feature.
[0051] The first discriminator group learns to classify the clear feature as class 1, classify the pseudo-enhanced feature as class 0, and provide feedback to the first generator G D2C . To effectively solve the discrimination challenge brought by high sparsity and prevent error accumulation in the propagation process of enhanced features in the deep network, a multiple adversarial mechanism is incorporated. Specifically, the first discriminator group includes multiple discriminators, and inserts the multiple discriminators in the first discriminator group into different convolutional layers in the DPL respectively, and discriminates the pseudo-enhanced feature and the clear feature after convolution processing by different convolutional layers. Therefore, in this embodiment, the multiple convolutional layers where the first discriminator group is located in the DPL are used as the feature extractor of the first discriminator group, and other models can also be selected as the feature extractor, which is not limited in this implementation.
[0052] In this embodiment, the processing process of the first discriminator group is illustrated by taking three discriminators as an example , and the value of N is 3. The first discriminator group includes the first discriminator the second discriminator and the third discriminator
[0053] Input the pseudo-enhanced feature( Figure 4 the Fake_CF in it) and the clear feature( Figure 4 the CF in it) into the first discriminator in the first discriminator group through one layer of convolution Obtain the first independent discrimination result, and calculate the first independent adversarial loss according to the first independent discrimination result and the clear feature. The calculation formula of the first independent adversarial loss is:
[0054]
[0055] where L adv_1 is the first independent adversarial loss, and G D2C (X D )1 are the features obtained by subjecting the clear feature and the pseudo-enhanced feature to one-layer convolution processing respectively.
[0056] Input the pseudo-enhanced feature (Fake_CF in Figure 4 ) and the clear feature (CF in Figure 4 ) into the second discriminator in the first discriminator group through two-layer convolution to obtain the second independent discrimination result, and calculate the second independent adversarial loss according to the second independent discrimination result and the clear feature. The calculation formula of the second independent adversarial loss is:
[0057]
[0058] where L adv_2 is the second independent adversarial loss, and G D2C (X D )2 are the features obtained by subjecting the clear feature and the pseudo-enhanced feature to two-layer convolution processing respectively.
[0059] Input the pseudo-enhanced feature (Fake_CF in Figure 4 ) and the clear feature (CF in Figure 4 ) into the third discriminator in the first discriminator group through three-layer convolution to obtain the third independent discrimination result, and calculate the third independent adversarial loss according to the third independent discrimination result and the clear feature. The calculation formula of the third independent adversarial loss is:
[0060]
[0061] where L adv_3 is the third independent adversarial loss, and G D2C (X D )3 are the features obtained by subjecting the clear feature and the pseudo-enhanced feature to three-layer convolution processing respectively.
[0062] Finally, the first adversarial loss can be extended to Equation (4), which helps to more clearly distinguish between degraded features and clear features and effectively address the challenges brought about by the high sparsity and low effective information ratio of large-sized features.
[0063] The calculation formula for the first adversarial loss is:
[0064]
[0065] To ensure a meaningful pairing of the input degraded feature DF and its corresponding pseudo-enhanced feature, the pseudo-enhanced feature is input into the third generator G C2D , and the cycle consistency loss L cyc is applied to constrain the cycle conversion result G C2D (G D2C (DF)) to match the input degraded feature DF. The calculation formula for the cycle consistency loss is:
[0066] L cyc = ||X D - G C2D (G D2C (X D ))||1; (5)
[0067] The identity preservation loss is introduced to ensure the stability and better convergence of the solution in this implementation. The calculation formula for the identity preservation loss is:
[0068] L idt = ||X D - G C2D (X D )||1; (6)
[0069] Among them, L mul_adv is the first adversarial loss, L cyc is the cycle consistency loss, L idt is the identity preservation loss, G D2C is the first generator, G C2D is the third generator, is the discriminator in the first discriminator group at the k-th layer of convolution, N is the total number of network layers of the feature extractor in the first discriminator group, W k is the weight of the discriminator at the k-th layer of convolution, X C is the clear feature, X D is the degraded feature, and G D2C (X D ) k are the features obtained by processing the clear feature and the pseudo-enhanced feature through k layers of convolution respectively, X C is the clear feature, X D is the degraded feature, ‖‖1 is the L1 norm loss.
[0070] As shown Figure 4 in the figure, the second discriminator group also includes a plurality of discriminators. In this embodiment, the second discriminator group includes a fifth discriminator a sixth discriminator and a seventh discriminator The third adversarial loss L is obtained by using the second discriminator group m_adv , which is used for training the third generator. The clear feature CF is input into the third generator G C2D to obtain a pseudo-clear feature (such as Fake_DF in Figure 4 ), and the degraded feature DF and the pseudo-clear feature Fake_DF are input into the second discriminator group for discrimination to obtain a third discrimination result. The second discriminator group will learn to classify the degraded feature DF as 1 and the pseudo-clear feature Fake_DF as 0. The multiple discriminators in the second discriminator group are similar to the multiple discriminators in the first discriminator group and will not be elaborated in this embodiment.
[0071] Step 304: Train the first generator based on the first loss function to obtain a trained first generator.
[0072] The calculation formula of the first loss function is:
[0073] L Stage1 = λ1·L mul_adv + λ2·L cyc + λ3·L idt ; (7)
[0074] where L Stage1 is the first loss, and λ1, λ2, and λ3 are weighting coefficients.
[0075] After the first stage of training, the trained first generator G D2C is obtained and retained, and then it is inserted into the existing model to restore the latent content and remove the additional artifacts in the degraded feature, correcting the content of the image.
[0076] Step 305: Calculate a second loss function based on the clear feature, the degraded feature, the second generator, and the fourth discriminator; the second loss function includes a correlation consistency loss, a content consistency loss, and a second adversarial loss.
[0077] As shown Figure 4 in the figure, the training process of the second generator G E2C is completed in the second stage. The second stage is a global modulation stage for restoring the detailed content of the image.
[0078] The degraded feature (i.e., the degraded feature) is input into the trained first generator G D2CObtain the pseudo-enhanced feature X E ( Figure 4 in the EF_stage1 in
[0079] In the second stage, input the pseudo-enhanced feature EF_stage1 and the clear feature CF into the DPL for training.
[0080] Input the pseudo-enhanced feature EF_stage1 into the second generator G E2C to obtain the double pseudo-enhanced feature EF_stage2; input the double pseudo-enhanced feature EF_stage2 and the clear feature CF into the fourth discriminator D F to obtain the second discrimination result. The fourth discriminator D F learns to classify the double pseudo-enhanced feature EF_stage2 as 0 and the clear feature CF as 1, and feeds back to the second generator G E2C Feedback.
[0081] Calculate the channel correlation matrix of the clear feature CF according to the feature map of the clear feature CF; calculate the channel correlation matrix of the double pseudo-enhanced feature EF_stage2 according to the feature map of the double pseudo-enhanced feature EF_stage2; calculate the correlation consistency loss according to the channel correlation matrix of the clear feature CF and the channel correlation matrix of the double pseudo-enhanced feature EF_stage2; calculate the content consistency loss according to the double pseudo-enhanced feature EF_stage2 and the clear feature; calculate the second adversarial loss according to the second discrimination result.
[0082] The calculation formula of the second adversarial loss is:
[0083] L op_adv =log(D F (X C ))+log(1-D F (G E2C (X E ));(8)
[0084] To better perceive the statistical properties related to degradation cues, the gap between the double pseudo-enhanced feature EF_stage2 and the clear feature CF is minimized during the training process. In the second-stage training, the feature maps extracted from the "Conv1_2" layer, "Conv2_2" layer, "Conv3_3" layer, and "Conv4_3" layer in VGG16 are selected to calculate the channel correlation matrix representations of the clear feature CF and the double pseudo-enhanced feature EF_stage2, and a correlation consistency loss function is used for constraint. In this embodiment, the "Conv1_2" layer, "Conv2_2" layer, "Conv3_3" layer, and "Conv4_3" layer in VGG16 are used as the channel correlation matrix feature extractors. Other convolutional layers can also be selected as the channel correlation matrix feature extractors, which is not limited in this implementation.
[0085] The calculation formula for the correlation consistency loss is:
[0086]
[0087] It is necessary to introduce a loss function during training to ensure the maximum retention of information in the pseudo-enhanced feature EF_stage1. Therefore, a content consistency loss is introduced to constrain the content consistency between the double pseudo-enhanced feature EF_stage2 and the pseudo-enhanced feature EF_stage1 to achieve semantic fidelity.
[0088] In the specific implementation process, the content consistency loss between the pseudo-enhanced feature EF_stage1 and the clear feature is calculated based on the pseudo-enhanced feature EF_stage1 and the clear feature, and the content consistency loss between the double pseudo-enhanced feature EF_stage2 and the clear feature is calculated based on the double pseudo-enhanced feature EF_stage2 and the clear feature. Through experimental verification, calculating the content consistency loss based on the double pseudo-enhanced feature EF_stage2 and the clear feature is more effective in enhancing the detailed content of the image. Therefore, in this embodiment, only the content consistency loss between the double pseudo-enhanced feature EF_stage2 and the clear feature is retained, and the calculation formula for the content consistency loss is:
[0089]
[0090] Among them, L correlation is the correlation consistency loss, L content is the content consistency loss, L op_adv is the second adversarial loss, G E2C is the second generator, D F is the fourth discriminator, X C is the clear feature, X E is the pseudo-enhanced feature, G l and They are the channel correlation matrix of the clear features and the channel correlation matrix of the double pseudo-enhanced features of the l-th layer, respectively, and W l is the weight of the l-th layer, a and b are the feature maps of the clear features and the double pseudo-enhanced features in the l-th layer, respectively, and V l and are the feature content representations of the clear features and the double pseudo-enhanced features in the l-th layer, respectively. L is the total number of network layers of the channel correlation matrix feature extractor, and i and j are the pixel coordinate position indices in the l-th layer.
[0091] Step 306: Train the second generator based on the second loss function to obtain the trained second generator.
[0092] Expression of the second loss function:
[0093] L Stage2 = λ3·L correlation + λ4·L op_adv + λ5·L content ; (11)
[0094] where L Stage2 is the second loss, and λ3, λ4, and λ5 are weighting coefficients.
[0095] After the above training, only the trained first generator G D2C and the trained second generator G E2C are retained and frozen. The trained first generator G D2C and the trained second generator together form an unsupervised feature enhancement module UEFM.
[0096] Insert UFEM into VGG16 to perform recognition tests on the degraded images in the existing dataset. Specifically, as Figure 4 shown, UFEM is seamlessly inserted after the "Conv3_2" layer of VGG16. That is, after extracting features from the shallow pre-training layer, UFEM is used to enhance the features in the degraded images, thereby enhancing the image features.
[0097] Extensive experiments have been carried out on three advanced vision tasks in autonomous driving to evaluate the processing effects of UFEM on three advanced vision tasks, namely image classification tasks, object detection tasks, and semantic segmentation tasks, so as to verify the generalization ability of the designed UFEM of this application for computer vision tasks.
[0098] In the degraded image classification task, this application conducts a quantitative analysis of UFEM with 20 IR (Image Restoration) and UDA (Unsupervised Domain Adaption) methods to fully verify the effectiveness and generality of UFEM. The quantitative analysis results show that the method of this application is significantly superior to IR and UDA methods in terms of classification accuracy on synthetic and real degradation datasets, and can significantly improve the accuracy of the pre-trained network. The method of this application is compared with five other different methods in terms of feature correction: the method of this application significantly enhances the feature response of the image discriminant region. A comparison is made between the UFEM and IR methods in this application in terms of accurately attracting the detector's attention using grad-CAMs: UFEM better enhances the discriminant region around the object, achieving focused attention and accurate prediction, which is beneficial for deep network recognition. The distributions of degraded features, clear features, and enhanced features are visualized using T-SNE and compared with the feature distributions of other IR methods: the distributions between the enhanced features and clear features enhanced by UFEM in this application are effectively aligned, thus improving the performance of existing models on degraded images. A comparison is made of the feature distributions between different layers: the distributions between the enhanced features and clear features enhanced by UFEM in this application are effectively aligned, reducing the error accumulation in forward propagation.
[0099] In the foggy object detection task, a quantitative analysis is conducted with six IR methods on two real fog object detection datasets, RTTS and DAWN: the method of this application shows consistent performance improvement in two real foggy scenarios, outperforming other methods. A comparison is made between UFEM and IR methods in terms of accurately attracting the detector's attention using grad-CAMs: UFEM better enhances the discriminant region around the object, enabling the detector to focus attention and predict the correct category and spatial location of the object. A comparison is made between UFEM and the YOLOv5 method in terms of object detection results: UFEM significantly reduces false alarms in environmental perception.
[0100] In the nighttime semantic segmentation task, this application is compared with six IR methods on three real dark segmentation datasets, ACDC Nighttime, Dark_Zurich, and Nighttime Driving: the method of this application achieves excellent feature correction at various resolutions, improving the performance of all three real datasets; a qualitative comparison is made between UFEM and IR methods: the method of this application can obviously produce more accurate segmentation results for distant and dense objects in the dark.
[0101] This application evaluates the method proposed in this application through the ablation experiment analysis results. Ablation experiment of the multi-adversarial mechanism: Ablation experiments were conducted on two datasets, ExDARK and RTTS, by inserting different numbers of discriminators into three different methods, VGG16, ResNet50, and YOLOv5. The results verified the positive effect of the multi-adversarial mechanism in stage 1 on enhancing the robustness and generality of UFEM; Ablation experiment of two-stage correction: Ablation experiments were conducted on the ExDARK and RTTS datasets for three different methods, VGG16, ResNet50, and YOLOv5, using different stages. The results showed that sequentially performing content restoration and correlation modulation is more beneficial for feature correction; Ablation experiment of the inserted layer of UFEM: The ablation experiment structure for evaluating the performance of inserting UFEMs of different scales into different layers of VGG16 or ResNet50 on Fog3 of ImageNet-C is as Figure 5 shown. The results indicate that inserting UFEM into the shallow network of the network is a better choice; Ablation experiment of content consistency and correlation consistency: The ablation experiment results of three different implementation forms (distance, KL divergence, and cosine similarity) in the second stage highlight the necessity in stage 2 and minimizing the L1 distance of the matrix always provides the best performance; Ablation experiment of the computational overhead of UFEM: Three UFEM variants were designed in the U-Net architecture, each containing 1, 2, and 3 downsampling layers, and then these variants were integrated into the existing YOLOv5 for evaluation. The results are as Figure 6 shown. To balance the improvement in accuracy and computational overhead, this application selected the UFEM structure with 2 downsampling layers in all experiments; Finally, considering the challenge of collecting large-scale degraded images in the real world, this application conducted a data sensitivity experiment by training UFEM with different numbers of images. The results are as Figure 7 shown, Figure 7 where (a) in Figure 7 is to select different numbers of images in the Fog3 dataset for training, and
[0102] where (b) in
[0102] is to select different numbers of images in the ExDARK and RTTS datasets for training. The method in this application shows significant improvements in all aspects.
[0102] In summary, the unsupervised feature enhancement model UFEM proposed in this application can effectively utilize the dual learning architecture. In the first stage, it fully restores the latent content and corrects the image by eliminating additional artifacts through the dual learning architecture. In the second stage, it fully realizes the fine modulation of channel correlation and restores the detailed content of the image.
[0103] This application evaluated the effectiveness of UFEM for 5 different methods on 3 tasks using a large synthetic dataset and seven real-world degradation datasets. The pre-trained VGG16, ResNet50, YOLOv5, and DeepLabv3 were used as baseline models, and UFEM was trained for 200 epochs using the Adam optimizer with a batch size of 5. The initial learning rates of the first generator, the second generator, the first discriminator group, and the second discriminator group were set to 2e -4 and 1e -4。The learning rate is multiplied by 0.1 every 10 epochs. All experiments were run on an Nvidia A100 GPU. To obtain a classifier and detector with good performance on clear images, this application fine-tuned the model (i.e., pre-trained the model). For image classification on Haze-20, 100 clear images per class (80 for training, 20 for validation, and the rest for testing) were randomly selected from Haze Clear-20 (the clear class in the Haze-20 dataset) to fine-tune the pre-trained classifier, obtaining a classifier for 20 classes; for image classification on ExDARK, its fine-tuning was similar to the above, but there was no clear set corresponding to ExDARK. Therefore, images of the same classes as ExDARK were manually selected from the ImageNet, PASCAL VOC, and COCO datasets, and obvious low-light images were filtered out, finally obtaining the dataset ExDark-clean of clear images for fine-tuning; for object detection on RTTS and DAWN, the Cityscapes dataset was selected as the clear image set, and its classes were filtered to be consistent with RTTS (5 classes) and DAWN (6 classes) respectively. Then, the clear image set was fine-tuned to obtain a detector with good performance on the clear image set. Different from the above, for ImageNet-C and the three nighttime semantic segmentation datasets, whose classes are consistent with the official pre-trained model, this application did not perform fine-tuning but selected the pre-trained model as the baseline. For the UIR method on the synthetic dataset, this application randomly selected 1200 unpaired image pairs from each degradation level for model training and verified on the test set of 50,000 images from ImageNet-C; for the UDA method on the synthetic dataset, the model needed to be re-trained with unpaired but semantically labeled images. This application selected 10 unpaired image pairs from each class in ImageNet and ImageNet-C to form a dataset of 10K image pairs for re-training; for the UIR method on the real dataset, this application selected approximately 500 unpaired image pairs from Haze-20 and ExDARK for the training of the UIR method. The weights obtained from Haze-20 were subsequently used for testing on RTTS and DAWN. However, for the nighttime segmentation datasets, due to the lack of a clear set, this application only used the pre-trained weights released by the UIR method. In addition, when the number of images exceeds 100, the performance gain tends to saturate. Therefore, this application defaulted to using 100 unpaired image pairs to train UFEM. All results confirmed the effectiveness of UFEM in various degradation scenarios and tasks, demonstrating its ability to endow autonomous vehicles with efficient and highly robust visual perception under real degradation conditions.
[0104] Based on the same inventive concept, an embodiment of the present application further provides an unsupervised image quality enhancement device for implementing the above-mentioned unsupervised image quality enhancement method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the unsupervised image quality enhancement device provided below can refer to the limitations on the unsupervised image quality enhancement method in the above text, and will not be repeated here.
[0105] In an exemplary embodiment, an unsupervised image quality enhancement device is provided, including:
[0106] An acquisition module, configured to acquire video frame images of the original vehicle environment.
[0107] A feature extraction module, configured to perform feature extraction on the video frame images by using a feature extractor to obtain initial image features.
[0108] A network construction module, configured to construct an image enhancement network and train the unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a second discriminator group; the unsupervised feature enhancement module includes a first generator and a second generator.
[0109] An enhancement module, configured to input the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; wherein, the trained first generator is used to remove artifacts in the video frame images, and the trained second generator is used to restore the detailed content of the video frame images.
[0110] A fusion module, configured to determine a video frame image with enhanced features based on the enhanced image features.
[0111] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store video frame images of the original vehicle environment. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an unsupervised image quality enhancement method.
[0112] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0113] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0114] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0115] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0117] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0118] The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0120] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An unsupervised image quality enhancement method, characterized in that, The unsupervised image quality enhancement method includes: Obtain video frame images of the original vehicle environment; Use a feature extractor to extract features from the video frame images to obtain initial image features; Construct an image enhancement network and train the unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a fourth discriminator; the unsupervised feature enhancement module includes a first generator and a second generator; Input the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; among them, the trained first generator is used to correct the content of the video frame images, and the trained second generator is used to restore the detailed content of the video frame images; Determine the video frame images with enhanced features based on the enhanced image features.
2. The unsupervised image quality enhancement method according to claim 1, wherein The training process of the unsupervised feature enhancement module includes: Construct a training set; the training set includes multiple training samples, and the training samples include clear images and degraded images; Use a feature extractor to extract features from the clear images and degraded images respectively to obtain clear features and degraded features; Calculate a first loss function based on the clear features, the degraded features, the first generator, the third generator, and the first discriminator group; the first loss function includes a first adversarial loss, a cycle consistency loss, and an identity preservation loss; Train the first generator based on the first loss function to obtain the trained first generator; Calculate a second loss function based on the clear features, the degraded features, the second generator, and the fourth discriminator; the second loss function includes a correlation consistency loss, a content consistency loss, and a second adversarial loss; Train the second generator based on the second loss function to obtain the trained second generator.
3. The unsupervised image quality enhancement method according to claim 2, wherein Calculating the first loss function based on the clear features, the degraded features, the first generator, the third generator, and the first discriminator group specifically includes: Input the degraded features into the first generator to generate pseudo-enhanced features; Input the pseudo-enhanced features into the third generator to generate double pseudo-degraded features; Input the degraded features into the third generator to generate degraded maintenance features; Input the pseudo-enhanced features and the clear features into the first discriminator to obtain a first discrimination result; Calculate the first adversarial loss according to the first discrimination result; Calculate the cycle consistency loss according to the degraded features and the double pseudo-degraded features; Calculate the identity preservation loss according to the degraded features and the degraded maintenance features.
4. The unsupervised image quality enhancement method according to claim 2, wherein Calculating the second loss function based on the clear features, the degraded features, the second generator, and the fourth discriminator specifically includes: Input the pseudo-enhanced features into the second generator to obtain double pseudo-enhanced features; Input the double pseudo-enhanced features and the clear features into the fourth discriminator to obtain a second discrimination result; Calculate the channel correlation matrix of the clear features according to the feature maps of the clear features; Calculate the channel correlation matrix of the double pseudo-enhanced features according to the feature maps of the double pseudo-enhanced features; Calculate the correlation consistency loss based on the channel correlation matrix of the clear feature and the channel correlation matrix of the dual pseudo-enhanced feature; Calculate the content consistency loss based on the dual pseudo-enhanced feature and the clear feature; Calculate the second adversarial loss according to the second discrimination result.
5. The unsupervised image quality enhancement method according to claim 3, wherein The first discriminator group includes multiple discriminators; The calculation formula for the first adversarial loss is: The calculation formula for the cycle consistency loss is: L cyc = ||X D - G C2D (G D2C (X D ))||1; The calculation formula for the identity preservation loss is: L idt = ||X D - G C2D (X D )||1; Among them, L mul_adv is the first adversarial loss, L cyc is the cycle-consistency loss, L idt is the identity-preserving loss, G C2D is the third generator, G D2C is the first generator, is the discriminator in the k-th layer convolution of the first discriminator group, N is the total number of network layers of the feature extractor of the first discriminator group, W k is the weight of the discriminator in the k-th layer convolution, X C is the clear feature, X D is the degraded feature, and G D2C (X D ) k are the features obtained by processing the clear feature and the pseudo-enhanced feature through k-layer convolution respectively, and || ||1 is the L1 norm loss.
6. The unsupervised image quality enhancement method according to claim 2, wherein The calculation formula for the second adversarial loss is: L op_adv = log(D F (X C )) + log(1 - D F (G E2C (X E )); The calculation formula for the correlation consistency loss is: The calculation formula for the content consistency loss is: Among them, L correlation is the relevant consistency loss, L content is the content consistency loss, L op_adv is the second adversarial loss, G E2C is the second generator, D F is the fourth discriminator, X C is the clear feature, X E is the pseudo-enhanced feature, G l and are the channel correlation matrices of the clear feature and the double pseudo-enhanced feature of the l-th layer respectively, W l is the weight of the l-th layer, a and b are the feature maps of the clear feature and the double pseudo-enhanced feature in the l-th layer respectively, V l and are the feature content representations of the clear feature and the double pseudo-enhanced feature in the l-th layer respectively, L is the total number of network layers of the channel correlation matrix feature extractor, i and j are the pixel coordinate position indices in the l-th layer, and ||||1 is the L1 norm loss.
7. An unsupervised image quality enhancement device, characterized in that, The unsupervised image quality enhancement device includes: An acquisition module for acquiring video frame images of the original vehicle environment; A feature extraction module for extracting features from the video frame images by using a feature extractor to obtain initial image features; A network construction module for constructing an image enhancement network and training the unsupervised feature enhancement module in the image enhancement network; the image enhancement network includes an unsupervised feature enhancement module, a third generator, a first discriminator group, and a second discriminator group; the unsupervised feature enhancement module includes a first generator and a second generator; An enhancement module for inputting the initial image features into the trained unsupervised feature enhancement module to obtain enhanced image features; wherein, the trained first generator is used to correct the content of the video frame images, and the trained second generator is used to restore the detailed content of the video frame images; A fusion module for determining the video frame images with enhanced features based on the enhanced image features.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the unsupervised image quality enhancement method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the unsupervised image quality enhancement method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the unsupervised image quality enhancement method according to any one of claims 1-6.