Object Detection Method in Severe Weather Scenarios Based on Degradation Consistency
By using a global consistency network for style migration in bad weather scenarios, combining basic detection networks and local consistency networks, the problem of low target detection accuracy in bad weather is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202310881935.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-07-18
AI Technical Summary
In severe weather scenarios, existing object detection methods face the problems of lack of annotation images, loss of targets by two-stage detectors and damaged local features of targets, resulting in low detection accuracy.
A target detection method in bad weather scenarios based on degradation consistency is proposed, and style migration is performed through global consistency network (GCN). The basic detection network (BDN) adopts a single-stage detector, and the local consistency network (LCN) is used to fine-tune the features to make it closer to the low-quality images in the original bad weather scenario.
It effectively alleviates the problem of insufficient image labeling in bad weather scenarios, avoids target loss caused by regional mechanisms, and improves the accuracy of target detection, especially in rainy and foggy weather conditions, which perform better than the most advanced detection algorithms.
Smart Images

Figure CN117037050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically provides a target detection method in a bad weather scenario based on degradation consistency, which can be used for tasks such as target detection, target tracking, and autonomous driving in a bad weather scenario. Background Art
[0002] Target detection algorithms can identify and locate valuable foreground information in the captured images, and are widely used in autonomous driving, video surveillance, and augmented reality. However, in the actual application process, the detection performance of target detectors trained on datasets with good image quality will be significantly reduced in images of bad weather scenarios. By collecting a large amount of image data of bad weather scenarios and annotating the targets, the above problems can be alleviated. However, this workload requires a large amount of manpower, material resources, and financial resources. Common bad weather includes rain, fog, snow, and haze. Existing datasets such as Foggy Cityscape, RTTS, and UFDD are usually smaller than datasets with good large-scale image quality, and cannot fully train the detector, and often suffer from overfitting.
[0003] Domain adaptation technology or domain transfer technology is an effective method to alleviate the shortage of target detection datasets. The source domain consists of a large number of images with good image quality and annotations, while the target domain consists of small-scale images of bad weather scenarios. Then, in domain adaptation technology, there are often problems such as missing or damaged target features. To correct the damaged features, Prior-Adversarial Loss (PAL) and Residual Feature Recovery Block (RFRB) are proposed. However, the prior loss uses the two-stage Faster R-CNN as the baseline detector. Faster R-CNN extracts proposals through the Region Proposal Network. Bad weather can cause the region proposals of Faster R-CNN to be lost, thus reducing the final detection performance. Another shortcoming of the prior loss is that feature enhancement only targets the global features of the backbone network. However, target detection is more sensitive to local features, which is also reflected in bad weather. Even if the region proposal process can locate the object, bad weather will affect the appearance of the object. After the ROI set features are fed into the classifier, this effect may lead to misclassification of the category.
[0004] Through the above analysis, target detection in bad weather mainly faces three problems at present: lack of annotated images, loss of targets by two-stage detectors, and damage to target local features. To solve the above problems, it is necessary to improve the existing target detection and recognition methods in bad weather scenarios. Summary of the Invention
[0005] To solve the problems in object detection in adverse weather scenarios, such as the lack of annotated images, the loss of objects by two-stage detectors, and the damage to local features of objects, which lead to low object detection accuracy, the present invention provides an object detection method for adverse weather scenarios based on degradation consistency.
[0006] The present invention is implemented through the following technical solutions: An object detection method for adverse weather scenarios based on degradation consistency, which is implemented in an object detection network. The object detection network DCN includes a global consistency network GCN, a basic detection network BDN, and a local consistency network LCN. The specific steps are as follows:
[0007] 1) Images of good weather scenarios are represented by I n , including images of good weather in real scenarios and images of good weather synthesized based on computer graphics. Images of real adverse weather scenarios are represented by I i ; Since there are obvious visual differences between image I n and image I i , before feeding I n into the object detection network DCN, the global consistency network GCN (i.e., the style transfer network) is used to translate the style of I n into I i . The image generated by style transfer is represented by I ns , and I ns is the synthesized adverse weather image;
[0008] 2) Jointly train the basic detection network BDN and the local consistency network LCN: I ns and I i are the synthesized adverse weather image and the image of the real adverse weather scenario respectively; Image I ns is input into the object detection network DCN to generate the score map of the detector and optimize the detector; At the same time, the features generated by the detector are input into the discriminator of the DCN network to generate the prediction value of the discriminator;
[0009] 3) Image I i is input into the object detection network DCN, and the generated object features are directly input into the discriminator to generate a prediction value; The discriminator is used to distinguish whether the features come from image I ns or I i ;
[0010] 4) Repeat steps 2) and 3), and alternately feed I ns and I i into the object detection network DCN; When the performance verification index starts to decline or the number of training rounds reaches 1000 times, stop training;
[0011] 5) After the network training is completed, in the test phase, images in adverse weather scenarios are input into the target detector to obtain the detection categories and positions of the targets, and the experimental results are statistically analyzed and the metric scores including the evaluation accuracy AP for target recognition are calculated.
[0012] The present invention proposes a target detection method (DCN) for adverse weather scenarios based on degradation consistency. DCN consists of three parts: a Global Consistency Network (GCN), a Base Detection Network (BDN), and a Local Consistency Network (LCN). Synthetic data is used as a fully annotated data set and supplemented with a small number of adverse weather images in real scenarios to alleviate the scarcity of annotated data. GCN is an image-to-image style conversion technology, and its effectiveness has been thoroughly tested in the domain adaptation of target detection. DBN is a single-stage detector used to eliminate the adverse effects of RPN. LCN is a network derived from unpaired defogging or de-raining algorithms. However, LCN is different from traditional image enhancement algorithms: traditional image enhancement algorithms enhance images in the direction of improving image visibility, while LCN enhances images in the direction of reducing image visibility. In the experiment, DCN outperforms the state-of-the-art detection algorithms and the combination of BDN and the state-of-the-art enhancement algorithms.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The target detection method for adverse weather scenarios based on degradation consistency provided by the present invention uses an image-image style conversion algorithm in the GCN network to convert a large number of synthetic data sets into images in adverse weather scenarios, effectively alleviating the problem of insufficient image annotation in adverse weather scenarios. The BDN network uses an anchor-free target detector as the basic detector, thus avoiding the problem of target loss caused by the region mechanism in adverse weather scenario images. LCN is used to fine-tune the features generated by synthetic images to make them closer to the low-quality images in the original adverse weather scenarios in terms of details. The calculations of the three networks solve the problem of low target detection accuracy in actual adverse weather scenarios. Description of the Drawings
[0014] Figure 1 It is the network framework diagram of the present invention.
[0015] Figure 2 It is the target detection result diagram of the present invention in adverse weather scenarios. Detailed Embodiments
[0016] The present invention will be further described below in conjunction with specific embodiments.
[0017] A target detection method in bad weather scenarios based on degradation consistency, as follows Figure 1 shown: The target detection method is implemented in a target detection network. The target detection network DCN includes a global consistency network GCN, a basic detection network BDN, and a local consistency network LCN, and specifically includes the following steps:
[0018] 1) Images of good weather scenarios are represented by I n , including images of good weather in real scenarios and images of good weather synthesized based on computer graphics. Images of real bad weather scenarios are represented by I i ; Since there are obvious visual differences between image I n and image I i , before feeding I n into the target detection network DCN, the global consistency network GCN is used to translate the style of I n into I i . The image generated by style transfer is represented by I ns , and I ns is a synthesized bad weather image;
[0019] 2) Jointly train the basic detection network BDN and the local consistency network LCN: I ns and I i are the synthesized bad weather image and the image of the real bad weather scenario respectively; Image I ns is input into the target detection network DCN to generate a score map of the detector and optimize the detector; At the same time, the features generated by the detector are input into the discriminator of the DCN network to generate the prediction value of the discriminator; The description of the adaptation network loss for the entire target detection network DCN is as follows:
[0020] L(I n ,I i ) = L det (I ns ) + λ det I adv (I i )
[0021] where λ adv is used to balance and enhance the subtasks; The adversarial loss is L adv , and the binary cross-entropy loss is used.
[0022] 3) Image I i is input into the target detection network DCN, and the generated target features are directly input into the discriminator to generate a prediction value; The discriminator is used to distinguish whether the features come from image I ns or I i ;
[0023] 4) Repeat steps 2) and 3), and alternately feed I ns and I i into the target detection network DCN; when the performance verification index starts to decline or the number of training rounds reaches 1000 times, stop the training; I ns is used to optimize BDN; L det (I n ) is the loss function of BDN; the features created from the synthesized adverse weather images are called F n w,h,c ; the corresponding features created from the real adverse weather images are F i w ,h,c ; where w and h are the width and height of the feature map of the second-to-last layer of the head network; there are a total of 256 channels; the classification and regression sub-networks of each feature layer are exactly the same; the discriminator is alternately connected to F n w,h,c and F i w,h,c ; the discriminator loss function L adv can be expressed as:
[0024]
[0025] z = 0 indicates normal weather sensing; z = 1 indicates adverse weather sensing; if z is 0, the feature map of the obtained normal sensing image is fed into the discriminator;
[0026] After that, all discriminators optimize the following loss function:
[0027]
[0028] C and R represent the classification and regression sub-networks respectively; the multi-scale hierarchy of the target detection method is represented by L; Dis and Det represent the recognition and detection networks respectively.
[0029] 5) After the network training is completed, test the network; in the test stage, input the images in the adverse weather scenario into the target detector, obtain the detection categories and positions of the targets, and count the experimental results and calculate the index scores including the evaluation accuracy AP of target recognition.
[0030] The effects of the present invention can be further illustrated by the following simulation experiments of this embodiment.
[0031] 1. Simulation conditions
[0032] This embodiment conducts the simulation on the central processing unit of Intel(R) Xeon(R) CPU E5-2650 V4@2.20GHz, with 500G of memory and the Ubuntu 14 operating system, using Python and other related toolkits.
[0033] 2. Simulation Content
[0034] To verify the effectiveness of this embodiment, tests were conducted on the following datasets.
[0035] CIR: 2,285 rainy-day images were collected from the Internet, mostly from short video news. To be compatible with the SIM 10K target categories, vehicles were annotated, with a total of 8,795 cars. The ratio of the training set, validation set, and test set was 799:800:686.
[0036] RTTS: RTTS is a collection of foggy images containing a large number of targets. This dataset has 11,366 people, 2,590 buses, 1,232 motorcycles, and 698 bicycles. The ratio of the training, validation, and test sets was 1273:852:913.
[0037] SIM 10K: The SIM 10K dataset is a computer-graphics-based synthetic dataset. 10K and 200K are two capacity options; the label format is the same as PASCAL VOC. There is no test dataset because testing accuracy on synthetic datasets is meaningless.
[0038] Cityscapes: Cityscapes is a dataset of urban street scenes with semantic scene understanding. It mainly includes street scenes of 50 different cities, as well as 5,000 pixel-level annotated images of driving scenes in high-quality urban environments. People, riders, cars, trucks, buses, trains, motorcycles, and bicycles are the target categories.
[0039] In the experiment, first, it was compared with the state-of-the-art detector, and then with the combination of BDN and the enhancement algorithm. This evaluation form was also used in previous methods. In rainy and foggy weather, SIM 10K was used as I n while CIR and RTTS were used as I i .
[0040] (1) Comparison with the state-of-the-art detector
[0041] The accuracies under two kinds of adverse weather conditions are shown in Table 1. SIM 10K is used as the training set, and RTTS is used as the test set for the haze condition. Experiments are conducted on detectors FCOS, RetinaNet, EfficientNet, YOLOv4, SSD, and Faster R-CNN. ExtremeNet is the final best detector with an AP of 23.38. The AP of the proposed DCN is 27.81. In a foggy environment, the AP of DCN is 4.43 higher than that of ExtremeNet. SIM 10K is used as the training set, and CIR is used as the test set to study the rain scene. Experiments are also conducted on FCOS, RetinaNet, EfficientNet, YOLOv4, SSD, and Faster R-CNN. ExtremeNet is the final best detector with an AP of 20.38. The AP is 25.67. In the rain scene, the AP of DCN is 5.29 higher than that of ExtremeNet.
[0042] Table 1 Comparison of DCN with state-of-the-art object detectors under adverse weather
[0043]
[0044]
[0045] (2) Comparison with state-of-the-art enhancement algorithms
[0046] Table 2 shows the experimental results. The state-of-the-art DCN is compared with a combination of baseline detection and enhancement algorithms, including defogging and de-raining in rainy scenarios. In foggy weather scenarios, PMHLD, JSTASR, AOD, and DCPDN are compared. In rainy scenarios, PReNet, DID-MDN, and ID-CGAN are compared. Compared with the enhancement algorithms below FCOS, the experimental results show that DCN has the best detection effect.
[0047] Table 2 DCN and state-of-the-art enhancement algorithms
[0048]
[0049] The scope of protection required by the present invention is not limited to the above specific embodiments. Moreover, for those skilled in the art, the present invention can have various deformations and changes. Any modifications, improvements, and equivalent replacements made within the concept and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A target detection method in a bad weather scenario based on degradation consistency, characterized in that: The target detection method is implemented in a target detection network. The target detection network DCN includes a global consistency network GCN, a basic detection network BDN, and a local consistency network LCN, and specifically includes the following steps: 1) The image of a good weather scene is denoted by I n and includes the good weather images in real scenes and the good weather images synthesized based on computer graphics. The image of a real bad weather scene is denoted by I i ; Since there are obvious visual differences between the image I n and the image I i , before sending I n into the target detection network DCN, the global consistency network GCN is used to translate the style of I n into I i . The image generated by style transfer is denoted by I ns , and I ns is the synthesized bad weather image; 2) Jointly train the basic detection network BDN and the local consistency network LCN: Image I ns is input into the object detection network DCN to generate the score map of the detector and optimize the detector; meanwhile, the features generated by the detector are input into the discriminator of the object detection network DCN to generate the prediction value of the discriminator; It also includes a description of the adaptation network loss for the entire target detection network DCN: L(I n ,I i ) = L det (I ns ) + λ adv L adv (I i ) where λ adv represents the balance coefficient of the adversarial loss function; L adv (I i ) represents the adversarial loss function, using binary cross-entropy loss; L det (I ns ) is the detection loss function of the BDN; 3) Image I i is input into the object detection network DCN, and the generated object features are directly input into the discriminator to generate prediction values; the discriminator is used to distinguish whether the features come from Image I ns or I i ; 4) Repeat steps 2) and 3), and alternately feed I ns and I i into the object detection network DCN; When the performance verification metric starts to decline or the number of training rounds reaches 1000 times, stop training; Specifically: I ns Used to optimize BDN; the features created from the synthesized adverse weather images are called F n w,h,c ; the corresponding features created from the real adverse weather images are F i w,h,c ; where w and h are the width and height of the feature map of the penultimate layer of the head network; c represents the number of channels of the feature map, with a total of 256 channels; the classification and regression sub-networks of each feature layer are exactly the same; the discriminator is alternately connected to F n w,h,c and F i w,h,c ; the adversarial loss function L adv (I i ) is expressed as: z = 0 indicates normal weather sensing; z = 1 indicates bad weather sensing; if z is 0, the feature map of the obtained normal sensing image is sent into the discriminator; After that, all discriminators optimize the following loss function: C and R respectively represent the classification and regression sub-networks; L represents the multi-scale level of the target detection method; Dis and Det respectively represent the recognition and detection networks; 5) After the network training is completed, test the network; In the test phase, input the images in the bad weather scenario into the target detector to obtain the detection categories and positions of the targets, count the experimental results, and calculate the metric scores including the evaluation accuracy AP of target recognition.
Citation Information
Patent Citations
Small sample license plate detection method in severe weather
CN115050028A
Method and apparatus for image enhancement
WO2022174908A1