A real-time obstacle detection method in a full foggy day scenario
By using the Transformer classification model to classify foggy days scenes and selecting appropriate detection models according to the level, the problems of poor obstacle detection effect and poor real-time performance in all foggy days scenes are solved, and more efficient obstacle detection accuracy and real-time performance in foggy days scenes are achieved.
Patent Information
- Application Number
- CN202210855171.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-07-19
AI Technical Summary
The prior art has poor obstacle detection effect and poor real-time performance in all foggy scenes, which cannot meet the real-time requirements of computer vision target obstacle detection algorithms.
The Transformer classification model is used to classify foggy scenes of different levels, and appropriate detection models are selected according to the foggy day level. For example, using the YOLOv5 detection model in light foggy scenes, multi-scale fusion CycleGAN network is used to generate foggy day image data sets in medium foggy scenes, and using infrared detection network for obstacle detection in dense foggy scenes.
The adaptability and reliability of the computer visual obstacle detection method in foggy scenes is improved, and more efficient obstacle detection accuracy and real-time performance in foggy scenes is achieved.
Smart Images

Figure CN115240069B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to a fog detection method in image processing technology and multi-sensor applications, and specifically is a real-time obstacle detection method in a full foggy scene. Background Art
[0002] Currently, in recent years, computer vision recognition has become a research hotspot in the fields of target recognition and target tracking, and is widely applied in multiple fields such as autonomous driving, medical image segmentation, human-computer interaction, and robot vision. When actually working outdoors with the aid of computer vision, it is inevitable to encounter some foggy weather, which reduces the effect of visual recognition and causes the corresponding recognition work to be unable to proceed normally.
[0003] In a foggy scene, local information of the image is lost, which not only affects the recognition of human eye vision, but also increases the recognition difficulty of computer vision. The accuracy of computer vision target recognition is determined by the quality of the image, and foggy weather has an adverse impact on image recognition, restricting the development of computer vision. Currently, significant progress has been made in target detection based on computer vision, which is widely applied in outdoor application scenarios, but there is a lack of a reasonable method in the field of computer vision to deal with obstacle detection in foggy scenes.
[0004] Currently, the solution for a foggy scene is to use a defogging algorithm to correct the noise of the foggy image, and use a detection algorithm to detect the obstacle target for the defogged image.
[0005] The image defogging algorithms in a foggy scene are mainly divided into three categories:
[0006] 1) Defogging algorithms based on image contrast enhancement. The algorithms based on contrast enhancement are early image enhancement image processing methods, which improve the quality of foggy images by adjusting the difference between the target and environmental information, mainly including methods such as histogram equalization, wavelet transform, homomorphic filtering, log transform, power transform, and gamma correction. These methods rely on information such as contours, can highlight the color difference of the target, and then determine the position and contour information of the target, and have a certain effect on target recognition in a foggy scene, but are prone to losing the target information in the image and have poor adaptability in a foggy scene.
[0007] 2) Image restoration algorithm enhanced based on physical models. The physical model defogging algorithm uses physical model simulation of the principle of natural foggy scene generation to determine environmental information and perform reverse inference on the defogged image to obtain a clear image. In the physical model of foggy scene simulation, the atmospheric scattering model is mainly used. Using the atmospheric scattering model, a clear image can be obtained based on information such as atmospheric illumination and distance. In 1975, McCartney et al. proposed the original atmospheric scattering model. Based on the atmospheric scattering model, Narasimhan proposed an atmospheric scattering model for foggy scenes, constructed a physical model for generating foggy images in foggy scenes, and laid an important physical model foundation for subsequent image defogging and other work.
[0008] 3) Defogging algorithm based on deep learning. The defogging algorithm based on deep learning does not rely on contrast enhancement and physical models. It trains a defogging network model through a large number of foggy scene images and clear scene images to achieve the effect of image defogging. The deep learning defogging algorithm is not limited to a specific physical model. It uses a large amount of data to build a foggy scene model and effectively performs deep defogging on foggy images to achieve a better defogging effect. The defogging algorithm for foggy scenes generally uses the Unet network architecture to extract effective features in the foggy image through Encode and construct the defogged image through Decode. The generated image not only retains the features in the original image but also reduces the interference of foggy information.
[0009] For the above three defogging method models, the deep learning defogging network shows better defogging effects in its own test environment dataset. However, due to the influence of the randomness and variability factors of the real foggy environment, the complexity of the actual working conditions in the foggy environment is greatly affected, which significantly affects the defogging effect of the deep learning defogging algorithm. The existing computer vision obstacle detection methods for foggy scenes more rely on defogging methods. Since the defogging methods have poor applicability in actual scenes, it increases the difficulty of model deployment and cannot meet the real-time requirements of the computer vision target obstacle detection algorithm. In addition to the problems existing in the defogging algorithm itself, a more notable problem is the multi-level nature of the foggy scene. For different levels of foggy scenes, the applicability of the solution must be more reliable and stable.
[0010] After using the defogging method to perform defogging processing on the image, the detection algorithm is then used for obstacle target detection. Currently, target detection algorithms are mainly divided into two categories: one is the target detection algorithm based on target candidate regions, namely the two-stage detection method, and typical algorithms include Fast R-CNN, Faster R-CNN, R-FCN, etc.; the other is the target detection algorithm based on regression, namely the one-stage detection method, and typical algorithms include the YOLO series, SSD, RetinaNet, etc. Summary of the Invention
[0011] The object of the present invention is to provide a real-time obstacle detection method in a full foggy day scenario, so as to solve the problems of poor obstacle detection effect and poor real-time performance in the existing technology in the full foggy day scenario, and improve the adaptability and reliability of the computer vision obstacle detection method in the foggy day scenario.
[0012] To achieve the above object, the technical solution of the present invention is as follows: A real-time obstacle detection method in a full foggy day scenario, including the following steps:
[0013] A. Establish a Transformer classification model
[0014] A1. Use a fog visibility detection device and a vision sensor to collect image information of different levels of fog concentration, and classify the vision foggy day scenario according to horizontal visibility:
[0015] 1000m ≤ visibility < 10000m is light fog;
[0016] 500m ≤ visibility < 1000m is moderate fog;
[0017] Visibility < 500m is thick fog;
[0018] A2. After completing the collection of classification data sets for light fog, moderate fog and thick fog, use foggy day images of different levels to train the Transformer classification network;
[0019] B. Detect obstacle targets in foggy day scenarios of different concentration levels
[0020] If it is a light foggy day scenario, go to step B1; if it is a moderate foggy day scenario, go to step B2; if it is a thick foggy day scenario, go to step B3;
[0021] B1. For a light foggy day scenario, use a vision sensor to collect image information, use the YOLOv5 detection model, and use the ImageNet pre-trained weights to detect obstacle targets; go to step C;
[0022] B2. For a moderate foggy day scenario, use a vision sensor to collect image information, use the multi-scale fusion CycleGAN network, namely DF-CycleGAN, to generate a foggy day image data set, and use the mixed data set of clear images and foggy day images to train the YOLOv5 detection model. Finally, perform obstacle detection in the foggy day scenario; go to step C;
[0023] B3. For a thick foggy day scenario, use an infrared sensor to collect infrared image information, use the infrared images to train the YOLOv5 detection network, and perform obstacle target detection in the foggy day scenario;
[0024] C. Output the obstacle detection target result in step B.
[0025] Further, the method for generating the foggy image dataset by the multi-scale fusion CycleGAN network in step B2 includes the following steps:
[0026] B21. Enhance the Encode structure in the generator Unet of the original CycleGAN network in the scale direction, and use the gap between the upper-layer feature map and the lower-layer feature map to make up for the lost feature information of the lower-layer feature map. The calculation formula is as follows:
[0027]
[0028] J n = D n (J n-1 )
[0029] In the formula, J n represents the enhanced feature information from the nth layer of the decoder; represents the feature information after feature fusion; represents the fusion module of the nth layer, which is composed of N residual modules; D n is the downsampling method of the nth layer, which is composed of a 3×3 convolution with a stride of 2, ↓ represents downsampling by 2 times; ↑ represents upsampling by 2 times, using the method of transposed convolution;
[0030] B22. Use an enhancement module for the Decode structure in Unet, and use multiple residual structures to fuse the Encode information and the Decode information. The calculation formula is as follows:
[0031]
[0032] In the formula, represents the enhanced feature information of the nth layer of Decode; represents the enhanced feature information of the ith layer of Encode. n and i correspond to each other. The size of the nth layer feature map in Decode is equal to the size of the ith layer feature map in Decode, and is equal to 2 times the size of the (i - 1)th layer feature map in Decode; F n represents the enhancement module of the nth layer, which is composed of N residual modules; ↑ represents upsampling by 2 times, using the bilinear interpolation method;
[0033] B23. Use the G structure at the connection of Encode and Decode to extract information features from M residual modules.
[0034] Further, in step B23, N = 10 and M = 20.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention uses a Transformer model to construct a foggy weather scene classification network at different levels, which can effectively classify the fog levels in foggy weather scenes and provide corresponding real-time solutions for the usage requirements of different foggy weather scenes.
[0037] 2. The present invention conducts research on medium fog weather scenes, generates a realistic foggy weather scene dataset using a multi-scale fusion DF-CycleGAN network, and trains an object detection network with a mixture of clear images and foggy images, which can effectively improve the adaptability to foggy weather scenes.
[0038] 3. The present invention conducts research on thick fog weather scenes, uses the advantages of infrared data in foggy weather scenes to train an infrared detection network and detect obstacles, which can effectively improve the detection effect in thick fog scenes and is a necessary process to ensure the feasibility of the full foggy weather scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a schematic flowchart of the present invention.
[0040] Figure 2 is a schematic flowchart of the multi-scale fusion module of the present invention.
[0041] Figure 3 is a schematic flowchart of the DF-CycleGAN network structure of the present invention.
[0042] Figure 4 is a schematic flowchart of the obstacle detection by the infrared sensor of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the advantages of the present invention from the content disclosed in this specification. The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. To provide a deep understanding of the present invention, many specific details will be included in the following description. The present invention can also be implemented without these details. In addition, some specific details will be omitted in the description to avoid confusing or obscuring the key points of the present invention.
[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the implementation manners of the present invention will be further described in detail below with reference to the accompanying drawings:
[0045] Figure 1It is an overall obstacle detection framework designed by the present invention for foggy scenarios. By using a Transformer classification network with better detection performance, it can effectively consider the global information of foggy images and then accurately predict the fog level. For different levels of foggy scenarios, using different obstacle detection schemes can effectively improve the obstacle detection accuracy in foggy scenarios: in light fog scenarios, the image clarity is relatively high, the visibility is greater than 1 kilometer, and the visual sensor can effectively collect target information. Using the YOLOv5 detection network with good real-time performance and detection effect can complete the obstacle target detection. For thick fog scenarios, the visual sensor cannot effectively collect target information, which causes great trouble to obstacle detection and is difficult to effectively meet the detection requirements in foggy weather scenarios. However, the infrared sensor is less affected by foggy scenarios and has a similar imaging effect to the visual sensor, and can effectively identify obstacle targets. Therefore, based on the images collected by the infrared sensor, using the YOLOv5 detection network can meet the requirements of foggy obstacle detection. Figure 4 It is the overall application process of the infrared sensor; for medium fog scenarios, the visual images are affected to a certain extent, mainly due to the white fog noise in the collected images. Since in most scenarios, special data needs to be collected from visual images, the infrared sensor cannot directly replace the visual sensor, and the cooperation of the two will lead to the synchronous use of multiple sensors, resulting in waste of multi-device resources. Therefore, the present invention only performs obstacle target detection in medium fog scenarios based on the visual sensor, and uses simulated foggy scenario images to train the detection network, which can not only meet the obstacle detection effect of the detected images but also ensure the real-time performance of foggy scenarios. The generation of simulated images uses the DF-CycleGAN network ( Figure 2 and Figure 3 are the multi-scale fusion module and the DF-CycleGAN network structure respectively). Compared with the CycleGAN network, the DF-CycleGAN network conducts deeper information interaction in the generator part, can effectively extract more image feature information, generate more realistic and reliable foggy scenario images, and use the simulated foggy images to train the YOLOv5 detection network, which can enhance the adaptability of the YOLOv5 detection network in foggy scenarios and ensure the real-time performance of the network, meeting the obstacle detection requirements in actual foggy scenarios.
[0046] The above describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A real-time obstacle detection method in a full foggy day scenario, characterized in that: It includes the following steps: A. Establish a Transformer classification model A1. Use a fog visibility detection device and a vision sensor to collect image information of fog concentrations at different levels, and classify the vision foggy day scenarios according to horizontal visibility: 1000m ≤ visibility < 10000m is light fog; 500m ≤ visibility < 1000m is moderate fog; visibility < 500m is thick fog; A2. After completing the collection of classification datasets for light fog, moderate fog, and thick fog, use foggy day images at different levels to train the Transformer classification network; B. Detect obstacle targets in foggy day scenarios with different concentration levels If it is a light foggy day scenario, go to step B1; If it is a moderate foggy day scenario, go to step B2; If it is a thick foggy day scenario, go to step B3; B1. For a light foggy day scenario, use a vision sensor to collect image information, use the YOLOv5 detection model, and use the ImageNet pre-trained weights to detect obstacle targets; Go to step C; B2. For a moderate foggy day scenario, use a vision sensor to collect image information, use the multi-scale fusion CycleGAN network, namely DF-CycleGAN, to generate a foggy day image dataset, and use the mixed dataset of clear images and foggy day images to train the YOLOv5 detection model. Finally, perform obstacle detection in the foggy day scenario; Go to step C; B3. For a thick foggy day scenario, use an infrared sensor to collect infrared image information, use the infrared images to train the YOLOv5 detection network, and perform obstacle target detection in the foggy day scenario; C. Output the obstacle detection target results in step B.
2. The real-time obstacle detection method in a full foggy day scenario according to claim 1, characterized in that: The method for the multi-scale fusion CycleGAN network to generate a foggy day image dataset in step B2 includes the following steps: B21. Enhance the Encode structure in the generator Unet of the original CycleGAN network in the scale direction, and use the gap between the upper feature map and the lower feature map to make up for the lost feature information of the lower feature map. The calculation formula is as follows: J n = D n (J n-1 ) Where J n represents the enhanced feature information from the n-th layer of the decoder; represents the feature information after feature fusion; represents the fusion module of the n-th layer, which is composed of N residual modules; D n is the downsampling method of the n-th layer, which is composed of a 3×3 convolution with a stride of 2. ↓ represents downsampling by a factor of 2; ↑ represents upsampling by a factor of 2 using the method of transposed convolution; B22. Use an enhancement module for the Decode structure in Unet, and use multiple residual structures to fuse the Encode information and the Decode information. The calculation formula is as follows: In the formula, represents the enhanced feature information of the n-th layer of Decode; represents the enhanced feature information of the i-th layer of Encode. n and i correspond to each other. The size of the feature map of the n-th layer in Decode is equal to the size of the feature map of the i-th layer in Decode, and is equal to 2 times the size of the feature map of the (i - 1)-th layer in Decode; F n represents the enhancement module of the n-th layer, which is composed of N residual modules; ↑ represents upsampling by 2 times, using the method of bilinear interpolation; B23. Use a G structure at the connection of Encode and Decode to extract information features of M residual modules.
3. The real-time obstacle detection method in a full foggy day scenario according to claim 2, characterized in that: In step B23, N = 10 and M = 20.
Citation Information
Patent Citations
Vehicle detection method for intelligent vehicle under severe weather conditions
CN111369541A
Lightweight night infrared image pedestrian detection method and system based on model optimization
CN113313078A