Visual odometer system for automatic driving in foggy days
By introducing a combination solution of graphics processing module, defogging network model and LEAP-VO module in the visual odometer system, the problems of image quality decline and trajectory drift in foggy conditions are solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510344509.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional visual odometer technology reduces image quality in foggy conditions, resulting in limited accuracy of feature extraction and matching, affecting the stability of pose estimation, and errors are prone to gradually accumulate, leading to aggravation of trajectory drift problems.
The combination scheme of graphics processing module, defogging network model and LEAP-VO module is adopted, including image preprocessing, feature extraction, multi-stage defogging and dynamic trajectory estimation, ensuring the generation of clear and semantic fogging-free images under foggy weather conditions, and improving the robustness and accuracy of the visual odometer.
Through the joint solution of the defogging network model and the LEAP-VO module, the system provides more accurate and reliable pose estimation in foggy days and dynamic scenarios, improving the quality and stability of visual perception.
Smart Images

Figure CN120219847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving vision navigation, and specifically to a vision odometer system for autonomous driving in foggy weather. Background Art
[0002] With the rapid development of autonomous driving vision navigation technology, visual odometry has gradually become one of the important methods to solve the problems of positioning and navigation. Although traditional visual odometry technology performs well in specific application scenarios, its adaptability is limited in dynamic environments and foggy weather conditions. In addition, existing depth dehazing algorithms usually do not consider the consistency of objects in the image before and after dehazing, and it is difficult to ensure the stability of image semantics during the dehazing process.
[0003] Specifically, traditional visual odometry usually relies on feature point matching in static scenes to estimate relative motion. However, in foggy weather conditions, the image quality significantly deteriorates, which limits the accuracy of feature extraction and matching, and seriously affects the stability of pose estimation. In addition, due to the fact that visual odometry algorithms rely on the cumulative calculation of relative motion between frames, errors are prone to gradually accumulate, leading to an exacerbation of the trajectory drift problem. Therefore, in order to improve the robustness and accuracy of visual odometry in adverse weather conditions such as fog, improving the dehazing effect has become a key research direction. At the same time, enhancing the feature tracking ability in dynamic environments is also the current research focus to ensure stability and positioning accuracy in complex environments. Summary of the Invention
[0004] In view of the above problems, the present invention provides a vision odometer system for autonomous driving in foggy weather.
[0005] The technical solution adopted by the present invention to solve its technical problems is:
[0006] A vision odometer system for autonomous driving in foggy weather, comprising
[0007] A graphics processing module: including an image preprocessing unit, a feature extraction unit, a threshold judgment unit, and a comprehensive judgment unit, for identifying whether the image is a foggy image;
[0008] A dehazing network model: including a generator G, a generator F, and a discriminator D, using the generator G and the generator F for multi-stage dehazing to generate a clear and semantically consistent fog-free image, and optimizing the dehazing effect through the discriminator D;
[0009] A LEAP-VO module: this module includes feature point extraction, key point trajectory generation, trajectory filtering, and local bundle adjustment, and uses a sliding window method for bundle adjustment to continuously optimize the pose estimation of the camera within a local range.
[0010] Further, the image preprocessing unit first grayscales the color image, converts it into a grayscale image to reduce the data dimension, and applies mean filtering to the grayscale image.
[0011] Further, the feature extraction unit converts the image to the HSV color space, calculates the mean value of its saturation component, and determines that an image with a low saturation value is a foggy scene.
[0012] Further, the generator G includes a multi-stage progressive image restoration network module and a densely connected pyramid defogging network module.
[0013] Further, the discriminator D is used to determine whether the defogged image output by the generator reaches the expected defogging effect.
[0014] Further, the discriminator D receives the generated defogged image as input, uses a binary classification model to determine whether the image is successfully defogged, the discriminator D learns the difference between the foggy image and the defogged image, and outputs a binary label, where 0 represents that there is still fog in the image, and 1 represents that the image has been successfully defogged.
[0015] Further, the LEAP-VO module extracts key points frame by frame in a given image sequence, constructs the position information of these key points into a continuous trajectory, and thereby performs a more stable spatial estimation of the scene.
[0016] The beneficial effects of the present invention are:
[0017] The present invention introduces a generative adversarial network, making the defogged image clearer than the original image. At the same time, it ensures that the defogged image has the same semantics as the original image. Then, a dynamic trajectory estimation method for long-term valid points is used to improve the adaptability of the system in dynamic scenarios. Through the joint scheme of the defogging network model and dynamic tracking, the system will be able to provide more accurate and reliable pose estimation in dynamic and foggy weather environments, thus better meeting the high requirements of autonomous driving and robot navigation systems for visual perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the system principle flowchart of the present invention;
[0019] Figure 2 is the flowchart for judging an image of the present invention;
[0020] Figure 3 is the schematic diagram of the defogging network structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To better understand the present invention, the following is combined with Figures 1-3The technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0022] The vision odometry system for autonomous driving in foggy days of the present invention includes a graphics processing module, a defogging network model, and a LEAP-VO module. The following will respectively elaborate on these three modules in detail.
[0023] As Figure 1 shown, this project designs a vision odometry system with defogging detection and defogging capabilities to improve its vision odometry performance in foggy and dynamic scenarios. First, through image preprocessing, feature extraction, threshold judgment, and comprehensive judgment, it is identified whether the image is a foggy image, and the foggy features (such as brightness, contrast, saturation) are analyzed. The defogging network model adopts a multi-stage defogging structure, including a generator F based on CycleGAN, a dense connection pyramid defogging network DCPDN, and a multi-stage progressive image restoration network MPRNet, to generate clear and semantically consistent fog-free images, and optimizes the defogging effect through a discriminator D to ensure the quality of the generated images. In addition, the LEAP-VO module robustly estimates the pose of the camera in dynamic scenarios through feature extraction, feature encoding, and hierarchical sampling strategies, realizes high-precision trajectory estimation, thereby improving the adaptability and robustness of the system in dynamic environments.
[0024] The graphics processing module includes an image preprocessing unit, a feature extraction unit, a threshold judgment unit, and a comprehensive judgment unit to identify whether the image is a foggy image. To achieve the research goal of the vision odometry system, this project uses the SOTS (Synthetic Object Training Set) in the RESIDE dataset as the preliminary training dataset. SOTS provides paired fog-free and foggy images, and first uses this dataset to train the defogging network. The goal of this stage is to make the model perform well in the basic effect of removing haze. During the training process, the generator and discriminator of the generative adversarial network (GAN) are used to optimize the defogging effect to ensure that the image after defogging is not only clear but also retains the semantic consistency of the original image.
[0025] The defogging network trained by SOTS can have certain defogging capabilities. However, due to the different foggy characteristics in actual application scenarios, the model may show certain deviations on real data. At this stage, the real foggy dataset collected by the trolley is used to fine-tune the weights of the model to enhance its adaptability to the real foggy environment.
[0026] Judging whether an image is a foggy image is carried out through multiple steps, including image preprocessing, feature extraction, threshold judgment, and comprehensive judgment. The overall method is as Figure 2 shown.
[0027] (1) Image preprocessing
[0028] First, the color image is grayscale converted to reduce the data dimension. Subsequently, mean filtering is applied to the grayscale image to reduce noise interference, making the overall features clearer and more prominent.
[0029] (2) Feature extraction
[0030] In a foggy scene, an image usually exhibits the characteristics of high brightness and low contrast. Therefore, the foggy weather features can be identified by calculating the average brightness and contrast of the image. Specifically, the brightness mean in foggy weather is higher and the fluctuation is smaller, while the contrast tends to be low, and the kurtosis and width of the histogram are concentrated. In addition, the color saturation of a foggy image is also low. The image can be converted to the HSV color space, and the mean value of its saturation component is calculated. Usually, an image with a lower saturation value is more likely to be a foggy scene.
[0031] (3) Threshold judgment
[0032] A reasonable threshold can be set based on the mean value or distribution of brightness, contrast, and saturation features. If the image has high brightness, low contrast and saturation, and meets the preset threshold conditions, then the image can be judged as a foggy image.
[0033] (4) Comprehensive judgment
[0034] By combining the analysis results of brightness, contrast, and saturation, multi-level feature fusion can be carried out to judge whether the image is a foggy image. If the analysis results meet the preset foggy conditions, the output is "foggy image"; otherwise, the output is "non-foggy image".
[0035] Through the above steps, it is possible to effectively judge whether an image is a foggy image, which is particularly suitable for improving image quality and recognition accuracy in environmental perception tasks such as autonomous driving.
[0036] Defogging network model: It includes generator G, generator F, and discriminator D. The multi-stage defogging of generator G and generator F is adopted to generate clear and semantically consistent fog-free images, and the defogging effect is optimized by discriminator D. The overall architecture is as Figure 3As shown. This project is based on a depth dehazing algorithm with an invariant object, combined with a generative adversarial network (GAN) architecture to ensure that the dehazed image not only has high clarity but also maintains semantic consistency. Specifically, the model uses a dense connection pyramid dehazing network (DCPDN) and a multi-stage progressive image restoration network (MPRNet) to achieve multi-stage dehazing processing, generating a haze-free image through layer-by-layer feature extraction and progressive restoration. During the training process, evaluation metrics such as contrast and detail retention are introduced to optimize the dehazing effect and ensure that the generated image is more visually suitable for the visual odometry task. At the same time, the generator F in the model is based on the CycleGAN architecture, combined with the multi-stage dehazing method of DCPDN and MPRNet, to achieve layer-by-layer feature extraction and image restoration, thus ensuring the clarity and semantic consistency of the dehazed image. The discriminator D is used to evaluate whether the generated haze-free image has achieved the expected dehazing effect. Its input is the generated haze-free image, and the output is the label of the success or failure of the dehazing effect, thereby further improving the dehazing performance and visual consistency of the system.
[0037] (1) Input image and depth feature extraction
[0038] In this project, the input is a set of dynamic scene images with haze. The dehazing network model will first process the input RGB image to extract the depth features in the image. Through the feature extractor of the deep neural network, multi-layer convolutions are performed on the feature information in the image to obtain feature maps regarding different scales and haze distributions. These feature maps contain the spatial information of the image and the preliminary dehazing effect, and will be used for further dehazing processing and feature enhancement.
[0039] (2) Generator G
[0040] The generator G of the dehazing network model consists of two main modules. First is the multi-stage progressive image restoration network (MPRNet), which performs preliminary dehazing processing on the input hazy image (the input image is X1). Next is the dense connection pyramid dehazing network (DCPDN), which further refines the dehazing process, gradually removing the residual haze in the image and enabling more detailed processing of the image at different scales.
[0041] (3) Generator F
[0042] The generator F adopts the structure of CycleGAN. CycleGAN is an unsupervised image translation model commonly used for mapping between two different domains. Its generator adopts an encoder-decoder architecture, extracts features through convolutional layers, and uses residual blocks for processing to ensure content consistency during image translation. The core of CycleGAN lies in ensuring that the image mapped from the original domain to the target domain, when mapped back to the original domain, is consistent with the input image, thereby effectively retaining the object and semantic information of the defogged image. In the defogging task, the generator F utilizes this cycle consistency to maintain the stability and consistency of objects in the images before and after defogging.
[0043] Discriminator D
[0044] The main function of the discriminator D is to determine whether the fog-free image output by the generator achieves the expected defogging effect. The discriminator D receives the generated fog-free image as input and uses a binary classification model to determine whether the image has been successfully defogged. Specifically, the discriminator D learns the differences between foggy images and fog-free images and outputs a binary label, where "0" represents that the image still has fog, and "1" represents that the image has been successfully defogged. By introducing the discriminator D, the model can more effectively optimize the defogging effect and ensure that the generated fog-free image is visually clearer and reliable for practical applications.
[0045] LEAP-VO module: This module includes feature point extraction, key point trajectory generation, trajectory filtering, and local bundle adjustment, and uses a sliding window method for bundle adjustment to continuously optimize the camera pose estimation within a local range. Based on the clear image after defogging, a visual odometry system is constructed, and the LEAP-VO (Long-term Effective Any PointTracking for Visual Odometry) method is used to enhance the point tracking ability of the system in dynamic scenes. LEAP-VO uses an anchor-assisted dynamic trajectory estimation model to improve the robustness and uncertainty evaluation of feature points by combining spatial and temporal information, reducing the impact of dynamic objects on the VO accuracy. And optimize the pose estimation of the system in dynamic scenes. By comparing the predicted trajectory and the true trajectory frame by frame, optimize the performance of the system in dynamic scenes and reduce error accumulation. Introduce a probability model to further improve the reliability of feature points in occluded or dynamic environments.
[0046] (1) Feature extraction and trajectory estimation
[0047] In the LEAP-VO module, key-point features are first extracted from the RGB image sequence, and the trajectories of these key points are tracked. The extracted features are used to estimate the motion between images. By combining the optimized trajectory data, more accurate pose estimation is obtained. Specifically, LEAP-VO extracts key points frame by frame in the given image sequence and constructs the position information of these key points into continuous trajectories to perform a more stable spatial estimation of the scene.
[0048] In this process, the trajectory of each key point represents the moving path of the camera during that time period. By fusing more feature information (such as direction, speed, etc.) in the trajectory, the motion trend of the camera can be better captured, thereby further improving the accuracy of pose estimation. To improve the computational efficiency, the LEAP-VO module adopts a sliding window approach, only considering a part of the image sequence each time instead of all the image data to reduce the computational amount.
[0049] (2) Feature Encoding
[0050] Although the LEAP-VO module can estimate the camera motion well by tracking the key-point trajectories, directly using the original feature point information for estimation is likely to lead to inaccurate results, especially in complex scenes. Therefore, to better capture the detailed changes in the scene and improve the estimation accuracy, we introduce a feature encoding method to encode the key-point information into a higher-dimensional feature vector
[24] .
[0051] The feature encoding process maps each key-point position to a higher-frequency feature space, thus better capturing complex scene information. The encoded information contains multiple features of space and direction, enabling the LEAP-VO module to more accurately fit the data containing high-frequency changes when performing motion estimation. The encoded feature information will be fed into the deep learning network so that the network can learn more complex spatial representations.
[0052] (3) Hierarchical Sampling and Optimization
[0053] A large number of samplings are required in the calculation process of LEAP-VO, and most of these sampling points do not substantially contribute to the final estimation result (such as some background noise or occluded areas). Therefore, to reduce the computational overhead and improve the efficiency, a hierarchical sampling strategy is introduced. This strategy adopts a "coarse-to-fine" structure, first performing sparse sampling on the coarse network, and then performing fine sampling based on the coarse sampling to further improve the accuracy of sampling.
[0054] First, the coarse network sparsely samples the trajectories between image frames to roughly estimate the overall trend of motion. On this basis, inverse transform sampling is performed based on the sampled distribution to obtain more fine sampling points, and then the trajectory optimization is carried out on the fine sampling points, and finally the fine estimation result of LEAP-VO is obtained.
[0055] Through multi-level voxel sampling and fine trajectory optimization, the LEAP-VO module can improve the accuracy of pose estimation while reducing ineffective calculations.
[0056] It should be noted that although this specification is described according to the implementation manners, not every implementation manner only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.
Claims
1. A visual odometer system for autonomous driving in foggy weather, characterized in that: include Graphics processing module: including image preprocessing unit, feature extraction unit, threshold judgment unit and comprehensive judgment unit, to identify whether the image is foggy; Defogging network model: It includes generator G, generator F and discriminator D. Generator G and generator F are used for multi-stage defogging to generate clear and semantically consistent defogging images, and the defogging effect is optimized through discriminator D. LEAP-VO module: This module includes feature point extraction, key point trajectory generation, trajectory filtering and local bundle adjustment, and uses a sliding window to perform bundle adjustment to continuously optimize the camera's pose estimation in a local range.
2. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The image preprocessing unit first grayscales the color image, converts it into a grayscale image to reduce the data dimension, and applies a mean filter to the grayscale image.
3. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The feature extraction unit converts the image into the HSV color space, calculates the mean value of the saturation component thereof, and determines the image with a lower saturation value as a foggy scene.
4. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The generator G includes a multi-stage progressive image restoration network module and a densely connected pyramid defogging network module.
5. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The discriminator D is used to determine whether the haze-free image output by the generator achieves the expected dehazing effect.
6. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The discriminator D receives the generated fog-free image as input, and determines whether the image is successfully defogged through a binary classification model. The discriminator D learns the difference between the foggy image and the fog-free image, and outputs a binary label, where 0 represents that the image still has fog, and 1 represents that the image has been successfully defogged.
7. The visual odometer system for autonomous driving in foggy weather as claimed in claim 1, characterized in that: The LEAP-VO module extracts key points frame by frame in a given image sequence and constructs the position information of these key points into continuous trajectories to perform a more stable spatial estimation of the scene.