Cross-illumination image feature matching method and system for night visual positioning

Through monocular depth estimation and illumination re-rendering technology, night images are upgraded to three-dimensional point cloud space, and a two-stage feature matching network is designed to solve the problem of low nighttime visual positioning accuracy and achieve high-precision matching under different lighting conditions. It is suitable for scenarios such as unmanned vehicles, security monitoring, AR equipment and mobile robots.

CN120655946APending Publication Date: 2025-09-16BEIJING INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510753098.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing visual positioning methods suffer from low positioning accuracy and unstable image matching at night or in low-light environments. In particular, it is difficult to achieve accurate matching with daytime mapping results under conditions of significant lighting changes.

Method used

Through monocular depth estimation, night images are upgraded to three-dimensional point cloud space. Light re-rendering technology is combined to generate training data across lighting conditions. A two-stage feature matching network is designed to achieve high-precision alignment of night images and daytime maps.

Benefits of technology

Without the need for multi-view data or external sensors, the robustness and accuracy of nighttime visual positioning are improved, and it has good matching capabilities across lighting conditions, making it suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655946A_ABST
    Figure CN120655946A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-illumination image feature matching method and system for night visual positioning, and belongs to the field of computer vision. For matching failure caused by image appearance difference at night or under low illumination, an image enhancement method combining monocular depth estimation, three-dimensional point cloud reconstruction, view angle synthesis and illumination modeling is provided. According to the method, a dense depth map is generated through a single image, a point cloud is reconstructed, a new visual angle image is synthesized, occlusion completion and style conversion are carried out, and a cross-illumination image pair corresponding to a pixel level is constructed and used for training a feature encoder and a matching decoder with three-dimensional perception capability. The proposed two-stage neural network has geometric consistency and cross-illumination robustness, self-supervised training can be carried out without annotation data, and high-precision matching and positioning of day and night images can be realized in application. The method does not need multiple view angles, has the advantages of low cost, high adaptability and high precision, and is suitable for night positioning scenes such as automatic driving, robot navigation and security monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and intelligent perception, specifically a method and system for cross-illumination feature matching for nighttime visual positioning. More specifically, the present invention utilizes monocular depth estimation, re-rendering, and novel perspective image synthesis techniques to achieve highly robust visual matching between nighttime images and daytime mapping results without relying on multi-view images. This relates to the fields of visual positioning and image matching. Background Art

[0002] Visual localization is a key technology in fields such as autonomous driving, robotic navigation, and augmented reality. Its core is to estimate the vehicle's position by matching the current image with images in a known map of the environment. In most applications, maps are typically constructed under well-lit daytime conditions. Therefore, when devices operate at night or in low-light environments, differences in image appearance caused by varying lighting can severely impact the robustness and accuracy of image matching.

[0003] Traditional visual localization methods rely on hand-crafted image features (such as SIFT and ORB) or image descriptors extracted through deep learning (such as SuperPoint and LoFTR). Although these methods perform well under standard conditions, they still face significant performance degradation in cross-domain scenarios, both day and night. This is primarily due to the lack of consistent texture, brightness, and color distribution between images. Furthermore, existing methods are generally based on learning 2D image features and lack the ability to model 3D geometric structures, resulting in unstable matching in areas with occlusion, perspective changes, and low texture, which is particularly noticeable in the poor quality of nighttime imaging.

[0004] In recent years, some research has attempted to enhance the stability of image matching by introducing 3D geometric information through methods such as multi-view images or structured light. However, multi-view image acquisition is expensive, and achieving complete view coverage is difficult in dynamic or large-scale scenes. Furthermore, most of these methods rely on artificially constructed multi-view datasets, limiting their generalizability and scalability.

[0005] To address the issues of low image matching accuracy and unstable features in nighttime visual positioning, the present invention proposes a single-image 3D upscaling and feature enhancement method. By performing monocular depth estimation on nighttime images, upscaling them to a 3D point cloud space, and combining them with illumination re-rendering technology, paired data from different perspectives across lighting conditions is generated. A robust image matching model is trained by synthesizing a large number of data pairs across lighting conditions. This method breaks through the traditional visual positioning's reliance on multi-view data, achieving highly robust matching of daytime mapping results under nighttime conditions, and providing a new solution for visual positioning across lighting conditions. Summary of the Invention

[0006] The purpose of this invention is to address the problems of low positioning accuracy and unstable image matching in existing visual positioning methods at night or in low-light environments. This invention provides a nighttime visual positioning method and system based on three-dimensional dimensionality upscaling and feature enhancement. This method can accurately match daytime mapping results under conditions of significant illumination changes, thereby improving the robustness and practicality of visual positioning across lighting conditions. This method, without relying on multi-view image data or external sensors, utilizes only a single image and a depth estimation model to achieve high-quality construction of three-dimensional perception features, demonstrating high generalization capabilities and good deployment efficiency.

[0007] The innovation of the present invention is:

[0008] 1. We propose a feature enhancement method for upscaling single images to three-dimensional space. This method converts nighttime images into dense relative depth maps through monocular depth estimation. This is then projected onto a camera model to generate a three-dimensional point cloud, thereby upscaling the two-dimensional image into a three-dimensional structure with spatial perception capabilities. This process does not require an external depth sensor, offering the advantages of low cost and high adaptability.

[0009] 2. A training data synthesis strategy based on relighting and rerendering is designed. By performing depth-guided viewpoint transformation, occlusion repair, and illumination transformation on single-view images, a large number of cross-view and cross-illumination training image pairs are constructed to generate rich supervision signals, effectively solving the problem of lack of real labeled data for night images and providing strong data support for feature decoder training.

[0010] 3. We propose a two-stage feature matching network architecture for day-night image matching. In the first stage, a 3D-aware encoder is used to extract multi-view consistency features. In the second stage, a decoder is used to predict pixel-level transformations and matching confidences, achieving high-precision alignment between nighttime images and daytime maps. This network can be trained end-to-end without requiring annotation of day-night image pairs and demonstrates good generalization across multiple cross-domain tests.

[0011] 4. This invention boasts excellent system compatibility and cross-platform adaptability. Relying solely on image input and a depth estimation module, this method can be widely applied to low-cost camera systems, such as unmanned vehicles, security surveillance systems, augmented reality devices, and mobile robots. Furthermore, this method can be seamlessly integrated with existing SLAM systems and map-building frameworks, providing an integrated solution for visual positioning across multiple time periods and environments.

[0012] In order to achieve the above objectives, the present invention adopts the following technical solutions:

[0013] Step 1: The present invention first uses a pre-trained monocular depth estimation model to process nighttime images to generate a dense relative depth map. This depth map can provide spatial relationship information between pixels in the image without the need for real-world depth annotation. Through the camera intrinsic parameter model, the depth map is converted into a corresponding three-dimensional point cloud, thereby achieving dimensionality upgrade of the two-dimensional nighttime image to three-dimensional space. This step provides a structural foundation for subsequent feature enhancement and multi-view image synthesis, effectively improving the spatial understanding ability of low-light images;

[0014] Step 2: The present invention further designs a strategy for synthesizing training data pairs. Based on the night image and the depth map generated by it, a perspective transformation is performed to generate a new perspective image. At the same time, an image restoration algorithm is used to process the occluded area to ensure the integrity of the view. Subsequently, by introducing physical illumination modeling, the image is re-illuminated to simulate the lighting conditions of the daytime image, thereby obtaining night-day image pairs and their corresponding pixel-level matching labels. This synthetic data covers a wide range of scenes, angles, and lighting changes, greatly enhancing the diversity of training samples;

[0015] Step 3: Build a two-stage neural network model consisting of a 3D-aware encoder and a feature matching decoder. The first-stage encoder receives nighttime image input and outputs a 3D-consistent feature map. The second-stage decoder predicts pixel-level correspondences and matching confidences based on feature pairs between nighttime and daytime map images. This architecture can be trained on large-scale synthetic datasets and supports end-to-end inference on real nighttime images, eliminating the need for annotated image pairs.

[0016] Step 4: In practical applications, the trained model can be directly used for nighttime visual localization tasks. By feeding nighttime images into the model, extracting 3D perceptual features and matching them with daytime map images, robust cross-temporal image alignment can be achieved. This method is suitable for tasks such as visual relocalization in GPS-free environments, nighttime robot navigation, and day / night visual loop detection. It maintains high-precision matching even under strong light differences and texture loss.

[0017] Beneficial effects

[0018] Compared with the prior art, the present invention has the following advantages:

[0019] 1. No reliance on multi-view data or additional sensors, reducing system complexity and cost. This invention introduces monocular depth estimation and 3D feature modeling mechanisms to achieve dimensionality increase and multi-view feature learning based solely on a single image. This effectively reduces reliance on traditional multi-view data collection and annotation, significantly lowering the threshold for practical deployment, and is particularly suitable for mobile devices and resource-constrained platforms.

[0020] 2. Significantly improve the matching accuracy and robustness between nighttime images and daytime maps. By constructing a 3D perceptual encoder and extracting geometrically consistent image features, this invention achieves stable feature alignment in nighttime scenes with extremely low illumination and missing textures, enabling high-precision matching with daytime map images. This effectively addresses the problem of traditional methods failing to match images under both daytime and nighttime conditions.

[0021] 3. Innovative data synthesis and self-supervised training mechanisms enhance the model's cross-scenario generalization capabilities. This paper designs perspective transformation, occlusion repair, and relighting strategies based on depth maps, constructing large-scale, highly diverse training image pairs. This enables the model to learn matching strategies that adapt to complex lighting and perspective changes, and exhibits excellent zero-shot generalization capabilities.

[0022] 4. Balancing accuracy, efficiency, and system adaptability, the proposed two-stage feature matching network structure has strong practical application potential. It has good inference efficiency and can be seamlessly integrated with existing SLAM frameworks and visual relocalization systems. It is suitable for typical nighttime application scenarios such as unmanned driving, security monitoring, drone nighttime inspections, and robot navigation.

[0023] The innovative technical solution of the present invention takes into account the real-time and robustness requirements in practical applications while ensuring high precision, providing strong technical support for the application of SLAM technology in the fields of autonomous driving, robot navigation, and drone inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flow chart of the method of the present invention.

[0025] Figure 2 It is a detailed schematic diagram of the present invention. DETAILED DESCRIPTION

[0026] In order to better illustrate the purpose and advantages of the present invention, the present invention is further described below with reference to the accompanying drawings.

[0027] A vision and lidar fusion SLAM method based on depth completion, such as Figure 1 As shown, the following steps are included:

[0028] Step 1: Monocular depth estimation and 3D reconstruction.

[0029] Step 1.1: Monocular depth estimation.

[0030] First, the present invention first collects input images under natural lighting scenes. Starting from this, we use a monocular depth estimation model (such as Depth Anything V2) to process it and output the corresponding dense depth map:

[0031] D syn =a·M mono (I night )+b

[0032] Among them, M mono Denotes the depth estimation network, and ab are random scaling and offset factors, respectively, used to enhance depth diversity during training. The resulting depth map D syn The relative depth information of each pixel in the image is captured.

[0033] Step 1.2: 3D reconstruction of monocular depth map.

[0034] Next, combined with the camera intrinsic parameter matrix The depth d of each pixel u,v is d=D syn (u,v) is projected into a three-dimensional point (X,Y,Z) in the camera coordinate system:

[0035]

[0036] This constructs a dense 3D point cloud. To achieve subsequent new-perspective image rendering and lighting modeling, the discrete point cloud needs to be converted into a continuous mesh model. Therefore, a 3D mesh reconstruction algorithm based on geometric topology relationships, such as Poisson surface reconstruction or Delaunay triangulation, is used to triangulate the point cloud and generate a mesh model with normal vectors, topological structure, and patch information.

[0037] Step 2: New perspective image synthesis and lighting rendering.

[0038] Step 2.1: New perspective synthesis and content filling.

[0039] Based on the above 3D point cloud, multiple virtual camera perspectives (including pose transformation and translation) are sampled, and the original image is re-rendered under the new perspective to obtain a set of synthetic perspective images. The process includes the following steps: first, forward projection is performed using the 3D Mesh to obtain the position of each pixel under the new perspective. Then, visibility is determined by depth buffer sorting, and an occlusion mask M is used for the invisible area. occ Mark. Use image restoration model M inpaint Complete the image under the new perspective v:

[0040]

[0041] This allows us to obtain geometrically consistent images from new perspectives for training feature matching models.

[0042] Step 2.2: Re-render the lighting image.

[0043] Considering the significant lighting differences between nighttime images and daytime map images, and the difficulty in collecting nighttime images, direct matching is difficult. Therefore, the present invention performs a relighting process on the original natural lighting image to generate an image version with a nighttime lighting style.

[0044] This process is implemented by the rendering function R, where L represents the simulated vector of the new lighting conditions (such as light source direction, intensity, color):

[0045]

[0046] Relighting rendering constructs the scene geometry based on depth information and re-estimates reflections, shadows, and colors, making the output image close to the appearance of real nighttime low-light images, providing cross-lighting data support for subsequent matching.

[0047] Step 3: Feature matching model training.

[0048] Step 3.1: Feature consistency encoder training.

[0049] After obtaining the original natural lighting image, the night-time lighting new perspective image and its depth map, the present invention uses the natural lighting image (such as the daytime collected image) and the synthesized night-style new perspective image to construct training sample pairs, and designs a self-supervised feature consistency loss function to train the feature encoder E, thereby extracting image representations with three-dimensional geometric consistency and cross-illumination robustness.

[0050] Specifically, given the original natural lighting image I day and its synthesized new perspective night illumination image Input the encoder to be trained to extract its two-dimensional feature map:

[0051]

[0052] After aligning the image features from multiple perspectives using 3D point cloud coordinates, multiple loss functions are introduced for joint training to ensure that the features extracted by the encoder are invariant to perspective and illumination. First, reconstruction loss (L1) is used to perform pixel-level alignment between the original and new perspective features:

[0053]

[0054] Where u,v and m,n are the corresponding matching pixel locations of the original image and the new view image. A local block similarity metric (such as cosine or NCC) is used on the feature map to maintain the local geometric structure:

[0055]

[0056] The training process is based on synthetic perspectives and self-supervisory signals, and has good scalability and migration capabilities. After training, the encoder has strong three-dimensional structure perception and cross-illumination robustness, providing a stable foundation for subsequent feature matching decoder training.

[0057] Step 3.2: Freeze the encoder and train the feature matching decoder.

[0058] After the encoder is trained, its parameters are frozen and a decoder D is constructed to perform feature matching predictions on image pairs. During training, the real pixel correspondences of the synthetic data pairs are used as supervision signals, and the decoder performance is optimized using methods such as optical flow loss and confidence-weighted loss, resulting in a matching model that is robust across viewpoints and illumination conditions.

[0059] Step 4: Practical deployment in nighttime visual positioning tasks.

[0060] After training, the feature encoder and decoder proposed in this invention have good three-dimensional geometric consistency feature modeling capabilities and strong cross-illumination matching capabilities. They can be directly integrated into the visual positioning module of the autonomous driving system. They are particularly suitable for cross-time period repositioning tasks such as building maps during the day and positioning at night.

[0061] Specifically, during the operation of the autonomous driving mission, the forward-facing camera on the vehicle collects the current night image frame in real time, and inputs the image into the trained feature encoder to extract its deep feature map with three-dimensional perception capabilities. In addition, when the system is running during the day, it has completed the map construction (SLAM) process and retained the key frame image as a map library. For the current night image, its feature map will be matched with several candidate daytime image features in the map. With the assistance of the decoder, the pixel matching transformation field and matching confidence between the two images are calculated. Then, the system uses a robust geometric algorithm (such as 5-point algorithm+RANSAC) to estimate the relative pose of the night image with respect to the map image. And further combined with the absolute pose of the map image, the global position and orientation of the current image can be inferred, achieving accurate visual positioning under night conditions.

[0062] The above specific description further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A cross-illumination image feature matching method for nighttime visual positioning, characterized in that: The following steps are involved: Generate dense relative depth maps for nighttime images using a pre-trained monocular depth estimation model; The depth map is converted into a 3D point cloud through the camera intrinsic parameter model to achieve dimensionality upgrade of the 2D image to 3D space, and then a new perspective image synthesis is performed based on the 3D point cloud, including perspective transformation and occlusion repair; Perform illumination re-rendering on night images to generate images that simulate daytime lighting conditions. Then, use the natural illumination map and the new perspective night images to construct training data pairs. Construct a two-stage feature matching neural network. The first stage encoder extracts 3D perception features, and the second stage decoder predicts pixel-level matching and confidence. In practical applications, the trained feature matching model is used to perform feature matching between night images and daytime map images to achieve robust cross-illumination visual positioning.

2. The method according to claim 1, characterized in that The monocular depth estimation model is a pre-trained network that does not require real depth annotations and enhances depth diversity through random scaling and offset factors.

3. The method according to claim 1, characterized in that The three-dimensional point cloud is reconstructed using a Poisson surface or a Delaunay triangulation algorithm to generate a continuous triangular mesh, which is used to support the synthesis and occlusion processing of new perspective images.

4. The method according to claim 1, wherein The lighting re-rendering adopts a physical lighting modeling method based on depth information to simulate the direction, intensity and color of daytime light sources, thereby achieving style migration from nighttime images to daytime images.

5. The method according to claim 1, characterized in that In the feature matching neural network, the encoder is trained through feature consistency self-supervision loss, and the decoder is optimized using pixel-level matching supervision signals, supporting end-to-end training without real labeled data.

6. The method according to claim 1, wherein The practical applications include but are not limited to nighttime visual positioning of unmanned vehicles, nighttime robot navigation, nighttime security monitoring repositioning, and visual repositioning in nighttime inspection tasks.

Citation Information

Cited By

  • Robot multi-modal data enhancement system for industrial close-range grabbing scene

    CN121424404A

  • Robust single-frame structured light three-dimensional imaging method and system based on neural feature decoding

    CN121505174A