A method and device for depth estimation of a night scene fusion of heterogeneous images

By fusing heterogeneous images collected by RGB and infrared cameras and utilizing multi-scale depth feature extraction and pose estimation, the accuracy and real-time performance issues of depth estimation under low light conditions at night are solved, and real-time acquisition of dense depth information is achieved, thereby improving the safety and intelligent decision-making capabilities of the autonomous driving system.

CN119904498BActive Publication Date: 2025-10-17SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311406876.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-10-17
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

Under low-light conditions at night, existing technologies have insufficient depth estimation accuracy of monocular RGB images, poor infrared image quality, high-cost lidar and sparse depth information, and poor real-time performance of binocular stereo vision, resulting in insufficient depth estimation accuracy and real-time performance of autonomous driving systems in night scenes.

Method used

An RGB camera and an infrared camera are used to synchronously capture heterogeneous images. The heterogeneous image features are fused through a multi-scale deep feature extraction network with a U-Net structure and a CNN network. Combined with pose estimation and epipolar geometry constraints, a depth estimation device that fuses heterogeneous images is established to achieve pixel-by-pixel dense depth estimation.

Benefits of technology

It improves the depth estimation accuracy in night scenes, overcomes the shortcomings of a single sensor, reduces costs, realizes the real-time acquisition of dense depth information, assists the autonomous driving system in intelligent decision-making in low light conditions, and ensures safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904498B_ABST
    Figure CN119904498B_ABST
Patent Text Reader

Abstract

The application relates to a kind of depth estimation method and device for night scene fusion of heterogeneous image, including heterogeneous image acquisition module, depth estimation module of fusion heterogeneous image and attitude estimation module.Heterogeneous image acquisition module carries out the real-time synchronous acquisition of heterogeneous image.In the depth estimation module of fusion heterogeneous image, RGB image feature extraction unit and infrared image feature extraction unit respectively realize the extraction of depth-related features in RGB image and infrared image;heterogeneous feature fusion unit realizes the real-time depth estimation of scene by fusing the above two features.The camera attitude obtained by the attitude estimation module is used to constrain the depth estimation result in the depth estimation module, so that it satisfies the distortion transformation in epipolar geometry.The application establishes the real-time depth estimation for night scene, synchronously acquires RGB image and infrared image, and fuses the corresponding depth features, to realize the real-time perception of surrounding environment under night scene, with strong flexibility and practicality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image processing, and particularly relates to a real-time depth estimation method and device for night scene, which is applied to real-time depth estimation of targets including obstacles, pedestrians and the like in a night low-light scene of automatic driving, and assists intelligent decision-making of subsequent automatic driving. BACKGROUND

[0002] Real-time perception is an important direction in the field of automatic driving at present, especially real-time depth estimation of various targets on the road under night low-light conditions, which can effectively guarantee the safety of automatic driving. At the same time, real-time depth estimation of night scene can also optimize the night travel route in real time and perfect the intelligent decision-making of the automatic driving system under low light. Therefore, depth estimation for night scene has very important significance for the safety of the entire automatic driving system.

[0003] For depth estimation for night scene, the common methods mainly include: the first method is to directly use the RGB image under low light as input, and adopt a self-supervised monocular depth estimation method for depth estimation. The disadvantage of this method is that the RGB image in the night scene lacks a lot of detail information, and there are also some image blur conditions, resulting in low precision of the estimated depth value; another method is to directly use the image collected by an infrared camera as input, and adopt a self-supervised monocular depth estimation model for depth estimation. Although the infrared camera can avoid the loss of details caused by low light, the image itself has poor quality, mainly in low contrast, blurred visual effect, low signal-to-noise ratio and unobvious local features. These shortcomings make the depth estimation precision of this method insufficient; the third method is to adopt a laser radar or the like for depth estimation. The disadvantage of this method is that the cost is too high, and the obtained depth information is sparse, i.e. not pixel by pixel, in addition, this method has strict limitations on the measured distance. In addition, there is a depth estimation method based on binocular stereo vision. This method not only has poor real-time performance due to high computational complexity, but also has the same shortcomings in the method, i.e. limited by the loss of scene information caused by low light. SUMMARY

[0004] In order to solve the above-mentioned deficiencies in the prior art, the present application proposes a depth estimation method and device for night scene fusion of heterogeneous images, which is used to solve the problem of insufficient depth map precision in the prior art of depth estimation for night scene, which only adopts monocular depth estimation based on RGB image and only adopts monocular depth estimation based on infrared image, and also solves the problems of high cost and insufficient dense depth information of laser radar and the like sensors, and overcomes the shortcomings of poor real-time performance and insufficient precision of the prior art of night scene depth estimation based on binocular stereo vision.

[0005] The technical means adopted by the present application are as follows:

[0006] A depth estimation device for night scene fusion of heterogeneous images is provided, and the following modules are set up: an offline network model is established and iteratively trained to obtain an optimized device for depth estimation of night scene heterogeneous images, and the device comprises:

[0007] A heterogeneous image acquisition module, which synchronously acquires night scene images of heterogeneous images by controlling the heterogeneous camera, and establishes training set data;

[0008] A depth estimation module for fusion of heterogeneous images, which comprises an RGB image feature extraction unit, an infrared image feature extraction unit and a heterogeneous feature fusion unit, is used to extract and fuse image features of heterogeneous images for depth estimation; and simultaneously combines the relative pose estimation T t→t′ The night target scene is reconstructed, and the training is performed through the corresponding photometric loss to realize night scene depth estimation based on heterogeneous images;

[0009] A pose estimation module, which comprises an RGB camera pose estimation unit, an infrared camera pose estimation unit and a pose consistency constraint unit, is used to estimate the camera pose of the heterogeneous camera and to strengthen convergence through consistency constraint; and finally combines the relative pose estimation result T t→t′ with the estimation result of the depth estimation module to realize simultaneous training of the pose network and the depth estimation network.

[0010] The heterogeneous camera is an RGB camera and an infrared camera, which are used to acquire real-time monocular images in a night scene; and the training set data is a pair of heterogeneous images of the current frame, the previous frame and the next frame.

[0011] The heterogeneous image acquisition module further comprises an acquisition controller for synchronously triggering the synchronous acquisition of the heterogeneous camera; and the RGB camera and the infrared camera are on the same horizontal line and are close to each other for shooting the same angle of view.

[0012] The depth estimated by the depth estimation module for fusion of heterogeneous images is based on the scene in the RGB camera to obtain the depth value of each pixel in the RGB image.

[0013] Both the RGB image feature extraction unit and the infrared image feature extraction unit adopt a multi-scale depth feature extraction network based on the U-Net structure, and the internal encoder of each unit adopts a network structure based on MobileNetV3; the network input is the corresponding synchronous image, and the network output is the corresponding depth-related feature map.

[0014] After the infrared image features are extracted, mask processing is further performed, which is used to mask the identification area that only appears in the infrared image, and to retain the identification area that simultaneously appears in the heterogeneous images as a mask template area.

[0015] The mask screen masks the corresponding position on the feature map obtained by the infrared image feature extraction unit to 0; the mask template area is obtained by neural network training.

[0016] The RGB camera pose estimation unit and the infrared camera pose estimation unit both adopt a CNN network, estimate the camera poses between adjacent RGB image frames and adjacent infrared image frames respectively, and are constrained by the pose consistency constraint unit, so that the camera pose estimation results of the two heterogeneous cameras should be consistent.

[0017] The camera pose estimated by combining the epipolar geometry constraint and the depth information satisfies the twist transformation in the epipolar geometry, and the output relative pose estimation T t→t′ It includes the offset of the camera in three directions and three angle changes.

[0018] The reconstructed night scene and the final output of the dense depth map include:

[0019] Combined with the locally differentiable bilinear interpolation, the source image I t′ The reconstructed target reconstructed image I t′→t ;

[0020] The per-pixel luminosity loss after reconstruction is established for iterative training, and the network model parameters of the depth estimation module and the pose estimation module of the fused heterogeneous image are optimized.

[0021] A depth estimation method for fused heterogeneous images in a night scene is used to perform the following steps in a night driving scene to perform real-time night scene depth estimation:

[0022] 1) The heterogeneous image acquisition module controls two heterogeneous cameras to perform synchronous image acquisition in real time;

[0023] 2) In the depth estimation module of the fused heterogeneous image, the RGB image feature extraction unit and the infrared image feature extraction unit receive the real-time RGB image and the real-time infrared image respectively, and perform corresponding feature extraction, and the heterogeneous feature fusion unit will fuse the above two features, and finally output the estimated dense depth map.

[0024] The beneficial effects of the present application are:

[0025] 1. The present application establishes the depth perception of the automatic driving system in the night scene, which can solve the problem of insufficient accuracy when using a single source image for depth estimation in low light to a certain extent, and overcome the shortcomings of high cost and sparse estimation of using laser sensors, and can improve the intelligent decision-making of the automatic driving system in low light, ensure the safety of the automatic driving in the night scene, and has strong practicality.

[0026] 2. The application discloses a night scene fusion heterogeneous image depth estimation method and device, which comprises a heterogeneous image acquisition module, a fusion heterogeneous image depth estimation module and a pose estimation module. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 Figure 1 is a schematic diagram of the night scene fusion heterogeneous image depth estimation device of the application;

[0028] Figure 2 Figure 2 is a flowchart of the fusion heterogeneous image depth estimation module of the application;

[0029] Figure 3 Figure 3 is a flowchart of the pose estimation module of the application. DETAILED DESCRIPTION

[0030] In order to make the above objectives, features and advantages of the application more apparent, specific implementation methods of the application are described in detail below in combination with the drawings. In the following description, a large number of specific details are set forth in order to facilitate a full understanding of the application. However, the application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the application, so the application is not limited by the specific implementation disclosed below.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application belongs. The terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit the application.

[0032] The application establishes a real-time depth estimation method for night scene fusion heterogeneous images, which is used for automatic driving systems to perceive real-time depth information of the surrounding environment in low-light night scenes, including obstacles, pedestrians and other targets in the night low-light scene of the automatic driving system. The application can effectively improve the accuracy of real-time depth estimation, thereby better assisting subsequent decision-making tasks, and improving the safety of the automatic driving system. Therefore, the application has strong practicability.

[0033] A night scene fusion heterogeneous image depth estimation device, as shown in Figure 1As shown, it comprises: a heterogeneous image acquisition module, a depth estimation module for fused heterogeneous images and a pose estimation module; the heterogeneous image acquisition module controls two heterogeneous cameras to synchronously acquire by an acquisition controller, and inputs the acquired heterogeneous images to the depth estimation module for fused heterogeneous images and the pose estimation module; the depth estimation module outputs real-time depth information, and the pose estimation module outputs corresponding camera pose information according to the input heterogeneous image sequence in the training process. In addition, the pose estimation module and the depth estimation module are mutually constrained by performing twist transformation in epipolar geometry in the training process, so that simultaneous training is realized.

[0034] As shown in the figure, Figure 2 The flowchart of the depth estimation module for fused heterogeneous images is shown, and the acquired real-time heterogeneous images will be input into the corresponding feature extraction unit. Since the finally output dense depth map is pixel by pixel corresponding to the RGB real-time image, and the RGB real-time image and the infrared real-time image have a certain offset, the features obtained by the infrared image feature extraction unit need to be further processed by performing dot product with a binary mask, so as to reduce the influence of noise regions such as occlusion. The binary mask is also obtained by training. The processed features are concatenated with the features Figure 2 in the channel, and then a series of convolutions are performed to realize feature fusion, and finally real-time depth information is output. Figure 2 Figure 2 Figure 1

[0035] As shown in the figure, Figure 3 The flowchart of the pose estimation module is shown, and the module is only used in the training process. In the training process, the module inputs the corresponding heterogeneous image sequence, and simultaneously realizes the pose estimation of the two heterogeneous cameras through the corresponding camera pose estimation unit. Since the two cameras are fixed on the same rigid support and are close to each other, they can be constrained by the pose consistency unit, that is, the pose results of the two cameras should be consistent. The module finally outputs the camera pose estimation result. In the training process, the camera pose estimation result and the depth estimation result are constrained by the twist transformation in epipolar geometry, so as to realize the mutual assistance of the pose estimation module and the depth estimation module in the training process.

[0036] The depth estimation device for fused heterogeneous images in night scene of the present application in real-time depth estimation, the specific steps are as follows:

[0037] S1 When the depth estimation device receives the instruction of collecting training images, the acquisition controller will control two heterogeneous cameras to synchronously acquire, and save the obtained heterogeneous image sequence;

[0038] ​​​S2 preprocesses the acquired image sequences to create a corresponding training set. Each subsequence in the training set contains an RGB image sequence and a synchronized infrared image sequence, and both sequences contain three frames of images. Here, the middle frame is named the target image, and the previous and next frames are named the source image. Since the training process uses a self-supervised method, no real depth information is required for annotation.

[0039] During the S3 training process, the heterogeneous target image in the subsequence is input into the depth estimation module that fuses heterogeneous images to obtain the depth map corresponding to the RGB target image; the complete subsequence is input into the posture estimation module, and after passing through two heterogeneous camera posture estimation units and a posture consistency constraint unit, the corresponding camera posture estimation result is obtained.

[0040] During the S4 training process, it is necessary to combine the distortion transformation to establish the relationship between the original RGB target image and the source image. t and p t′ Represent the original target image I t and source image I t′ Therefore, the following formula can be established based on epipolar geometry:

[0041] p t′ ~KT t→t′ D t (p t )K -1 p t

[0042] Among them, K represents the camera internal parameter, D t Represents the depth estimation result obtained by the depth estimation module, T t→t′ is the relative pose estimation result output by the pose estimation module. According to the above formula, combined with the local differentiable bilinear interpolation, we can get the relative pose estimation result of the source image I t′ Reconstructed target image I t′→t , by establishing the pixel-by-pixel photometric loss after reconstruction, training is carried out. After the loss function value converges, the training is terminated, and finally the corresponding network model parameters of the depth estimation module and the pose estimation module of the heterogeneous image fusion are obtained;

[0043] When S5 is running in the test phase, when it receives the perception signal, the acquisition controller will synchronously collect real-time RGB images and real-time infrared images of the night scene and transmit them to the trained depth estimation module that fuses heterogeneous images;

[0044] In the depth estimation module S6, the RGB image feature extraction unit and the infrared image feature extraction unit receive the RGB real-time image and the infrared real-time image respectively, and perform corresponding feature extraction to obtain feature Figure 1 and featuresFigure 2 ;

[0045] S7 after feature extraction, feature Figure 2 First, combined with the trained binary mask dot product operation, and then with the feature Figure 1 In the channel direction, cascade and after multi-layer convolution operation, finally realize the fusion of heterogeneous features, and output the corresponding real-time depth information.

[0046] The depth estimation method and device for fusing heterogeneous images in night scene, through the heterogeneous image acquisition module, the depth estimation module of the fused heterogeneous image and the attitude estimation module, a set of depth estimation method and device based on heterogeneous monocular vision is constructed, the dense depth information under the night scene can be estimated in real time, so as to better assist the corresponding decision task of the subsequent automatic driving system, and the safety of the automatic driving system is improved and guaranteed.

[0047] The above is the preferred embodiment of the present application, it should be pointed out that, for those skilled in the technical field, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should be regarded as the protection scope of the present application.

Claims

1. A depth estimation device for fusing heterogeneous images for night scenes, characterized by: The following modules are set up to establish a network model offline and iteratively train it to obtain an optimized device for depth estimation of heterogeneous images of night scenes. The device includes: The heterogeneous image acquisition module controls heterogeneous cameras to synchronously acquire heterogeneous night scene images and establish training set data; The depth estimation module for fusing heterogeneous images includes an RGB image feature extraction unit, an infrared image feature extraction unit, and a heterogeneous feature fusion unit, which is used to extract and fuse heterogeneous image features for depth estimation. At the same time, the relative posture estimation T output by the posture estimation module is combined with the t→t′ Reconstruct the night target scene, train through the corresponding photometric loss, realize the night scene depth estimation based on the heterogeneous image, and finally output the dense depth map; wherein, after extracting the infrared image features, it also includes mask processing, which is used to mask the identification area that only appears in the infrared image, and retain the identification area that appears in the heterogeneous image at the same time as the mask template area; wherein, the mask shielding sets the corresponding position on the feature map obtained by the infrared image feature extraction unit to 0; the mask template area is obtained by neural network training; wherein, the final output of the dense depth map includes: combining the locally differentiable bilinear interpolation to obtain the source image I t′ Reconstructed target reconstructed image I t′→t Establish a pixel-by-pixel photometric loss after reconstruction for iterative training to optimize the corresponding network model parameters of the depth estimation module and the pose estimation module that fuse heterogeneous images; The pose estimation module includes an RGB camera pose estimation unit, an infrared camera pose estimation unit, and a pose consistency constraint unit. It is used to estimate the pose of heterogeneous cameras and strengthen the convergence through consistency constraints. Finally, the relative pose estimation result T is combined with the epipolar geometry. t→t′ Combined with the estimation results of the depth estimation module, the simultaneous training of the pose network and the depth estimation network is realized; wherein, the RGB camera pose estimation unit and the infrared camera pose estimation unit both adopt CNN networks to estimate the camera pose between adjacent RGB image frames and the camera pose between adjacent infrared image frames, respectively, and are constrained by the pose consistency constraint unit so that the camera pose estimation results of the two heterogeneous cameras should be consistent; wherein, the camera pose and depth information estimated by combining the epipolar geometry constraint satisfy the distortion transformation in the epipolar geometry, and the relative pose estimation T t→t′ Including the camera offset in three directions and three angle changes.

2. The depth estimation device for fusing heterogeneous images for nighttime scenes according to claim 1, characterized in that: The heterogeneous camera is an RGB camera and an infrared camera, which are used to collect real-time monocular images in night scenes; the training set data is a heterogeneous image pair of the current frame and the previous and next adjacent frames; The heterogeneous image acquisition module also includes an acquisition controller for synchronously triggering heterogeneous cameras to acquire images synchronously; the RGB camera and the infrared camera are on the same horizontal line and are close to each other, so as to capture the same viewing angle.

3. The depth estimation device for fusing heterogeneous images for nighttime scenes according to claim 1, characterized in that: The depth estimation module for fusing heterogeneous images estimates the depth based on the scene in the RGB camera, and obtains the pixel-by-pixel depth value in the RGB image.

4. The depth estimation device for fusing heterogeneous images for nighttime scenes according to claim 3, characterized in that: The RGB image feature extraction unit and the infrared image feature extraction unit both adopt a multi-scale deep feature extraction network based on the U-Net structure, and their internal encoders both adopt a network structure based on MobileNetV3; the network input is the corresponding synchronized image, and the network output is the corresponding depth-related feature map.

5. A depth estimation method for fusing heterogeneous images in nighttime scenes, implemented based on the device according to any one of claims 1 to 4, characterized in that: It is used to perform the following steps in night driving scenarios to perform real-time depth estimation of night scenes: 1) The heterogeneous image acquisition module controls two heterogeneous cameras in real time to perform synchronous image acquisition; 2) In the depth estimation module that fuses heterogeneous images, the RGB image feature extraction unit and the infrared image feature extraction unit receive the RGB real-time image and the infrared real-time image respectively and perform corresponding feature extraction. The heterogeneous feature fusion unit will fuse the above two features and finally output the estimated dense depth map.

Citation Information

Patent Citations

  • Impeller quality detection method and device based on multi-dimensional monocular depth estimation

    CN113012091A

  • Night vision anti-halation method for multi-region fusion of different-source images

    CN116934812A