A Depth Estimation Method Based on Adaptive Structured Light

Through an end-to-end adaptive active stereo vision system, combined with real optical systems and optical simulation models, structured light generation and depth estimation are optimized, which solves the problem that fixed structured light mode cannot adapt to scene changes, and improves the accuracy and robustness of depth estimation.

CN120125633BActive Publication Date: 2025-07-22NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510621767.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The fixed structured light mode in the existing active stereo vision system cannot dynamically adapt to scene changes, and the depth estimation network does not fully integrate multimodal information, and there is a gap between simulation and reality.

Method used

Design an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically-aware differentiable imaging module and a parallax attention depth estimation module, optimize structured light generation and depth estimation through joint training, and combine real optical systems and optical simulation models to achieve hardware-model collaborative optimization.

Benefits of technology

The structured light mode is adjusted in real time according to the scene, which improves the accuracy and robustness of depth estimation, reduces the gap between simulation and reality, and improves the overall performance and real-time nature of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125633B_ABST
    Figure CN120125633B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of stereoscopic vision technology, and discloses a depth estimation method based on adaptive structured light. The RGB image of the camera view in the scene with the projector turned off is collected, and the coarse disparity map of the camera view is obtained through the semi-global matching algorithm. The RGB image of the projector view and the coarse disparity map of the projector view are generated by projective transformation and input into the adaptive structured light generation module to generate an adaptive structured light projection pattern. The adaptive structured light projection pattern is input into the physically-aware differentiable imaging module to generate the structured light image of the camera view. The disparity attention depth estimation module receives the RGB image of the camera view and the structured light image of the camera view and outputs the depth estimation result. Taking the depth error as the loss function, the adaptive structured light generation module and the disparity attention depth estimation module are jointly trained and optimized, and depth estimation is performed after optimization. The present invention dynamically generates the optimal structured light pattern according to the real-time information of the scene, and solves the problem that the traditional fixed-mode structured light cannot adapt to the scene change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of stereoscopic vision technology, and in particular to a depth estimation method based on adaptive structured light. Background Art

[0002] Active stereo vision technology combines structured light projection with binocular cameras, showing significant advantages in the field of 3D environment perception, and is widely used in autonomous driving, robot navigation, and 3D scene reconstruction. Unlike passive depth sensors that rely on the object's own texture, active stereo systems project specially designed structured light patterns onto the surface of objects, providing additional corresponding clues for binocular matching, thereby improving the stability and accuracy of depth estimation.

[0003] Most of the existing active stereo techniques use fixed structured light projection patterns. Some techniques design high-frequency structured light based on imaging theory to combat the effects of global illumination and defocus. For example, "M. Gupta and N. Nakhate, "Ageometric perspective on structured light coding," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 87–102." proposed a new paradigm of structured light coding based on geometric analysis, Hamiltonian coding. This technology uses coding curves as the geometric representation of structured light coding, maps the K-dimensional projection pattern into a continuous trajectory in the hypercube space, and then establishes a negative correlation between the length of the coding curve and the decoding error by deriving a proxy metric to optimize the structured light coding design. However, this type of method has the defect of scene adaptability, which is specifically manifested in the lack of a scene dynamic feedback correction mechanism. Other technical solutions propose data-driven structured light design methods to optimize structured light patterns through scene data. For example, "S.-H. Baek and F. Heide, "Polkalines: Learning structured illumination and reconstruction for active stereo," in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753–5763." proposed a customized stripe projection pattern that can learn to adapt to different environments (indoor, outdoor, and general scenes). By inputting a large amount of scene data to optimize the structured light end-to-end, the depth estimation effect of high-dynamic scenes is improved. However, this type of method uses an offline training mode, which leads to parameter solidification. After the structured light design is completed, it cannot be optimized again, so it cannot adapt to new scenes online.

[0004] In terms of depth estimation networks, current technologies typically only use structured light patterns as inputs and rely on the artificial textures of structured light images for matching. For example, "Y. Zhang, S. Khamis, C. Rhemann, J. Valentin, A. Kowdle, V. Tankovich, M. Schoenberg, S. Izadi, T. Funkhouser, and S. Fanello, 'Active stereo net: End-to-end self-supervised learning for active stereo systems,' in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 784–801." proposed a structured light stereo matching network that estimates depth end-to-end from a pair of structured light inputs, improving the performance of traditional methods in noisy, low-texture, and occluded regions to a certain extent. However, such methods may introduce noise in regions with rich natural textures and ignore the rich texture information in natural images, limiting the robustness of depth estimation. Summary of the Invention

[0005] In view of the problems in the existing active stereo vision system, such as the fixed structured light pattern being unable to dynamically adapt to the scene, the depth estimation network not fully integrating multi-modal information, and the gap between simulation and reality, the present invention proposes a depth estimation method based on adaptive structured light, designs an end-to-end adaptive active stereo vision system, and realizes the collaborative optimization of structured light and depth estimation through an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module, improving the depth perception performance in complex scenes.

[0006] The technical solution of the present invention is as follows: A depth estimation method based on adaptive structured light designs an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module;

[0007] Collect the RGB image from the camera's perspective in the scenario of turning off the projector, and obtain the rough disparity map from the camera's perspective through the semi-global matching algorithm; the RGB image from the camera's perspective and the rough disparity map from the camera's perspective are used to generate the RGB image from the projector's perspective and the rough disparity map from the projector's perspective through projective transformation; the RGB image from the projector's perspective and the rough disparity map from the projector's perspective are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; the adaptive structured light projection pattern is used by the physically-aware differentiable imaging module to generate the structured light image from the camera's perspective; the disparity attention depth estimation module simultaneously receives the RGB image from the camera's perspective and the structured light image from the camera's perspective, and outputs the depth estimation result; using the depth error as the loss function, jointly train and optimize the learning parameters of the adaptive structured light generation module and the learning parameters of the disparity attention depth estimation module, and use the optimized end-to-end adaptive active stereo vision system for depth estimation.

[0008] The projective transformation is based on the calibration parameters of the binocular camera and the calibration parameters of the projector, and maps the pixel points in the binocular camera coordinate system to the projector coordinate system through the geometric transformation formula.

[0009] In the adaptive structured light generation module, structured light encoding is performed;

[0010] The specific structured light encoding is as follows: an encoder is used to extract features from the RGB image from the projector's perspective and the rough disparity map from the projector's perspective; the extracted features are respectively used by the SegNet decoder for local texture decoding to generate local texture patterns and by the SegNet decoder for global brightness decoding to generate global brightness patterns , and the adaptive structured light projection pattern is calculated :

[0011] .

[0012] The encoder consists of several convolutional layers; the convolutional layers use convolutional kernels of different sizes and different strides.

[0013] The physically-aware differentiable imaging module includes a real optical system and an optical simulation model ; during the joint training and optimization process, the real optical system is used during forward propagation, and the optical simulation model is used during gradient backpropagation.

[0014] During forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern onto the scene. After the adaptive structured light projection pattern is reflected by the objects in the scene, the binocular camera collects and obtains the structured light image from the camera's perspective; turn off the real projector to obtain the RGB image from the camera's perspective;

[0015]

[0016] Represents a scene; Represents the structured light image from the camera's perspective; Represents the RGB image from the camera's perspective.

[0017] During backpropagation, an optical simulation model is used to calculate the loss function of the disparity attention depth estimation module With respect to the adaptive structured light projection pattern Gradient estimation:

[0018]

[0019] Where Represents the loss function Gradient estimation of the adaptive structured light projection pattern Gradient estimation, Represents the gradient estimation of the adaptive structured light projection pattern by the optical simulation model Gradient estimation, Represents matrix transpose operation, Is the simulated structured light image of the dual-camera perspective;

[0020] The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model;

[0021] The point spread function of the simulation projector sub-model is expressed as the squared magnitude of the Fourier transform of the complex-valued pupil function Of:

[0022]

[0023] Where, Is the amplitude of the structured light, Is the wavelength of the structured light; phase Is the scene depth Function of, characterizing the defocus degree caused by the mismatch between the projector focus depth and the scene depth of the scene point Mismatch; Is the Fourier transform; the output of the simulation projector sub-model is the structured light pattern from the projector's perspective Generated by convolving the point spread function with the adaptive structured light projection pattern output by the adaptive structured light generation module Obtained:

[0024]

[0025] Where Represents 2D convolution;

[0026] The simulation camera sub-model, through the coarse disparity map of the camera's perspective Transform the structured light pattern under the projector's perspective to the left camera's perspective and the right camera's perspective to obtain the illumination images under the camera's perspective :

[0027]

[0028] where is the image transformation operation;

[0029] Render based on the Lambertian reflection model to obtain the output of the simulated camera sub-model, and get the simulated structured light images from the dual-camera perspective , and the formula is as follows:

[0030]

[0031] where represents the influence coefficients of camera exposure, aperture, and sensitivity on pixel brightness; is the result of grayscale conversion of the RGB image from the camera's perspective, representing the contribution from ambient light to the brightness of the simulated structured light images from the dual-camera perspective ; is the power coefficient of the projector; is the reflectivity map, obtained from the RGB image from the camera's perspective by the RGB-inversion method, representing the contribution from the projector light to the brightness of the simulated structured light images from the dual-camera perspective ; is Gaussian noise, and , is the brightness clipping function.

[0032] The disparity attention depth estimation module includes a feature extraction sub-module, a cost volume construction sub-module, a cost aggregation sub-module, and a disparity prediction sub-module; the RGB image from the camera's perspective and the structured light image from the camera's perspective are respectively processed by the feature extraction sub-module to extract corresponding features, obtaining the corresponding features and of the RGB image from the camera's perspective, as well as the corresponding features and of the structured light image from the camera's perspective; the corresponding features are jointly input into the cost volume construction sub-module to obtain the combined attention cost volume, and then sequentially pass through the cost aggregation sub-module and the disparity prediction sub-module to output the depth estimation result.

[0033] In the cost volume construction sub-module, each disparity level of and is concatenated to construct the concat cost volume :

[0034]

[0035] Among them represents and the pixel positions, d is the disparity level, and Concat is the splicing operation;

[0036] and are divided into feature groups along the channel dimension, each feature group has channels, is the total number of channels of the features; Denote the th feature group as and , and the corr cost volume is obtained through the following formula:

[0037]

[0038] Among them represents the inner product of matrices;

[0039] The hourglass module is used to extract the first disparity attention weight from the corr cost volume , and extract the second disparity attention weight from the concat cost volume ;

[0040] The hourglass module is divided into a downsampling part and an upsampling part. The downsampling part first passes through a 3D convolutional layer to extract preliminary features, and then continuously uses 3D convolutional layers to gradually reduce the spatial resolution and disparity resolution of the extracted features; The upsampling part uses 3D transposed convolutional layers to gradually restore the spatial resolution and disparity resolution, and then passes through a 3D convolutional layer to obtain the output; The hourglass module gradually reduces the resolution in the downsampling part to form a "narrow part", and then restores the feature map to the original resolution in the upsampling part to form a "wide part";

[0041] Through the first disparity attention weight and the second disparity attention weight, filter the corr cost volume and the concat cost volume , and obtain the Concat attention cost volume and the Corr attention cost volume :

[0042]

[0043]

[0044] Among them is and The number of channels; finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume :

[0045] .

[0046] The cost aggregation sub-module is connected to the disparity prediction sub-module; the cost aggregation sub-module first uses four 3D convolutions to preliminarily aggregate the combined attention cost volume, and then uses two hourglass modules to learn context information through an encoder-decoder structure to obtain the aggregated cost;

[0047] The disparity prediction sub-module uses disparity regression to obtain the predicted disparity map from the aggregated cost; the weight of the disparity level d is obtained through a softmax operation from the aggregated cost and then the disparity value is calculated using the following formula:

[0048]

[0049] where is the maximum value of the disparity level; the disparity value is converted to the depth value :

[0050]

[0051] where B is the baseline length of the binocular camera and f is the focal length of the binocular camera;

[0052] The loss function of the disparity attention depth estimation module is as follows:

[0053]

[0054] .

[0055] Advantages of the present invention: The present invention can dynamically generate an optimal structured light pattern according to the real-time information of the scene, solving the problem that the traditional fixed-pattern structured light cannot adapt to scene changes.

[0056] The present invention realizes the collaborative optimization of hardware-model by combining data acquisition through a real optical system and an optical simulation model, effectively reducing the simulation-reality gap.

[0057] The disparity attention depth estimation module of the present invention adopts a bidirectional disparity attention mechanism, fully integrating the information of the structured light image and the RGB image, improving the accuracy and robustness of depth estimation.

[0058] The present invention realizes end-to-end joint learning of structured light generation and depth estimation, avoids the limitations of independent training of each module in traditional methods, and improves the overall performance of the system. Description of the Drawings

[0059] Figure 1 is the overall technical route of the present invention;

[0060] Figure 2 is a schematic diagram of the adaptive structured light generation module;

[0061] Figure 3 is a schematic diagram of the disparity attention depth estimation module;

[0062] Figure 4 is a schematic diagram of the feature extraction sub-module;

[0063] Figure 5 is a schematic diagram of the cost aggregation sub-module. Detailed Embodiment

[0064] The overall technical solution of the present invention is as Figure 1 shown. An end-to-end adaptive active stereo vision system is designed, which mainly includes three parts: an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module. The RGB image of the camera view in the scene with the projector turned off is collected, and the coarse disparity map of the camera view is obtained through the semi-global matching algorithm; the RGB image of the camera view and the coarse disparity map of the camera view are used to generate the RGB image of the projector view and the coarse disparity map of the projector view through projective transformation, and are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; the adaptive structured light projection pattern is used to generate the structured light image of the camera view through the physically-aware differentiable imaging module; the disparity attention depth estimation module simultaneously receives the RGB image of the camera view and the structured light image of the camera view and outputs the depth estimation result; taking the depth error as the loss function, the learning parameters of the adaptive structured light generation module and the learning parameters of the disparity attention depth estimation module are jointly trained and optimized, and finally the optimized end-to-end adaptive active stereo vision system is used for depth estimation.

[0065] The adaptive structured light generation module is as Figure 2As shown, the adaptive structured light generation module first uses an encoder to extract features from the input. The encoder consists of a series of convolutional layers to capture the multi-scale feature information of the image. These convolutional layers use convolutional kernels and strides of different sizes to gradually reduce the size of the feature map while increasing the dimension of the features. Then, the SegNet decoder for local texture decoding "V. Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 12, pp. 2481–2495, 2017." and the SegNet decoder for global brightness decoding respectively generate the local texture pattern and the global brightness pattern , and finally the adaptive structured light projection pattern is calculated through the following formula :

[0066]

[0067] The physically-aware differentiable imaging module contains a real optical system and an optical simulation model in two parts. For this physically-aware differentiable imaging module, during the overall framework training process, the real optical system is used during forward propagation, and the optical simulation model is used during backpropagation.

[0068] During forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern into the scene. After being reflected by the objects in the scene, the adaptive structured light projection pattern is captured by the binocular camera to obtain the structured light image from the camera's perspective ; the real projector is turned off to obtain the RGB image from the camera's perspective ;

[0069]

[0070] represents the scene; represents the structured light image from the camera's perspective; represents the RGB image from the camera's perspective.

[0071] During backpropagation, the optical simulation model is used to calculate the loss function of the disparity attention depth estimation module with respect to the adaptive structured light projection pattern Gradient estimation of:

[0072]

[0073] where represents the loss function for the gradient estimation of the adaptive structured light projection pattern ; represents the gradient estimation of the adaptive structured light projection pattern by the optical simulation model ; represents the matrix transpose operation; is the structured light image of the dual-camera perspective simulation;

[0074] The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model;

[0075] The point spread function of the simulation projector sub-model is expressed as the squared magnitude of the Fourier transform of the complex-valued pupil function :

[0076]

[0077] where is the amplitude of the structured light, is the wavelength of the structured light; the phase is a function of the scene depth characterizing the defocus degree caused by the mismatch between the projector focus depth and the scene depth of the scene point ; is the Fourier transform; the output of the simulation projector sub-model is the structured light pattern from the projector perspective , obtained by convolving the point spread function with the adaptive structured light projection pattern output by the adaptive structured light generation module :

[0078]

[0079] where represents 2D convolution;

[0080] For the simulation camera sub-model, through the coarse disparity map of the camera perspective the structured light pattern from the projector perspective is transformed to the left camera perspective and the right camera perspective, obtaining the illumination image from the camera perspective :

[0081]

[0082] where is the image transformation operation;

[0083] Rendering based on the Lambertian reflection model yields the output of the simulated camera sub-model, obtaining the dual-camera perspective simulated structured light images from the left camera perspective and the right camera perspective , as shown in the following formula:

[0084]

[0085] Where represents the influence coefficients of camera exposure, aperture, and ISO on pixel brightness; is the result of grayscale conversion of the camera perspective RGB image, representing the dual-camera perspective simulated structured light image The contribution from ambient light to the brightness of; is the power coefficient of the projector; is the reflectivity map, obtained from the camera perspective RGB image by the RGB-inversion method "S.-H. Baek and F. Heide, “Polka lines: Learning structured illumination and reconstruction for active stereo,” in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753–5763" represents the dual-camera perspective simulated structured light image The contribution from the projector light to the brightness of; is Gaussian noise, and , is the brightness clipping function, which limits the pixel values to an appropriate range.

[0086] The schematic diagram of the disparity attention depth estimation module is as shown in Figure 3 , including a feature extraction sub-module, a cost volume construction sub-module, a cost aggregation sub-module, and a disparity prediction sub-module.

[0087] The feature extraction sub-module is as shown in Figure 4 . The feature extraction sub-module consists of two resnet branches. One feature extraction sub-module receives the camera perspective structured light images from the left camera perspective and the right camera perspective as inputs respectively, and the other feature extraction sub-module receives the camera perspective RGB images from the left camera perspective and the right camera perspective as inputs respectively. After the camera perspective RGB image is input into the feature extraction sub-module, the features and , the number of channels is 320 and the resolution is one-fourth of the input. The camera-view structured light image input is a single-channel grayscale image. After the corresponding feature extraction sub-module, there are two additional convolutional layers to compress the 320-channel feature map into a 32-channel feature and .

[0088] The cost volume construction sub-module constructs a concat cost volume and a corr cost volume. The unary features extracted from the camera-view structured light image of the left camera view and the camera-view structured light image of the right camera view are respectively denoted as and , and their number of channels is 32. The concat cost volume and is constructed by concatenating each disparity level of :

[0089]

[0090] where (x, y) represents the position of the pixel and d is different disparity levels. Concat is the concatenation operation. The size of .

[0091] The unary features extracted from the camera-view RGB image of the left camera view and the camera-view RGB image of the right camera view are denoted as and , and the number of channels is 320. These two unary features are divided into feature groups along the channel dimension, = 40, and each feature group has channels. The th feature group is denoted as and , then the corr cost volume can be obtained by the following formula:

[0092]

[0093] where represents the inner product of matrices.

[0094] The hourglass module is used to extract the first disparity attention weight from the corr cost volume , and extract the second disparity attention weight from the concat cost volume ;

[0095] The hourglass module is divided into a downsampling part and an upsampling part. In the downsampling part, a 3D convolutional layer is first used to extract preliminary features, and then 3D convolutional layers are continuously used to gradually reduce the spatial resolution and disparity resolution of the extracted preliminary features. In the upsampling part, 3D transposed convolutional layers are used to gradually restore the spatial resolution and disparity resolution, and then a 3D convolutional layer is used to obtain the output. The hourglass module gradually reduces the resolution in the downsampling part to form a "narrow part", and then restores the feature map to the original resolution in the upsampling part to form a "wide part". This process of first shrinking and then expanding is overall similar to the shape of a traditional hourglass.

[0096] Filter the corr cost volume and the concat cost volume through the first disparity attention weight and the second disparity attention weight to obtain the Concat attention cost volume and the Corr attention cost volume :

[0097]

[0098]

[0099] where is and the number of channels of; Finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume :

[0100] .

[0101] The cost aggregation sub-module is as Figure 5 shown. First, four 3D convolutions are used to preliminarily aggregate the combined attention cost volume. Then, two hourglass modules are used to learn more context information through an encoder-decoder structure, and each hourglass module consists of 4 3D convolutions and 2 3D transposed convolutions.

[0102] The disparity prediction sub-module uses disparity regression to obtain the predicted disparity map from the output of the cost aggregation sub-module. First, two 3D convolutions are used to generate a 4D volume with 1 channel, and then the 4D volume is upsampled to the same size as the original input. The weight for each disparity level d is calculated through a softmax operation from the aggregated cost , and then the disparity value is calculated using the following formula:

[0103]

[0104] where is the maximum possible value of the disparity level. The disparity value Convert to depth value through the following formula :

[0105]

[0106] where B is the baseline length of the binocular camera and f is the focal length of the binocular camera.

[0107] The loss function of the disparity attention depth estimation module is as follows:

[0108]

[0109] The loss function compares the predicted depth value with the true depth value to evaluate the accuracy of the predicted depth. The method of the present invention improves the depth estimation accuracy, enhances the scene adaptability, reduces the simulation-reality gap, and improves the real-time performance and efficiency.

[0110] The present invention has achieved significantly better depth estimation accuracy than the existing methods. For example, the metrics EPE is 0.34px and D1 is 1.57%, both of which are better than the existing methods. The specific metric comparison is shown in the following table:

[0111] Table 1 shows the quantitative comparison results between the present invention and the existing methods

[0112]

[0113] In Table 1, GD is the method proposed by “S. Schreiberhuber, J.-B. Weibel, T. Patten, and M. Vincze, “Gigadepth: Learning depth from structured light with branching neural networks,” in Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Ciss´e, G. M. Farinella, and T. Hassner, Eds. Springer Nature Switzerland, 2022, pp. 214–229.”; Polkaline is the method proposed by “S.-H. Baek and F. Heide, “Polka lines: Learning structured illumination and reconstruction for active stereo,” in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753– 5763.”; DIS is the method proposed by “M. M. Johari, C. Carta, and F. Fleuret, “Depth in space: Exploitation and fusion of multiple video frames for structured-light depth estimation,” in 2021 IEEE / CVF International Conference on Computer Vision (ICCV), Feb. 2021, pp. 6019–6028.”; IGEV-Stereo is the method proposed by “G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encoding volume for stereo matching,” in 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Aug. 2023, pp. 21 919–21 928."The proposed method; Selective-RAFT is the method proposed by 'X. Wang, G. Xu, H. Jia, and X. Yang, "Selective-stereo: Adaptive frequency information selection for stereo matching," in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 19 701–19 710.'"

[0114] The adaptive structured light generation module enables the end-to-end adaptive active stereo vision system to adjust the structured light pattern in real time according to different scenes, showing better performance in complex scenes such as low light and high dynamic range. Compared with the fixed-pattern structured light method, the error is greatly reduced.

[0115] The present invention decomposes the structured light pattern into two parts: a local texture pattern and a global brightness pattern. After being independently generated by different branches, they are multiplied element by element and fused. The local texture pattern provides highly discriminative local features for binocular matching, while the global brightness pattern enables the structured light to dynamically adapt to the scene illumination and object reflectivity, avoiding the matching failure of the fixed brightness pattern in high-light regions / low-light regions. Different from the traditional fixed structured light pattern design strategies (such as random speckles, Gray codes), the strategy of the present invention enables the structured light to have both local matching uniqueness and global brightness adaptability.

[0116] The present invention simultaneously considers the scene information contained in the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective. The former provides scene illumination and reflectivity information, and the latter provides geometric structure prior information. Then, the two are used to generate an adaptive structured light projection pattern suitable for the specific measurement scene through an encoder-decoder network. The present invention integrates the scene geometric information (disparity) into the structured light generation, enabling the structured light projection pattern generation to have the ability of scene depth perception, automatically enhancing the local texture density in the depth mutation region, and adjusting the global brightness to compensate for the light intensity attenuation in the far-distance region, solving the defect that the traditional method ignores the scene geometric characteristics.

[0117] The physically-aware differential imaging model reduces the difference between the simulation model and the actual system through hardware-model joint optimization. Compared with the pure simulation model, the EPE of this method is reduced by 37%, making the trained model more reliable in practical applications.

[0118] The present invention directly adds a real projector and a binocular camera to the forward propagation link, avoiding the optical errors caused by the existing differential model relying solely on pure simulation. The projector projects the structured light projection pattern output by the adaptive structured light generation module, and the camera captures the structured light image containing real defocus, noise, and ambient light interference, ensuring that the training data is consistent with the actual application scenario.

[0119] The optical simulation model designed in the present invention combines a hybrid reverse model of Fourier optics (simulation projector sub-model) and geometric optics (simulation camera sub-model), and for the first time embeds physical optical principles (such as diffraction and refraction) into backpropagation, avoiding the high computational complexity of the finite difference method while ensuring the analyticity and efficiency of gradient calculation. At the projector end, the point spread function is calculated based on Fourier optics to simulate the diffraction effect of light passing through the projector lens. At the camera end, based on the image transformation formula of geometric optics, the structured light pattern in the projector view is transformed into the left camera view and the right camera view using the coarse disparity map of the camera view, and the simulated structured light image of the dual camera view is rendered by combining the Lambertian reflection model.

[0120] Through the optimized network structure of the adaptive structured light generation module and the disparity attention depth estimation module, the computational complexity is reduced, and the real-time performance and efficiency of the end-to-end adaptive active stereo vision system are improved, which can better meet the requirements of actual applications.

[0121] The present invention constructs two different cost volumes according to the feature characteristics of the structured light image and the RGB image in the camera view. The concat cost volume is constructed using the features of the structured light image in the camera view, and the strong discriminability of the artificial texture is used to provide absolute depth cues. The corr cost volume is constructed using the features of the RGB image in the camera view to capture the continuity information of the natural texture and solve the matching interference problem of the structured light image in the camera view in the rich texture area. Different from the traditional method of directly splicing the structured light image and the RGB image at the input or only using the structured light image, the cost volume construction method proposed in the present invention realizes the complementary advantages of the strong discriminability of the artificial texture and the continuity of the natural texture.

[0122] The present invention adopts a bidirectional disparity attention mechanism to achieve the mutual enhancement of the two cost volumes. The disparity attention weights are extracted from the two cost volumes through the hourglass module structure and cross-filtered. The second disparity attention weight generated by the concat cost volume is used to enhance the corr cost volume and suppress the redundant artificial features of the corr cost volume in the texture-rich area; the first disparity attention weight generated by the corr cost volume is used to enhance the concat cost volume and compensate for the insufficient matching ability of the concat cost volume in the weak texture area. Finally, the two filtered cost volumes are merged into a combined attention cost volume, improving the expression ability of the cost volume.

[0123] The proposed end-to-end adaptive active stereo vision system realizes the end-to-end joint learning of structured light generation and depth estimation. Through the physically aware differentiable imaging module, the depth loss gradient is simultaneously backpropagated to the adaptive structured light generation module (to adjust the structured light parameters) and the disparity attention depth estimation module (to optimize the depth estimation network), realizing the collaborative optimization of the two, rather than separately and independently training. Compared with the existing methods that first design the structured light projection pattern and then train the depth estimation network, the present invention avoids the limitations of independent training of each module in the existing methods and improves the overall performance of the end-to-end adaptive active stereo vision system.

Claims

1. A depth estimation method based on adaptive structured light, characterized in that Design an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module; Collect the RGB image from the camera's perspective in the scenario where the projector is turned off, and obtain the coarse disparity map from the camera's perspective through the semi-global matching algorithm; The RGB image from the camera's perspective and the coarse disparity map from the camera's perspective are projected and transformed to generate the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective; The RGB image from the projector's perspective and the coarse disparity map from the projector's perspective are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; The adaptive structured light projection pattern is input into the physically-aware differentiable imaging module to generate the structured light image from the camera's perspective; The disparity attention depth estimation module simultaneously receives the RGB image from the camera's perspective and the structured light image from the camera's perspective, and outputs the depth estimation result; Taking the depth error as the loss function, jointly train and optimize the learning parameters of the adaptive structured light generation module and the learning parameters of the disparity attention depth estimation module, and use the optimized end-to-end adaptive active stereo vision system for depth estimation; The physical perception differentiable imaging module includes a real optical system f ph and an optical simulation model f sim ; During the joint training and optimization process, the real optical system is used during forward propagation, and the optical simulation model is used during gradient backpropagation.

2. The depth estimation method based on adaptive structured light according to claim 1, wherein The projection transformation is based on the calibration parameters of the binocular camera and the calibration parameters of the projector, and maps the pixel points in the binocular camera coordinate system to the projector coordinate system through the geometric transformation formula.

3. The depth estimation method based on adaptive structured light according to claim 1, characterized in that Structured light encoding is performed in the adaptive structured light generation module; The structured light encoding is specifically as follows: An encoder is used to extract features from the RGB image and the coarse disparity map at the projector's view angle; the extracted features are respectively input into a SegNet decoder for local texture decoding to generate a local texture pattern P local and input into a SegNet decoder for global brightness decoding to generate a global brightness pattern P global , and an adaptive structured light projection pattern P is calculated as follows: P = P local × P global 。 4. The depth estimation method based on adaptive structured light according to claim 3, wherein The encoder consists of several convolutional layers; different sizes of convolutional kernels and different sizes of strides are used in the convolutional layers.

5. The depth estimation method based on adaptive structured light according to claim 1, wherein During forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern P into the scene. After the adaptive structured light projection pattern is reflected by the objects in the scene, the binocular camera collects and obtains the structured light image from the camera's perspective; turn off the real projector to obtain the RGB image from the camera's perspective; Represents a scene; I s Represents a structured light image from the camera's perspective; I represents an RGB image from the camera's perspective.

6. The depth estimation method based on adaptive structured light according to claim 1, wherein During backpropagation, an optical simulation model is used to calculate the loss function of the disparity attention depth estimation module Gradient estimation with respect to the adaptive structured light projection pattern P: where g p represents the loss function is the gradient estimation of the adaptive structured light projection pattern P, represents the gradient estimation of the adaptive structured light projection pattern P by the optical simulation model, and T represents the matrix transpose operation, is the structured light image of the dual-camera view simulation; The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model; The point spread function of the simulation projector sub-model is expressed as the square magnitude of the Fourier transform of the complex-valued pupil function Aexp(iφ(z)): where A is the amplitude of the structured light, λ is the wavelength of the structured light; the phase φ is a function of the scene depth z, representing the degree of defocus caused by the mismatch between the focus depth of the projector and the scene depth z of the scene point; is the Fourier transform; the output of the simulated projector sub-model is the structured light pattern P from the perspective of the projector proj , which is obtained by convolving the point spread function with the adaptive structured light projection pattern P output by the adaptive structured light generation module: where * represents 2D convolution; The simulation camera sub-model, through the coarse disparity map D of the camera view L / R Transform the structured light pattern under the projector view to the left camera view and the right camera view to obtain the illumination images under the camera view wherein is an image transformation operation; Render the output of the simulation camera sub-model based on the Lambertian reflection model to obtain the simulated structured light image of the dual-camera perspective, and the formula is as follows: where α represents the influence coefficients of camera exposure, aperture, and sensitivity on pixel brightness; is the result of grayscale conversion of the RGB image of the camera's perspective, representing the contribution from ambient light to the brightness of the simulated structured light image of the dual-camera perspective; β is the power coefficient of the projector; ρ is the reflectivity map, representing the contribution from the projector's light to the brightness of the simulated structured light image of the dual-camera perspective; is Gaussian noise, and σ = 0.004, τ is the brightness clipping function.

7. The depth estimation method based on adaptive structured light according to claim 1, wherein The parallax attention depth estimation module includes a feature extraction sub-module, a cost volume construction sub-module, a cost aggregation sub-module, and a disparity prediction sub-module; the RGB image from the camera view and the structured light image from the camera view are respectively processed by the feature extraction sub-module to extract corresponding features, obtaining the corresponding feature f l and f r of the RGB image from the camera view, as well as the corresponding features and of the structured light image from the camera view. The corresponding features are jointly input into the cost volume construction sub-module to obtain a combined attention cost volume, and then sequentially passed through the cost aggregation sub-module and the disparity prediction sub-module to output the depth estimation result.

8. The depth estimation method based on adaptive structured light according to claim 7, wherein In the cost volume construction sub-module, splice and for each disparity level to construct the concatenated cost volume C concat : where (x, y) represents and the pixel positions, d is the disparity level, and Concat is the concatenation operation; f l and f r is divided into Ng feature groups along the channel dimension, each feature group has N / Ng channels, where N is the total number of channels of the features; the g-th feature group is denoted as f l g and the corr cost volume C corr is obtained through the following formula: where <·,·〉 represents the matrix inner product; The hourglass module is used to extract the first disparity attention weight W from the corr cost volume C corr and extract the second disparity attention weight W from the concat cost volume C corr ; concat concat ;​ The hourglass module is divided into a downsampling part and an upsampling part. The downsampling part first passes through a 3D convolutional layer to extract preliminary features, and then continuously uses 3D convolutional layers to gradually reduce the spatial resolution and disparity resolution of the extracted preliminary features; The upsampling part uses 3D transposed convolutional layers to gradually restore the spatial resolution and disparity resolution, and then passes through a 3D convolutional layer to obtain the output; Filter the corr cost volume C through the first disparity attention weight and the second disparity attention weight corr and the concat cost volume C concat to obtain the Concat attention cost volume C' concat and the Corr attention cost volume C' corr : where N s is and the number of channels; finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume C: C = Concat{C′ corr , C′ concat}.

9. The depth estimation method based on adaptive structured light according to claim 8, characterized in that, The cost aggregation sub-module is connected to the disparity prediction sub-module; the cost aggregation sub-module first uses four 3D convolutional layers to preliminarily aggregate the combined attention cost volumes, and then uses two hourglass modules to learn context information through the encoder-decoder structure to obtain the aggregated cost; The parallax prediction sub-module obtains a predicted disparity map from the aggregated cost using disparity regression; the weight of disparity level d is calculated from the aggregated cost c through the softmax operation δ(·), and then the disparity value is calculated using the following formula: where D max is the maximum value of the parallax level; the parallax value is converted to the depth value D: where B is the baseline length of the binocular camera, and f is the focal length of the binocular camera; Loss function of the parallax attention depth estimation module As follows: D gt is the true depth value.

Citation Information

Patent Citations

  • Fast stereo matching algorithm for adaptive iterative residual optimization

    CN115797674A

  • Infrared and visible light image fusion method for inspection of power transmission and transformation equipment

    CN119540702A