Depth estimation method based on adaptive structured light
By designing an adaptive structured light generation module, a physically perceived differentiable imaging module and a parallax attention depth estimation module in an active stereo vision system, the problem that the fixed structured light mode cannot adapt to scene changes is solved, the high accuracy and robustness of depth estimation is achieved, and the simulation-reality gap is reduced.
Patent Information
- Application Number
- CN202510621767.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The fixed structured light mode in the existing active stereo vision system cannot dynamically adapt to scene changes, the depth estimation network does not fully integrate multimodal information, and there is a gap between simulation and reality.
Design an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically-aware differentiable imaging module and a parallax attention depth estimation module, to achieve dynamic adjustment of structured light and depth estimation through collaborative optimization.
The dynamic adaptation of structured light mode is achieved, the accuracy and robustness of depth estimation are improved, the simulation-reality gap is reduced, and the overall performance of the system is improved.
Smart Images

Figure CN120125633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of stereoscopic vision technology, and in particular to a depth estimation method based on adaptive structured light. Background Art
[0002] Active stereo vision technology combines structured light projection with binocular cameras, showing significant advantages in the field of 3D environment perception, and is widely used in autonomous driving, robot navigation, and 3D scene reconstruction. Unlike passive depth sensors that rely on the object's own texture, active stereo systems project specially designed structured light patterns onto the surface of objects, providing additional corresponding clues for binocular matching, thereby improving the stability and accuracy of depth estimation.
[0003] Most of the existing active stereo techniques use fixed structured light projection patterns. Some techniques design high-frequency structured light based on imaging theory to combat the effects of global illumination and defocus. For example, "M. Gupta and N. Nakhate, "Ageometric perspective on structured light coding," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 87–102." proposed a new paradigm of structured light coding based on geometric analysis, Hamiltonian coding. This technology uses coding curves as the geometric representation of structured light coding, maps the K-dimensional projection pattern into a continuous trajectory in the hypercube space, and then establishes a negative correlation between the length of the coding curve and the decoding error by deriving a proxy metric to optimize the structured light coding design. However, this type of method has the defect of scene adaptability, which is specifically manifested in the lack of a scene dynamic feedback correction mechanism. Other technical solutions propose data-driven structured light design methods to optimize structured light patterns through scene data. For example, "S.-H. Baek and F. Heide, "Polkalines: Learning structured illumination and reconstruction for active stereo," in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753–5763." proposed a customized stripe projection pattern that can learn to adapt to different environments (indoor, outdoor, and general scenes). By inputting a large amount of scene data to optimize the structured light end-to-end, the depth estimation effect of high-dynamic scenes is improved. However, this type of method uses an offline training mode, which leads to parameter solidification. After the structured light design is completed, it cannot be optimized again, so it cannot adapt to new scenes online.
[0004] In terms of depth estimation networks, current technologies typically only use structured light patterns as input and rely on the artificial textures of structured light images for matching. For example, "Y. Zhang, S. Khamis, C. Rhemann, J. Valentin, A. Kowdle, V. Tankovich, M. Schoenberg, S. Izadi, T. Funkhouser, and S. Fanello, 'Active stereonet: End-to-end self-supervised learning for active stereo systems,' in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 784–801." proposed a structured light stereo matching network that estimates depth end-to-end from a pair of structured light inputs, improving the performance of traditional methods in noisy, low-texture, and occluded regions to a certain extent. However, such methods may introduce noise in regions with rich natural textures and ignore the rich texture information in natural images, limiting the robustness of depth estimation. Summary of the Invention
[0005] In view of the problems in the existing active stereo vision system, such as the fixed structured light pattern being unable to dynamically adapt to the scene, the depth estimation network not fully integrating multimodal information, and the gap between simulation and reality, the present invention proposes a depth estimation method based on adaptive structured light, designs an end-to-end adaptive active stereo vision system, and realizes the collaborative optimization of structured light and depth estimation through an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module, improving the depth perception performance in complex scenes.
[0006] The technical solution of the present invention is as follows: A depth estimation method based on adaptive structured light designs an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module;
[0007] Collect the RGB image from the camera's perspective in the scenario of turning off the projector, and obtain the coarse disparity map from the camera's perspective through the semi-global matching algorithm; the RGB image from the camera's perspective and the coarse disparity map from the camera's perspective are used to generate the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective through projective transformation; the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; the adaptive structured light projection pattern is used to generate the structured light image from the camera's perspective through the physically-aware differentiable imaging module; the disparity attention depth estimation module simultaneously receives the RGB image from the camera's perspective and the structured light image from the camera's perspective, and outputs the depth estimation result; using the depth error as the loss function, jointly train and optimize the learning parameters of the adaptive structured light generation module and the learning parameters of the disparity attention depth estimation module, and use the optimized end-to-end adaptive active stereo vision system for depth estimation.
[0008] The projective transformation is based on the calibration parameters of the binocular camera and the calibration parameters of the projector, and maps the pixel points in the binocular camera coordinate system to the projector coordinate system through the geometric transformation formula.
[0009] Structured light encoding is performed in the adaptive structured light generation module;
[0010] The specific structured light encoding is as follows: an encoder is used to extract features from the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective; the extracted features are respectively used by the SegNet decoder for local texture decoding to generate local texture patterns and the SegNet decoder for global brightness decoding to generate global brightness patterns , and the adaptive structured light projection pattern is calculated :
[0011] .
[0012] The encoder consists of several convolutional layers; the convolutional layers use convolutional kernels of different sizes and different strides.
[0013] The physically-aware differentiable imaging module includes a real optical system and an optical simulation model ; during the joint training and optimization process, the real optical system is used during forward propagation, and the optical simulation model is used during gradient backpropagation.
[0014] During forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern into the scene. After the adaptive structured light projection pattern is reflected by the objects in the scene, the binocular camera collects and obtains the structured light image from the camera's perspective; turn off the real projector and obtain the RGB image from the camera's perspective;
[0015]
[0016] Represents a scene; Represents the structured light image from the camera's perspective; Represents the RGB image from the camera's perspective.
[0017] During backpropagation, an optical simulation model is used to calculate the loss function of the disparity attention depth estimation module With respect to the adaptive structured light projection pattern Gradient estimation:
[0018]
[0019] Where Represents the loss function Gradient estimation of the adaptive structured light projection pattern Gradient estimation, Represents the gradient estimation of the adaptive structured light projection pattern by the optical simulation model Gradient estimation, Represents matrix transpose operation, Is the simulated structured light image of the dual-camera perspective;
[0020] The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model;
[0021] The point spread function of the simulation projector sub-model is expressed as the squared magnitude of the Fourier transform of the complex-valued pupil function Of:
[0022]
[0023] Where, Is the amplitude of the structured light, Is the wavelength of the structured light; the phase Is a function of the scene depth Characterizing the defocus degree caused by the mismatch between the projector focus depth and the scene depth of the scene point Mismatch; Is the Fourier transform; the output of the simulation projector sub-model is the structured light pattern from the projector's perspective Generated by convolving the point spread function with the adaptive structured light projection pattern output by the adaptive structured light generation module Obtained:
[0024]
[0025] Where Represents 2D convolution;
[0026] The simulation camera sub-model, through the coarse disparity map from the camera's perspective Transform the structured light pattern from the perspective of the projector to the perspectives of the left and right cameras to obtain the illumination images from the camera perspectives. :
[0027]
[0028] Where is the image transformation operation;
[0029] Render the output of the simulated camera sub-model based on the Lambertian reflection model to obtain the simulated structured light images from the dual-camera perspectives. , and the formula is as follows:
[0030]
[0031] Where represents the influence coefficients of camera exposure, aperture, and sensitivity on pixel brightness; is the result of graying the RGB image from the camera perspective, representing the contribution from ambient light to the brightness of the simulated structured light images from the dual-camera perspectives ; is the power coefficient of the projector; is the reflectivity map, obtained from the RGB image from the camera perspective by the RGB-inversion method, representing the contribution from the projector light to the brightness of the simulated structured light images from the dual-camera perspectives ; is Gaussian noise, and , is the brightness clipping function.
[0032] The disparity attention depth estimation module includes a feature extraction sub-module, a cost volume construction sub-module, a cost aggregation sub-module, and a disparity prediction sub-module; the RGB image from the camera perspective and the structured light image from the camera perspective are respectively processed by the feature extraction sub-module to extract corresponding features, obtaining the corresponding features and of the RGB image from the camera perspective, as well as the corresponding features and of the structured light image from the camera perspective; the corresponding features are jointly input into the cost volume construction sub-module to obtain a combined attention cost volume, and then successively through the cost aggregation sub-module and the disparity prediction sub-module to output the depth estimation result.
[0033] In the cost volume construction sub-module, each disparity level of and is concatenated to construct the concat cost volume :
[0034]
[0035] Among them denotes and the pixel positions, d is the disparity level, and Concat is the splicing operation;
[0036] and are divided into feature groups along the channel dimension, each feature group having channels, is the total number of feature channels; Denote the th feature group as and , and the corr cost volume is obtained through the following formula:
[0037]
[0038] Among them denotes the inner product of matrices;
[0039] The hourglass module is used to extract the first disparity attention weight from the corr cost volume , and extract the second disparity attention weight from the concat cost volume ;
[0040] The hourglass module is divided into a downsampling part and an upsampling part. The downsampling part first passes through a 3D convolutional layer to extract preliminary features, and then continuously uses 3D convolutional layers to gradually reduce the spatial resolution and disparity resolution of the extracted features; The upsampling part uses 3D transposed convolutional layers to gradually restore the spatial resolution and disparity resolution, and then passes through a 3D convolutional layer to obtain the output; The hourglass module gradually reduces the resolution in the downsampling part to form a "narrow part", and then restores the feature map to the original resolution in the upsampling part to form a "wide part";
[0041] Filter the corr cost volume and the concat cost volume through the first disparity attention weight and the second disparity attention weight to obtain the Concat attention cost volume and the Corr attention cost volume :
[0042]
[0043]
[0044] Among them is and The number of channels; finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume :
[0045] 。
[0046] The cost aggregation sub-module is connected to the disparity prediction sub-module; the cost aggregation sub-module first uses four 3D convolutions to preliminarily aggregate the combined attention cost volume, and then uses two hourglass modules to learn context information through an encoder-decoder structure to obtain the aggregated cost;
[0047] The disparity prediction sub-module uses disparity regression to obtain the predicted disparity map from the aggregated cost; the weight of disparity level d is obtained through a softmax operation from the aggregated cost and then the disparity value is calculated using the following formula:
[0048]
[0049] where is the maximum value of the disparity level; the disparity value is converted to the depth value :
[0050]
[0051] where B is the baseline length of the binocular camera and f is the focal length of the binocular camera;
[0052] The loss function of the disparity attention depth estimation module is as follows:
[0053]
[0054] 。
[0055] Advantages of the present invention: The present invention can dynamically generate an optimal structured light pattern according to the real-time information of the scene, solving the problem that traditional fixed-pattern structured light cannot adapt to scene changes.
[0056] The present invention realizes the collaborative optimization of hardware-model by combining data acquisition through a real optical system and an optical simulation model, effectively reducing the simulation-reality gap.
[0057] The disparity attention depth estimation module of the present invention adopts a bidirectional disparity attention mechanism, fully integrating the information of the structured light image and the RGB image, improving the accuracy and robustness of depth estimation.
[0058] The present invention realizes end-to-end joint learning of structured light generation and depth estimation, avoids the limitations of independent training of each module in traditional methods, and improves the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the overall technical route of the present invention;
[0060] Figure 2 is a schematic diagram of the adaptive structured light generation module;
[0061] Figure 3 is a schematic diagram of the disparity attention depth estimation module;
[0062] Figure 4 is a schematic diagram of the feature extraction sub-module;
[0063] Figure 5 is a schematic diagram of the cost aggregation sub-module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The overall technical solution of the present invention is as Figure 1 shown. An end-to-end adaptive active stereo vision system is designed, which mainly includes three parts: an adaptive structured light generation module, a physically-aware differentiable imaging module, and a disparity attention depth estimation module. The RGB image of the camera view in the scene with the projector turned off is collected, and the coarse disparity map of the camera view is obtained through the semi-global matching algorithm; the RGB image of the camera view and the coarse disparity map of the camera view are projected and transformed to generate the RGB image of the projector view and the coarse disparity map of the projector view, which are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; the adaptive structured light projection pattern is generated into the structured light image of the camera view through the physically-aware differentiable imaging module; the disparity attention depth estimation module simultaneously receives the RGB image of the camera view and the structured light image of the camera view and outputs the depth estimation result; taking the depth error as the loss function, the learning parameters of the adaptive structured light generation module and the learning parameters of the disparity attention depth estimation module are jointly trained and optimized, and finally the optimized end-to-end adaptive active stereo vision system is used for depth estimation.
[0065] The adaptive structured light generation module is as Figure 2As shown, the adaptive structured light generation module first uses an encoder to extract features from the input. The encoder consists of a series of convolutional layers to capture the multi-scale feature information of the image. These convolutional layers use convolutional kernels and strides of different sizes to gradually reduce the size of the feature map while increasing the dimension of the features. Then, the SegNet decoder for local texture decoding "V. Badrinarayanan, A. Kendall, and R. Cipolla, "Segnet: A deep convolutional encoder-decoder architecture for image segmentation," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 12, pp. 2481–2495, 2017." and the SegNet decoder for global brightness decoding respectively generate the local texture pattern and the global brightness pattern , and finally the adaptive structured light projection pattern is calculated through the following formula :
[0066]
[0067] The physically aware differentiable imaging module contains a real optical system and an optical simulation model . For this physically aware differentiable imaging module, during the overall framework training process, the real optical system is used during forward propagation, and the optical simulation model is used during backpropagation.
[0068] During forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern into the scene. After being reflected by the objects in the scene, the adaptive structured light projection pattern is captured by the binocular camera to obtain the structured light image from the camera's perspective ; the real projector is turned off to obtain the RGB image from the camera's perspective ;
[0069]
[0070] represents the scene; represents the structured light image from the camera's perspective; represents the RGB image from the camera's perspective.
[0071] During backpropagation, the optical simulation model is used to calculate the loss function of the disparity attention depth estimation module with respect to the adaptive structured light projection pattern Gradient estimation of:
[0072]
[0073] where represents the loss function for the gradient estimation of the adaptive structured light projection pattern ; represents the gradient estimation of the adaptive structured light projection pattern by the optical simulation model ; represents the matrix transpose operation; is the structured light image of the dual-camera view simulation;
[0074] The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model;
[0075] The point spread function of the simulation projector sub-model is expressed as the squared magnitude of the Fourier transform of the complex-valued pupil function :
[0076]
[0077] where is the amplitude of the structured light, is the wavelength of the structured light; the phase is a function of the scene depth characterizing the defocus degree caused by the mismatch between the projector focus depth and the scene depth of the scene point ; is the Fourier transform; the output of the simulation projector sub-model is the structured light pattern from the projector view obtained by convolving the point spread function with the adaptive structured light projection pattern output by the adaptive structured light generation module :
[0078]
[0079] where represents 2D convolution;
[0080] For the simulation camera sub-model, the structured light pattern from the projector view is transformed to the left camera view and the right camera view through the coarse disparity map of the camera view to obtain the illumination image in the camera view :
[0081]
[0082] where is the image transformation operation;
[0083] Rendering based on the Lambertian reflection model to obtain the output of the simulation camera sub-model, and obtaining the dual-camera perspective simulated structured light images under the left camera view and the right camera view , the formula is as follows:
[0084]
[0085] where represents the influence coefficients of camera exposure, aperture, and sensitivity on pixel brightness; is the result of grayscale conversion of the camera-view RGB image, representing the dual-camera perspective simulated structured light image the contribution from ambient light to the brightness of; is the power coefficient of the projector; is the reflectivity map, obtained from the camera-view RGB image by the RGB-inversion method “S.-H. Baek and F. Heide, “Polka lines: Learning structured illumination and reconstruction for active stereo,” in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753–5763” represents the dual-camera perspective simulated structured light image the contribution from the projector light to the brightness of; is Gaussian noise, and , is the brightness clipping function, which limits the pixel value to an appropriate range.
[0086] The schematic diagram of the disparity attention depth estimation module is as shown in Figure 3 , including a feature extraction sub-module, a cost volume construction sub-module, a cost aggregation sub-module, and a disparity prediction sub-module.
[0087] The feature extraction sub-module is as shown in Figure 4 . The feature extraction sub-module consists of two resnet branches. One feature extraction sub-module respectively receives the camera-view structured light images under the left camera view and the right camera view as inputs, and the other feature extraction sub-module respectively receives the camera-view RGB images under the left camera view and the right camera view as inputs. After the camera-view RGB image is input into the feature extraction sub-module, the features and , the number of channels is 320 and the resolution is one-fourth of the input. The input of the camera-view structured light image is a single-channel grayscale image. After the corresponding feature extraction sub-module, there are two additional convolutional layers to compress the 320-channel feature map into a 32-channel feature and .
[0088] The cost volume construction sub-module constructs a concat cost volume and a corr cost volume. The unary features extracted from the camera-view structured light image of the left camera view and the camera-view structured light image of the right camera view are respectively denoted as and , and their number of channels is 32. The concat cost volume and is constructed by concatenating each disparity level of :
[0089]
[0090] where (x, y) represents the position of the pixel and d is different disparity levels. Concat is the concatenation operation. The size of .
[0091] The unary features extracted from the camera-view RGB image of the left camera view and the camera-view RGB image of the right camera view are denoted as and , and the number of channels is 320. These two unary features are divided into feature groups along the channel dimension, = 40, and each feature group has channels. The th feature group is denoted as and , then the corr cost volume can be obtained by the following formula:
[0092]
[0093] where represents the matrix inner product.
[0094] The hourglass module is used to extract the first disparity attention weight from the corr cost volume , and the second disparity attention weight from the concat cost volume ;
[0095] The hourglass module is divided into a downsampling part and an upsampling part. In the downsampling part, a 3D convolutional layer is first used to extract preliminary features, and then 3D convolutional layers are continuously used to gradually reduce the spatial resolution and disparity resolution of the extracted preliminary features. In the upsampling part, 3D transposed convolutional layers are used to gradually restore the spatial resolution and disparity resolution, and then a 3D convolutional layer is used to obtain the output. The hourglass module gradually reduces the resolution in the downsampling part to form a "narrow part", and then restores the feature map to the original resolution in the upsampling part to form a "wide part". This process of first contracting and then expanding is overall similar to the shape of a traditional hourglass.
[0096] Filter the corr cost volume and the concat cost volume through the first disparity attention weight and the second disparity attention weight to obtain the Concat attention cost volume and the Corr attention cost volume :
[0097]
[0098]
[0099] where is and the number of channels of; finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume :
[0100] .
[0101] The cost aggregation sub-module is as Figure 5 shown. First, four 3D convolutions are used to preliminarily aggregate the combined attention cost volume. Then, two hourglass modules are used to learn more context information through an encoder-decoder structure, and each hourglass module consists of 4 3D convolutions and 2 3D transposed convolutions.
[0102] The disparity prediction sub-module uses disparity regression to obtain the predicted disparity map from the output of the cost aggregation sub-module. First, two 3D convolutions are used to generate a 4D volume with 1 channel, and then the 4D volume is upsampled to the same size as the original input. The weight for each disparity level d is calculated through a softmax operation from the aggregated cost , and then the disparity value is calculated using the following formula:
[0103]
[0104] where is the maximum possible value of the disparity level. The disparity value Convert to depth value through the following formula :
[0105]
[0106] where B is the baseline length of the binocular camera and f is the focal length of the binocular camera.
[0107] The loss function of the disparity attention depth estimation module is as follows:
[0108]
[0109] The loss function compares the predicted depth value with the true depth value to evaluate the accuracy of the predicted depth. The method of the present invention improves the depth estimation accuracy, enhances the scene adaptability, reduces the simulation-reality gap, and improves the real-time performance and efficiency.
[0110] The present invention has achieved significantly better depth estimation accuracy than the existing methods. For example, the metrics EPE is 0.34px and D1 is 1.57%, both of which are better than the existing methods. The specific metric comparisons are shown in the following table:
[0111] Table 1 shows the quantitative comparison results between the present invention and the existing methods
[0112]
[0113] In Table 1, GD is the method proposed by "S. Schreiberhuber, J.-B. Weibel, T. Patten, and M. Vincze, “Gigadepth: Learning depth from structured light with branching neural networks,” in Computer Vision – ECCV 2022, S. Avidan, G. Brostow, M. Ciss´e, G. M. Farinella, and T. Hassner, Eds. Springer Nature Switzerland, 2022, pp. 214–229."; Polkaline is the method proposed by "S.-H. Baek and F. Heide, “Polka lines: Learning structured illumination and reconstruction for active stereo,” in 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5753– 5763."; DIS is the method proposed by "M. M. Johari, C. Carta, and F. Fleuret, “Depth in space: Exploitation and fusion of multiple video frames for structured-light depth estimation,” in 2021 IEEE / CVF International Conference on Computer Vision (ICCV), Feb. 2021, pp. 6019–6028."; IGEV-Stereo is the method proposed by "G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encoding volume for stereo matching,” in 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Aug. 2023, pp. 21 919–21 928."The proposed method; Selective-RAFT is the method proposed by 'X. Wang, G. Xu, H. Jia, and X. Yang, "Selective-stereo: Adaptive frequency information selection for stereo matching," in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 19 701–19 710.'."
[0114] The adaptive structured light generation module enables the end-to-end adaptive active stereo vision system to adjust the structured light pattern in real time according to different scenes, showing better performance in complex scenes such as low light and high dynamic range. Compared with the fixed-pattern structured light method, the error is greatly reduced.
[0115] The present invention decomposes the structured light pattern into two parts: a local texture pattern and a global brightness pattern. After being independently generated by different branches, they are fused by element-wise multiplication. The local texture pattern provides highly discriminative local features for binocular matching, while the global brightness pattern enables the structured light to dynamically adapt to the scene illumination and object reflectivity, avoiding the matching failure of the fixed brightness pattern in high-light / low-light regions. Different from the traditional fixed structured light pattern design strategies (such as random speckles, Gray codes), the strategy of the present invention enables the structured light to have both local matching uniqueness and global brightness adaptability.
[0116] The present invention simultaneously considers the scene information contained in the RGB image from the projector's perspective and the coarse disparity map from the projector's perspective. The former provides scene illumination and reflectivity information, and the latter provides geometric structure prior information. Then, the two are used to generate an adaptive structured light projection pattern suitable for the specific measurement scene through an encoder-decoder network. The present invention integrates the scene geometric information (disparity) into the structured light generation, enabling the structured light projection pattern generation to have the ability to perceive scene depth, automatically enhancing the local texture density in the depth mutation region, and adjusting the global brightness to compensate for the light intensity attenuation in the far-distance region, solving the defect that traditional methods ignore the geometric characteristics of the scene.
[0117] The physically aware differential imaging model reduces the difference between the simulation model and the actual system through hardware-model joint optimization. Compared with the pure simulation model, the EPE of this method is reduced by 37%, making the trained model more reliable in practical applications.
[0118] The present invention directly adds a real projector and a binocular camera to the forward propagation link, avoiding the optical errors caused by the existing differential model relying on pure simulation. The projector projects the structured light projection pattern output by the adaptive structured light generation module, and the camera captures the structured light image containing real defocus, noise, and ambient light interference, ensuring that the training data is consistent with the actual application scenario.
[0119] The optical simulation model designed by the present invention combines a hybrid reverse model of Fourier optics (simulated projector sub-model) and geometric optics (simulated camera sub-model), and for the first time embeds physical optical principles (such as diffraction and refraction) into the backpropagation, avoiding the high computational complexity of the finite difference method, while ensuring the analyticity and efficiency of gradient calculation. At the projector end, the point spread function is calculated based on Fourier optics to simulate the diffraction effect of light passing through the projector lens. At the camera end, based on the image transformation formula of geometric optics, the structured light pattern in the projector view is transformed into the left camera view and the right camera view using the coarse disparity map of the camera view, and the simulated structured light images of the two camera views are rendered by combining the Lambertian reflection model.
[0120] Through the optimized network structure of the adaptive structured light generation module and the disparity attention depth estimation module, the computational complexity is reduced, and the real-time performance and efficiency of the end-to-end adaptive active stereo vision system are improved, which can better meet the requirements of actual applications.
[0121] The present invention constructs two different cost volumes according to the characteristics of the structured light image and the RGB image in the camera view. The concat cost volume is constructed using the features of the structured light image in the camera view, and the strong discriminability of the artificial texture is used to provide absolute depth cues. The corr cost volume is constructed using the features of the RGB image in the camera view to capture the continuity information of the natural texture and solve the matching interference problem of the structured light image in the camera view in the rich texture area. Different from the traditional method of directly splicing the structured light image and the RGB image at the input or only using the structured light image, the cost volume construction method proposed by the present invention realizes the complementary advantages of the strong discriminability of the artificial texture and the continuity of the natural texture.
[0122] The present invention adopts a bidirectional disparity attention mechanism to realize the mutual enhancement of the two cost volumes. The disparity attention weights are extracted from the two cost volumes through the hourglass module structure and cross-filtered. The second disparity attention weight generated by the concat cost volume enhances the corr cost volume, suppressing the redundant artificial features of the corr cost volume in the rich texture area; the first disparity attention weight generated by the corr cost volume enhances the concat cost volume, compensating for the insufficient matching ability of the concat cost volume in the weak texture area. Finally, the two filtered cost volumes are merged into a combined attention cost volume, improving the expression ability of the cost volume.
[0123] The proposed end-to-end adaptive active stereo vision system realizes the end-to-end joint learning of structured light generation and depth estimation. Through the physically-aware differentiable imaging module, the depth loss gradient is simultaneously backpropagated to the adaptive structured light generation module (adjusting the structured light parameters) and the disparity attention depth estimation module (optimizing the depth estimation network), achieving the collaborative optimization of the two, rather than separately and independently training. Compared with the existing method of first designing the structured light projection pattern and then training the depth estimation network, the present invention avoids the limitations of independent training of each module in the existing method and improves the overall performance of the end-to-end adaptive active stereo vision system.
Claims
1. A depth estimation method based on adaptive structured light, characterized in that: Design an end-to-end adaptive active stereo vision system, including an adaptive structured light generation module, a physically aware differentiable imaging module, and a parallax attention depth estimation module; Collect the camera perspective RGB image in the scene with the projector turned off, and obtain the camera perspective coarse disparity map through the semi-global matching algorithm; The camera perspective RGB image and the camera perspective coarse disparity map are transformed by projection to generate a projector perspective RGB image and a projector perspective coarse disparity map; The projector viewing angle RGB image and the projector viewing angle coarse disparity map are input into the adaptive structured light generation module to generate an adaptive structured light projection pattern; The adaptive structured light projection pattern generates a camera-perspective structured light image through a physically-aware differentiable imaging module; The parallax attention depth estimation module receives the camera's RGB image and the camera's structured light image at the same time, and outputs the depth estimation result; Taking depth error as the loss function, the learning parameters of the adaptive structured light generation module and the learning parameters of the parallax attention depth estimation module are jointly trained and optimized, and the optimized end-to-end adaptive active stereo vision system is used for depth estimation.
2. The depth estimation method based on adaptive structured light according to claim 1, characterized in that: The projection transformation is based on the calibration parameters of the binocular camera and the calibration parameters of the projector, and maps the pixel points in the binocular camera coordinate system to the projector coordinate system through a geometric transformation formula.
3. The depth estimation method based on adaptive structured light according to claim 1, characterized in that: The adaptive structured light generation module performs structured light encoding; The structured light encoding is specifically as follows: an encoder is used to extract features from the projector's viewing angle RGB image and the projector's viewing angle coarse disparity map; the extracted features are respectively used by a SegNet decoder for local texture decoding to generate a local texture pattern , the global brightness pattern is generated by the SegNet decoder for global brightness decoding , calculate the adaptive structured light projection pattern : 。 4. The depth estimation method based on adaptive structured light according to claim 3, characterized in that: The encoder is composed of a plurality of convolutional layers; the convolutional layers use convolution kernels of different sizes and steps of different sizes.
5. The depth estimation method based on adaptive structured light according to claim 3, characterized in that: The physical perception differentiable imaging module includes a real optical system and optical simulation models ; During the joint training and optimization process, the real optical system is used for forward propagation and the optical simulation model is used for gradient backpropagation.
6. The depth estimation method based on adaptive structured light according to claim 5, characterized in that: In the forward propagation, the real optical system uses a real projector to project the adaptive structured light projection pattern To the scene, after the adaptive structured light projection pattern is reflected by the objects in the scene, the binocular camera collects and obtains the structured light image from the camera perspective; the real projector is turned off to obtain the RGB image from the camera perspective; Indicates a scene; Represents the structured light image from the camera’s perspective; Represents the camera view RGB image.
7. The depth estimation method based on adaptive structured light according to claim 5, characterized in that: During back propagation, the optical simulation model is used to calculate the loss function of the parallax attention depth estimation module. Relative to the adaptive structured light projection pattern Gradient estimate of : in Represents the loss function Projecting patterns on adaptive structured light The gradient estimate of Representation of the optical simulation model for adaptive structured light projection pattern The gradient estimate of represents the matrix transpose operation, It is a dual-camera perspective simulated structured light image; The optical simulation model includes a simulation projector sub-model and a simulation camera sub-model; The point spread function of the simulated projector sub-model is expressed as a complex-valued pupil function The squared magnitude of the Fourier transform of : in, is the amplitude of the structured light, is the wavelength of the structured light; the phase is the scene depth A function that characterizes the scene depth due to the projector focal depth and the scene point The degree of defocus caused by the mismatch; is the Fourier transform; the output of the simulation projector sub-model is the structured light pattern at the projector's viewing angle , the adaptive structured light projection pattern output by the point spread function and the adaptive structured light generation module Convolution obtains: in, Represents 2D convolution; The simulated camera sub-model uses a rough disparity map of the camera's viewing angle Transform the structured light pattern from the projector's perspective to the left and right camera perspectives to obtain the illumination image from the camera's perspective : in is the image transformation operation; Based on the Lambert reflection model rendering, the output of the simulated camera sub-model is obtained, and the dual-camera perspective simulated structured light image is obtained. The formula is as follows: in Indicates the influence coefficient of camera exposure, aperture and sensitivity on pixel brightness; It is the result of graying the RGB image from the camera perspective, indicating the contribution of ambient light to the brightness of the dual-camera perspective simulated structured light image; is the power coefficient of the projector; is the reflectivity map, Represents the contribution of the projector light to the brightness of the dual-camera perspective simulated structured light image; is Gaussian noise, and , is the brightness clipping function.
8. The depth estimation method based on adaptive structured light according to claim 1, characterized in that: The disparity attention depth estimation module includes a feature extraction submodule, a cost volume construction submodule, a cost aggregation submodule and a disparity prediction submodule; The camera perspective RGB image and the camera perspective structured light image are respectively extracted by the feature extraction submodule to obtain the corresponding features of the camera perspective RGB image. and , and the corresponding features of the camera perspective structured light image and ; The corresponding features are jointly input into the cost volume construction submodule to obtain the combined attention cost volume, and then pass through the cost aggregation submodule and the disparity prediction submodule in sequence to output the depth estimation result.
9. The method for depth estimation based on adaptive structured light according to claim 8, characterized in that: The cost volume construction submodule is spliced and For each disparity level, construct a concat cost volume : in express and The pixel position, d is the disparity level, and Concat is the concatenation operation; and Along the channel dimension, it is divided into feature groups, each with channels, is the total number of feature channels; The first The feature group is recorded as and ,corr cost volume Obtained by the following formula: in represents the matrix inner product; Use the hourglass module to separate the corr cost volume Extracting the first disparity attention weight , from the concat cost volume Extract the second disparity attention weight ; The hourglass module is divided into a downsampling part and an upsampling part. The downsampling part first extracts preliminary features through a 3D convolution layer, and then uses 3D convolution layers continuously to gradually reduce the spatial resolution and disparity resolution of the extracted preliminary features. The upsampling part uses a 3D deconvolution layer to gradually restore the spatial resolution and disparity resolution, and then passes through a 3D convolution layer to obtain the output; Filter the corr cost volume by the first disparity attention weight and the second disparity attention weight and concat cost volume , get the Concat attention cost volume and Corr attention cost volume : in yes and The number of channels; Finally, the two attention cost volumes are concatenated to obtain the combined attention cost volume : 。 10. The depth estimation method based on adaptive structured light according to claim 9, characterized in that: The cost aggregation submodule is connected to the disparity prediction submodule; the cost aggregation submodule first uses four 3D convolutions to preliminarily aggregate the combined attention cost volume, and then uses two hourglass modules to learn context information through the encoder-decoder structure to obtain the aggregated cost; The disparity prediction submodule uses disparity regression to obtain the predicted disparity map from the aggregated cost; the weight of the disparity level d is calculated by the softmax operation From the aggregation cost Then use the following formula to calculate the disparity value: in is the maximum value of the parallax level; the parallax value Convert to depth value : Where B is the baseline length of the stereo camera, and f is the focal length of the stereo camera; Loss function of the parallax attention depth estimation module As shown below: 。
Citation Information
Patent Citations
Depth acquisition method for complex scene
CN104899882A
Depth perception method and device based on adaptive binocular structured light
CN114022529A
Fast stereo matching algorithm for adaptive iterative residual optimization
CN115797674A
Abdominal cavity reconstruction and focus positioning method, system and equipment based on binocular endoscope
CN119313824A
Infrared and visible light image fusion method for inspection of power transmission and transformation equipment
CN119540702A
Cited By
Fourier finite difference deconvolution imaging method based on inclination angle adaptive noise suppression
CN120871257A
A Fourier finite-difference deconvolution imaging method based on dip angle adaptive noise suppression
CN120871257B
Robust single-frame structured light three-dimensional imaging method and system based on neural feature decoding
CN121505174A
Binocular stereo matching depth estimation method and system for hyperspectral reconstruction feature enhancement
CN121962224A
Hyperspectral reconstruction feature enhanced binocular stereo matching depth estimation method and system
CN121962224B