Depth perception method based on joint optimization of speckle structured light and stereo matching network

The combined optimization of the differentiable speckle structured light and stereo matching network generated by Fourier transform solves the problem of structured light pattern predesign in active stereo technology, realizes high-precision depth perception, and reduces costs.

CN120451239AActive Publication Date: 2025-08-08NORTHEASTERN UNIV CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510532486.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The pre-design of structured light patterns of existing active three-dimensional technology results in the inability to effectively transmit environmental information, affecting the depth perception accuracy, and custom diffraction optical components are costly and difficult to achieve.

Method used

The speckle structured light generation scheme based on Fourier transform is used to optimize it with the dual-branch stereo matching network. The size, grayscale and density of speckle structured light can be controlled by microprocessing, and depth estimation is combined with a binocular camera and a projector to construct a micro-projection model and a dual-branch stereo matching network for joint optimization.

Benefits of technology

It realizes the optimization of structured light parameters, improves depth estimation performance, and does not require customized diffraction optical components, and has high-precision depth perception capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451239A_ABST
    Figure CN120451239A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of depth perception, and discloses a depth perception method based on joint optimization of speckle structured light and a stereo matching network. Constructing a speckle structured light generation scheme based on Fourier transform, performing parameterized characterization on the size, gray level and density of speckle structured light, and performing microprocessing on the speckle structured light generation process; constructing a micro-projection and acquisition model, and synthesizing an active three-dimensional image; constructing a double-branch stereo matching network, and taking the synthesized active stereo image pair and the RGB image pair as the input of the double-branch stereo matching network; in the training stage, combined optimization of the speckle structured light and the double-branch stereo matching network is completed; and in the test stage, the optimized speckle structured light is actually projected and collected, and the depth of an actual scene is obtained through the double-branch stereo matching network. According to the invention, the optimization of the structured light parameters is realized, the optimization of the structured light becomes possible, the active stereo image containing the structured light information is introduced into the network, and the high-precision depth estimation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of depth perception technology, and in particular to a depth perception method based on joint optimization of speckle structured light and a stereo matching network. Background Art

[0002] Depth sensing technology has become the cornerstone of the development of virtual reality, autonomous driving, facial recognition and other fields. Currently, mainstream depth sensing technologies include binocular vision, time of flight and structured light.

[0003] Binocular vision acquires depth by calculating disparity. In recent years, the development of binocular vision technology has been greatly promoted by advances in deep learning. However, there is inherent ambiguity in finding matching points in textureless and repetitively textured areas. Time-of-flight technology achieves depth perception by emitting light pulses into a scene and calculating the time difference or phase difference between their return pulses. However, time-of-flight technology suffers from multipath interference and low resolution. Structured light achieves depth perception by calculating distortion information in two-dimensional images or by marking three-dimensional space, but structured light is fragile to lighting.

[0004] Active stereo technology offers a low-cost depth perception solution to these problems. An active stereo system utilizes two cameras and a projection module. The projection module generates a pre-designed structured light pattern, which artificially adds a layer of texture to the measurement scene. The camera then calculates parallax to obtain scene depth information. This process leverages the spatial properties of structured light, enhancing the uniqueness of matching points and enabling reliable estimation of stereo correspondences. Due to its low-cost infrared projection module and CMOS sensor, active stereo technology has already been deployed in commercial products, such as the Intel RealSense D435.

[0005] Active stereo technology still faces some challenges. The structured light patterns projected by active stereo technology are all pre-designed, and environmental information cannot be effectively integrated into the structured light design process. This disconnects the structured light generation and depth estimation algorithms, affecting depth perception accuracy. Furthermore, the active stereo depth camera projection modules currently deployed on the market generally use diffractive optical elements. However, manufacturing customized diffractive optical elements requires advanced equipment and high costs, which is generally difficult to achieve. Summary of the Invention

[0006] The purpose of this invention is to propose a depth perception method based on the joint optimization of speckle structured light and a stereo matching network. By jointly optimizing digital speckle structured light and a stereo matching network, the depth estimation task can be completed in actual scenes using only a projector and a binocular camera, without the need for customized diffractive optical elements.

[0007] The technical solution of the present invention is as follows: a depth perception method based on the joint optimization of speckle structured light and a stereo matching network, comprising two stages: a training stage and an actual testing stage. During the training stage, a Fourier transform-based speckle structured light generation scheme is constructed, the size, grayscale, and density of the speckle structured light are parameterized, and the speckle structured light generation process is subjected to differentiable processing; a differentiable projection and acquisition model is constructed to synthesize active stereo images; a two-branch stereo matching network is constructed, and the synthesized active stereo image pair and RGB image pair are used as inputs of the two-branch stereo matching network for depth acquisition; the joint optimization of the speckle structured light and the two-branch stereo matching network is completed throughout the training stage; and during the actual testing stage, the optimized speckle structured light is actually projected, and the depth of the actual scene is acquired via the two-branch stereo matching network.

[0008] The Fourier transform-based speckle structured light generation scheme is specifically as follows: according to Fourier optical theory, the speckle structured light pattern is generated by superimposing multiple components with independent phases; discrete Fourier transform is used to generate the speckle structured light pattern, and an N×M phase matrix φ(x, y) is randomly generated. The elements in φ(x, y) are defined as:

[0009]

[0010] Where U(0,1) represents a uniform distribution between 0 and 1, K represents the speckle size, and φ(x,y) is used to control the density of the speckle structured light pattern;

[0011] Design a complex matrix E to represent the light field:

[0012] E(x,y)=e 2πiφ(x,y)

[0013] A differentiable mask based on the Sigmoid function is added before the complex matrix E to control the speckle size. The final light field is expressed as:

[0014] E′(x,y)=Q(x,y)×E(x,y)

[0015] The final light field E′(x,y) is subjected to Fourier transform and square operation in sequence to obtain the speckle structured light pattern:

[0016]

[0017] Pixel-by-pixel brightness control of speckle structured light patterns:

[0018] P′=I(x,y)×P

[0019] Where I(x,y) represents the grayscale value of the pixel;

[0020] Among them, φ(x,y) and I(x,y) are differentiable using the automatic differentiation mechanism, and Q(x,y) itself is differentiable, completing the differentiable processing of the speckle structured light generation process; the speckle structured light pattern P′ generated by φ(x,y), Q(x,y) and I(x,y) serves as the input of the differentiable acquisition model.

[0021] The differentiable projection model includes an optical projection process and a camera imaging process;

[0022] In the simulated optical projection process, the sampling factor of the speckle structured light pattern in the optical projection process is determined by calculating the ratio of the camera pixel size to the projector pixel size:

[0023]

[0024] where p cam and p proj represent the pixel pitch of the camera and the pixel pitch of the projector, respectively, and f cam and f proj Represent the focal length of the camera and the focal length of the projector respectively; the bicubic sampling factor completes the sampling of the speckle structured light pattern P final ←bicubic(P′,s);

[0025] After simulating the optical projection process, the Lambertian model is used to simulate the camera imaging process; according to the depth map D of the camera perspective L / R and the occlusion map O L / R After completing the perspective conversion operation, the sampled speckle structured light pattern P final Distorted to the binocular perspective, the speckle structured light pattern P under the binocular perspective is obtained L / R :

[0026] P L / R =O L / R ⊙warp(P final ,D L / R )

[0027] Where ⊙ represents element-by-element multiplication, and warp represents the warping operator;

[0028] After the perspective conversion, active stereo images are obtained according to the Lambertian model;

[0029] J L / R =γ(α+βP L / R )R L / R +η

[0030] Where γ is a scalar value describing the exposure and the sensor's spectral quantum efficiency, α represents the ambient light, β represents the projector's power, η is the noise, and R L / R Represents reflectivity.

[0031] The dual-branch stereo matching network includes a convolution-based backbone network, multiple attention layers, attention masks, optimal transfer, coarse disparity and occlusion regression, and context adaptation layers;

[0032] The processing of the dual-branch stereo matching network is as follows:

[0033] Taking the active stereo image pair and the RGB stereo image pair as input, the convolution-based backbone network is used to obtain the initial feature descriptor of the active stereo image pair and the feature descriptor of the RGB image pair respectively; the initial feature descriptors of the active stereo image pair are superimposed to obtain the local features of the active stereo image pair;

[0034] The local features of the active stereo image pair and the feature descriptors of the RGB image pair are input into multiple attention layers for processing; each attention layer includes: a main matching path, an active stereo path, and a fusion path; the local features of the active stereo image pair are input into the fusion path via the active stereo path; the feature descriptors of the RGB image pair are input into the fusion path via the main matching path;

[0035] The main matching path and the active stereo path each contain a self-attention layer and a cross-attention layer; the main matching path is used to obtain contextual information based on the RGB stereo image pair and the corresponding pixel similarity; the active stereo path is used to obtain contextual information based on the active stereo image pair and the corresponding pixel similarity; the fusion path is used to integrate the contextual information of the active stereo image pair into the main matching path; in the last attention layer, the features output by the main matching path are used as the input of the attention mask; then the coarse disparity map and occlusion map are obtained through optimal transmission; finally, after processing through the context adaptation layer, the final disparity map and occlusion map are obtained.

[0036] The convolution-based backbone network adopts an hourglass-shaped feature extraction architecture, and the generated feature descriptor is consistent with the size of the original image.

[0037] The initial feature descriptor superposition processing of the active stereo image pair is specifically as follows: after obtaining the initial feature descriptor of the active stereo image pair, for pixel point p, with point p as the center, the feature channels of the surrounding w×w pixels are spliced to form a w×w×c vector, where c represents the feature channel dimension; the above operation is performed on all pixels in the initial feature descriptor, and then a layer of 1×1 convolution and a layer of 3×3 convolution are used to reduce the dimension of the feature channel to the same dimension as the initial feature descriptor, and use it as the input of the attention layer.

[0038] The self-attention layer is used to aggregate information in features in the main matching path and the active stereo path, and the cross-attention layer is used to calculate the pixel similarity on each epipolar line in the RGB stereo image features and the active stereo image features; the fusion path splices the features of the main matching path and the active stereo path, and then uses two convolutional layers to aggregate the active stereo image features containing speckle structured light information into the RGB stereo image features; the fused features serve as the input of the main matching path of the next attention layer, and the output features of the active stereo path continue to serve as the input of the active stereo path of the next attention layer.

[0039] The attention mask not only ensures the uniqueness of the matching points between the two feature images, but also considers the position information of the matching points; let p and q be the matching points on an epipolar line, and the position coordinates increase from left to right, then the position coordinates of the matching points satisfy x p -x q ≥0, the range of the final matching points is a lower triangular matrix;

[0040] The optimal transmission takes the features output by the attention mask as input. In order to ensure the uniqueness of the matching between pixels, the optimal transmission based on entropy regularization is introduced. According to soft distribution and differentiability, the normal flow of gradients is achieved.

[0041] The coarse disparity and occlusion regression first adopts the winner-takes-all principle to find the best matching point k, then builds a 3-pixel window N3(k) around the matching point, normalizes the pixels in the pixel window, and finally calculates the weighted sum of the candidate disparities in the pixel window to obtain the coarse disparity

[0042]

[0043] where t i represents the probability of the candidate disparity, represents the normalized candidate disparity probability, d i represents the candidate disparity; the occlusion probability p of k point occ (k) is calculated as:

[0044]

[0045] The context adaptation layer uses convolution blocks and Sigmoid functions to complete the final estimation, and uses residual blocks and ReLu functions to complete the final disparity estimation, thereby achieving aggregation of context information on different epipolar lines.

[0046] In the joint learning process, the dual-branch stereo matching network parameters are marked as θ n , the parameters of speckle structured light φ(x,y), Q(x,y) and I(x,y) are collectively marked as θs , then the end-to-end joint optimization problem can be summarized as:

[0047]

[0048] in Denotes the predicted left disparity map, D L represents the true value of left disparity, represents the predicted left occlusion map, O L Represents the true value of the left occlusion; for the disparity loss, the relative response loss and L1 loss function are used, and for the occlusion loss, the binary cross entropy loss is used.

[0049] The optimized speckle structured light is actually projected and captured using two monocular cameras and a projector to project the optimized speckle structured light pattern onto the scene to be measured, and the camera is used to obtain active stereo images. The active stereo images captured by the camera are input into a trained dual-branch stereo matching network to complete depth acquisition.

[0050] Beneficial effects of the present invention:

[0051] (1) A differentiable structured light generation scheme is proposed, which realizes the optimizability of structured light parameters and makes structured light optimization possible, which is beneficial to improving depth estimation performance.

[0052] (2) A dual-branch stereo matching network is proposed, which introduces active stereo images containing structured light information into the network to achieve high-precision depth estimation.

[0053] (3) An experimental prototype was built and comprehensive experiments were conducted in datasets and the real world. The results showed that the method of the present invention has robust depth acquisition capabilities and is superior to existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 The overall flow chart of the present invention; (a) is the training phase, (b) is the testing phase;

[0055] Figure 2 This is a diagram of the dual-branch stereo matching network architecture;

[0056] Figure 3 Schematic diagram of the fusion path;

[0057] Figure 4 Schematic diagram of attention mask; (a) is a schematic diagram of the position of image matching points, and (b) is a schematic diagram of attention mask calculation;

[0058] Figure 5 A diagram of an apparatus for use in the method of the present invention;

[0059] Figure 6Qualitative evaluation graphs of accuracy; (a) is a qualitative evaluation graph of the accuracy of the active stereo method ActiveStereoNet, (b) is a qualitative evaluation graph of the accuracy of the structured light optimization method Polka Lines, (c) is a qualitative evaluation graph of the accuracy of the stereo matching method STTR, (d) is a qualitative evaluation graph of the accuracy of the stereo matching method CSTR, (e) is a qualitative evaluation graph of the accuracy of the stereo matching method ELFNet, and (f) is a qualitative evaluation graph of the accuracy of the present invention;

[0060] Figure 7 It is a quantitative evaluation chart of accuracy;

[0061] Figure 8 are qualitative estimation maps; (a) is the RGB image of the scene; (b) is the qualitative estimation map of the active stereo method ActiveStereoNet in different scenes, (c) is the qualitative estimation map of the structured light optimization method Polka Lines in different scenes, (d) is the qualitative estimation map of the stereo matching method STTR in different scenes, (e) is the qualitative estimation map of the stereo matching method CSTR in different scenes, (f) is the qualitative estimation map of the stereo matching method ELFNet in different scenes, (g) is the qualitative estimation map of the present invention in different scenes, and (g) is the true value image. DETAILED DESCRIPTION

[0062] Figure 1 The flowchart of the technical solution of the present invention is as follows. The present invention proposes a depth perception method based on the joint optimization of speckle structured light and stereo matching network, which includes two stages: a training stage and an actual testing stage. In the training stage, a speckle structured light generation scheme based on Fourier transform is constructed, and the size, grayscale and density of the speckle structured light are parameterized and characterized, and the speckle structured light generation process is subjected to differentiable processing. A differentiable acquisition model is constructed to synthesize active stereo images. A dual-branch stereo matching network is constructed, and the synthesized active stereo image pair and RGB image pair are used as inputs of the dual-branch stereo matching network for depth acquisition. During the entire training stage, the speckle structured light and the dual-branch stereo matching network are jointly optimized. In the actual testing stage, the optimized speckle structured light is actually acquired, and the depth of the actual scene is acquired through the dual-branch stereo matching network.

[0063] The specific implementation process includes the following steps:

[0064] (1) Structured light generation. According to Fourier optics theory, the generation of speckle patterns is formed by the superposition of multiple components with independent phases. Discrete Fourier transform can be used to simulate this process. First, assume that an N×M phase matrix φ(x,y) is randomly generated, and the elements in φ(x,y) are defined as:

[0065]

[0066] Where U(0,1) represents a uniform distribution between 0 and 1, K represents the speckle size, and φ(x,y) is used to control the density of the speckle structured light pattern. Then a complex matrix E is obtained to represent the light field:

[0067] E(x,y)=e 2πiφ(x,y)

[0068] Note that the range of φ(x,y) approximates a rectangular function, which results in non-differentiable speckle size. To ensure that the entire process of structured light and deep network is differentiable, we add a differentiable mask Q(x,y) that approximates a matrix function before the matrix E to control the speckle size. The light field is finally expressed as:

[0069] E′(x,y)=Q(x,y)×E(x,y)

[0070] The speckle structured light pattern can be obtained by performing Fourier transform and square operation on the matrix E′(x,y):

[0071]

[0072] In addition to controlling the phase and speckle size, we also control the brightness of the speckle structured light pattern pixel by pixel:

[0073] P′=I(x,y)×P

[0074] Where I(x,y) represents the grayscale value of the pixel.

[0075] So far, we have completed the encoding of the key parameters of structured light, namely density (achieved by controlling the phase), size, and brightness, and ensured the differentiability of the parameters.

[0076] (2) Differentiable projection and sampling model. The sampling factor of the structured light pattern during projection is determined by calculating the ratio of the camera pixel size to the projector pixel size:

[0077]

[0078] where p cam and p proj represent the pixel pitch of the camera and projector respectively, f cam and f proj Represent the focal lengths of the camera and projector respectively. Apply the sampling factor to the structured light pattern and use the bicubic sampling factor to complete the sampling of the structured light P final ←bicubic(P′,s). Since this process has nothing to do with depth information, P final Can be applied in any scenario.

[0079] After simulating the projection process, it is necessary to simulate the light transmission process from the scene to be measured to the stereo camera. This process is implemented using geometric optics.

[0080] Using the depth map D of the camera view L / R and the occlusion map O L / R Complete the perspective conversion operation and transform the structured light P final Distorted to binocular perspective:

[0081] P L / R =O L / R ⊙warp(P final ,D L / R )

[0082] Where ⊙ represents element-wise multiplication and warp represents the warping operator.

[0083] After the viewpoint conversion, the active stereo image is obtained using the Lambertian model.

[0084] J L / R =γ(α+βP L / R )R L / R +η

[0085] Where γ is a scalar value describing the exposure and the sensor's spectral quantum efficiency, α represents the ambient light, β represents the projector's power, η is the noise, and R L / R Represents reflectivity.

[0086] (3) Dual-branch stereo matching network. The dual-branch stereo matching network architecture is as follows: Figure 2 As shown in the figure, the entire process introduces active stereo images to enhance regional recognition. The active stereo image pair and the RGB stereo image pair are taken as input, and the convolution-based backbone network is used to obtain the initial feature descriptor of the active stereo image pair and the feature descriptor of the RGB image pair respectively. The initial feature descriptors of the active stereo image pair are superimposed to obtain the local features of the active stereo image pair.

[0087] The local features of the active stereo image pair and the feature descriptors of the RGB stereo image pair are input into multiple attention layers for processing; each attention layer includes: a main matching path, an active stereo path, and a fusion path; the local features of the active stereo image pair are input into the fusion path via the active stereo path; the feature descriptors of the RGB image pair are input into the fusion path via the main matching path;

[0088] The main matching path and the active stereo path each contain a self-attention layer and a cross-attention layer; the main matching path is used to obtain contextual information based on the RGB stereo image pair and the corresponding pixel similarity; the active stereo path is used to obtain contextual information based on the active stereo image pair and the corresponding pixel similarity; the fusion path is used to integrate the contextual information of the active stereo image pair into the main matching path; in the last attention layer, the features output by the main matching path are used as the input of the attention mask; then the coarse disparity map and occlusion map are obtained through optimal transmission; finally, after processing through the context adaptation layer, the final disparity map and occlusion map are obtained.

[0089] Regarding feature extraction, we use an hourglass-shaped feature extraction architecture, and the generated feature descriptor is the same size as the original image. In particular, we further process the feature descriptors of the active stereo image pair. Taking into account the local uniqueness brought by the structured light information, after obtaining the initial feature descriptor, we superimpose the feature descriptors. Specifically, for pixel p, we take point p as the center and splice the feature channels of the surrounding w×w pixels to form a w×w×c vector (c represents the feature channel dimension), where w is set to 5. The above operations are performed on all pixels in the feature descriptor. Then, after a layer of 1×1 convolution and a layer of 3×3 convolution, the dimension of the feature channel is reduced to the same dimension as the original feature descriptor and input into the attention layer. The purpose of using 1×1 convolution here is to integrate the local information of structured light into a single pixel, and the purpose of using 3×3 convolution is to increase the capture of local features of the active stereo image.

[0090] Regarding the attention layer, self-attention is used to aggregate information in the image in both the main matching path and the active stereo path, and cross-attention is used to calculate the pixel similarity along an epipolar line. Note that self-attention is calculated in both the horizontal and vertical directions in both the main matching path and the active stereo path to improve generalization performance. In order to fully utilize the features of the active stereo image pair, a fusion path is added in each layer, such as Figure 3 Specifically, the features of the main matching path and the active stereo path are concatenated, and then two convolutional layers are used to aggregate the features containing structured light information into the main features. The fused features serve as the input to the next layer of the main matching path.

[0091] Other modules mainly include attention mask, optimal transfer, original disparity and occlusion regression, and context adaptation layer. Regarding optimal transfer and attention mask, in order to ensure the uniqueness of matching, optimal transfer based on entropy regularization is introduced, and its soft distribution and differentiability are used to achieve normal gradient flow. In addition to ensuring the uniqueness of matching points, the location information of matching points also needs to be considered. Figure 4As shown, assuming that p and q are matching points on an epipolar line, and the position coordinates increase gradually from left to right, then the position coordinates of the matching points satisfy x p -x q ≥0, the possible range of the final matching point is a lower triangular matrix. Regarding the original disparity and occlusion regression, the winner-takes-all principle is first used to find the best matching point k. Then a 3-pixel window N3(k) is constructed around the matching point, the pixels in the window are normalized, and finally the weighted sum of the candidate disparities in the window is calculated to obtain the original disparity.

[0092]

[0093] where t i represents the probability of the candidate disparity, represents the normalized candidate disparity probability, d i Represents the candidate disparity. The above method improves the robustness of the model in multimodal distribution. The above is the calculation of non-occlusion disparity probability, then the occlusion probability p of point k is occ (k) can be calculated as:

[0094]

[0095] Regarding the context adaptation layer, we use convolution blocks and Sigmoid functions to complete the final estimation, and use residual blocks and ReLu functions to complete the estimation of the final disparity. This process achieves the aggregation of contextual information on different epipolar lines.

[0096] (4) Joint learning. The parameters of the dual-branch stereo matching network are marked as θ n , the parameters of speckle structured light φ(x,y), Q(x,y) and I(x,y) are collectively marked as θ s , then the end-to-end joint optimization problem can be summarized as:

[0097]

[0098] in Denotes the predicted left disparity map, D L represents the true value of the left disparity. Similarly, represents the predicted left occlusion map, O L Represents the true value of the left occlusion. For the disparity loss, we use the relative response loss and L1 loss function, and for the occlusion loss, we use the binary cross entropy loss.

[0099] The device for realizing the present invention is as follows Figure 5 As shown in the figure, the camera resolution is 1920*1200. The projector resolution is 1920*1080.

[0100] Compared with the existing technology, the present invention improves the depth perception accuracy. A standard plane with weak texture is measured multiple times in the range of 40-180 cm. The qualitative evaluation results are as follows Figure 6 As shown in Figure 2, compared with other methods, the planes reconstructed by the present invention at different depths have smooth surfaces and continuous depths. Figure 7 As shown in the figure, the present invention uses mean absolute error to measure reconstruction performance. The present invention maintains a low error overall, and even though the results at 60cm and 80cm are not optimal, they are suboptimal. Other methods have larger error fluctuations, and their maximum error is much greater than that of the present invention.

[0101] In a qualitative comparison with current mainstream depth perception methods, such as Figure 8 As shown, the present invention obtains a more complete depth map than other methods and has better matching effect in detail areas.

Claims

1. A depth perception method based on joint optimization of speckle structured light and stereo matching network, characterized in that: It includes two phases: training phase and actual testing phase; During the training phase, a Fourier transform-based speckle structured light generation scheme is constructed to parameterize the size, grayscale, and density of the speckle structured light, while performing differentiable processing on the speckle structured light generation process. Build a differentiable projection model to synthesize active stereo images; build a dual-branch stereo matching network, and use the synthesized active stereo image pairs and RGB image pairs as inputs of the dual-branch stereo matching network for depth acquisition; During the entire training phase, the joint optimization of speckle structured light and dual-branch stereo matching network is completed; During the actual testing phase, the optimized speckle structured light is actually projected and sampled, and the depth of the actual scene is obtained through a dual-branch stereo matching network.

2. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 1, characterized in that: The Fourier transform-based speckle structured light generation scheme is specifically as follows: according to Fourier optical theory, the speckle structured light pattern is generated by superimposing multiple components with independent phases; discrete Fourier transform is used to generate the speckle structured light pattern, and an N×M phase matrix φ(x, y) is randomly generated. The elements in φ(x, y) are defined as: Where U(0,1) represents a uniform distribution between 0 and 1, K represents the speckle size, and φ(x,y) is used to control the density of the speckle structured light pattern; Design a complex matrix E to represent the light field: E(x,y)=e 2πiφ(x,y) A differentiable mask based on the Sigmoid function is added before the complex matrix E to control the speckle size. The final light field is expressed as: E′(x,y)=Q(x,y)×E(x,y) The final light field E′(x,y) is subjected to Fourier transform and square operation in sequence to obtain the speckle structured light pattern: Pixel-by-pixel brightness control of speckle structured light patterns: P′=I(x,y)×P Where I(x,y) represents the grayscale value of the pixel; Among them, φ(x,y) and I(x,y) are differentiable using the automatic differentiation mechanism, and Q(x,y) itself is differentiable, completing the differentiable processing of the speckle structured light generation process; the speckle structured light pattern P′ generated by φ(x,y), Q(x,y) and I(x,y) serves as the input of the differentiable acquisition model.

3. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 1, characterized in that: The differentiable projection model includes an optical projection process and a camera imaging process; In the simulated optical projection process, the sampling factor of the speckle structured light pattern in the optical projection process is determined by calculating the ratio of the camera pixel size to the projector pixel size: where p cam and p proj represent the pixel pitch of the camera and the pixel pitch of the projector, respectively, and f cam and f proj Represent the focal length of the camera and the focal length of the projector respectively; the bicubic sampling factor completes the sampling of the speckle structured light pattern P final ←bicubic(P′,s); After simulating the optical projection process, the Lambertian model is used to simulate the camera imaging process; according to the depth map D of the camera perspective L / R and the occlusion map O L / R After completing the perspective conversion operation, the sampled speckle structured light pattern P final Distorted to the binocular perspective, the speckle structured light pattern P under the binocular perspective is obtained L / R : P L / R =O L / R⊙warp(P final ,D L / R ) Where ⊙ represents element-by-element multiplication, and warp represents the warping operator; After the perspective conversion, active stereo images are obtained according to the Lambertian model; J L / R =γ(α+βP L / R )R L / R +n Where γ is a scalar value describing the exposure and the sensor's spectral quantum efficiency, α represents the ambient light, β represents the projector's power, η is the noise, and R L / R Represents reflectivity.

4. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 1, characterized in that: The dual-branch stereo matching network includes a convolution-based backbone network, multiple attention layers, attention masks, optimal transfer, coarse disparity and occlusion regression, and context adaptation layers; The processing of the dual-branch stereo matching network is as follows: Taking the active stereo image pair and the RGB stereo image pair as input, we use the convolution-based backbone network to obtain the initial feature descriptor of the active stereo image pair and the feature descriptor of the RGB image pair respectively; Performing superposition processing on the initial feature descriptors of the active stereo image pair to obtain the local features of the active stereo image pair; The local features of the active stereo image pair and the feature descriptors of the RGB image pair are input into multiple attention layers for processing; In each attention layer, there are: main matching path, active stereo path and fusion path; the local features of the active stereo image pair are input to the fusion path through the active stereo path; the feature descriptor of the RGB image pair is input to the fusion path through the main matching path; The main matching path and the active stereo path each contain a self-attention layer and a cross-attention layer; the main matching path is used to obtain contextual information based on the RGB stereo image pair and the corresponding pixel similarity; the active stereo path is used to obtain contextual information based on the active stereo image pair and the corresponding pixel similarity; the fusion path is used to integrate the contextual information of the active stereo image pair into the main matching path; in the last attention layer, the features output by the main matching path are used as the input of the attention mask; then the coarse disparity map and occlusion map are obtained through optimal transmission; finally, after processing through the context adaptation layer, the final disparity map and occlusion map are obtained.

5. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 4, characterized in that: The convolution-based backbone network adopts an hourglass-shaped feature extraction architecture, and the generated feature descriptor is consistent with the size of the original image.

6. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 4, characterized in that: The initial feature descriptor superposition process of the active stereo image pair is specifically as follows: after obtaining the initial feature descriptor of the active stereo image pair, for a pixel point p, with point p as the center, the feature channels of the surrounding w×w pixels are spliced to form a w×w×c vector, where c represents the feature channel dimension; The above operation is performed on all pixels in the initial feature descriptor. Then, a layer of 1×1 convolution and a layer of 3×3 convolution are performed to reduce the dimension of the feature channel to the same dimension as the initial feature descriptor, and serve as the input of the attention layer.

7. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 6, characterized in that: The self-attention layer is used to aggregate information in features in the main matching path and the active stereo path, and the cross-attention layer is used to calculate the pixel similarity on each epipolar line in the RGB stereo image features and the active stereo image features; The fusion path concatenates the features of the main matching path and the active stereo path, and then uses two convolutional layers to aggregate the active stereo image features containing speckle structured light information into the RGB stereo image features. The fused features serve as the input of the main matching path of the next attention layer, and the output features of the active stereo path continue to serve as the input of the active stereo path of the next attention layer.

8. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 6, characterized in that: The attention mask not only ensures the uniqueness of the matching points between the two feature images, but also considers the position information of the matching points; let p and q be the matching points on an epipolar line, and the position coordinates increase from left to right, then the position coordinates of the matching points satisfy x p -x q ≥0, the range of the final matching points is a lower triangular matrix; The optimal transmission takes the features output by the attention mask as input. In order to ensure the uniqueness of the matching between pixels, the optimal transmission based on entropy regularization is introduced. According to soft distribution and differentiability, the normal flow of gradients is achieved. The coarse disparity and occlusion regression first adopts the winner-takes-all principle to find the best matching point k, then builds a 3-pixel window N3(k) around the matching point, normalizes the pixels in the pixel window, and finally calculates the weighted sum of the candidate disparities in the pixel window to obtain the coarse disparity where t i represents the probability of the candidate disparity, represents the normalized candidate disparity probability, d i represents the candidate disparity; the occlusion probability p of k point occ (k) is calculated as: The context adaptation layer uses convolution blocks and Sigmoid functions to complete the final estimation, and uses residual blocks and ReLu functions to complete the final disparity estimation, thereby achieving aggregation of context information on different epipolar lines.

9. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 6, characterized in that: In the joint learning process, the dual-branch stereo matching network parameters are marked as θ n , the parameters of speckle structured light φ(x,y), Q(x,y) and I(x,y) are collectively marked as θ s , then the end-to-end joint optimization problem can be summarized as: in Denotes the predicted left disparity map, D L represents the true value of left disparity, represents the predicted left occlusion map, O L Represents the true value of the left occlusion; for the disparity loss, the relative response loss and L1 loss function are used, and for the occlusion loss, the binary cross entropy loss is used.

10. The depth perception method based on joint optimization of speckle structured light and stereo matching network according to claim 1, characterized in that: The optimized speckle structured light is actually projected and captured using two monocular cameras and a projector to project the optimized speckle structured light pattern onto the scene to be measured, and the camera is used to obtain an active stereo image pair; the active stereo image pair captured by the camera is input into a trained dual-branch stereo matching network to complete depth acquisition.

Citation Information

Patent Citations

  • Material identification method and device based on laser speckle and modal fusion

    CN110942060A

  • Depth image acquisition method and device, terminal, imaging system and medium

    CN114511608A

  • Parallax prediction method and device, equipment and storage medium

    CN119762565A

  • Polka lines: learning structured illumination and reconstruction for active stereo

    US20220414913A1