Monocular depth recovery method and device based on structured light and electronic equipment

By using the structured light-based depth recovery method, the target acquisition light is constructed using the monocular acquisition device and the system calibration parameters, the rendering grayscale value is calculated and the voxel grid is optimized, and the problems of low depth recovery accuracy and slow calculation speed in the prior art are solved, thereby achieving high-precision and efficient depth recovery.

CN120014009AActive Publication Date: 2025-05-16UNIV OF SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510041213.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing depth recovery methods based on structured light rely on matching algorithms between projection patterns and captured images, resulting in low accuracy and slower calculation speed when processing edges and occlusion areas.

Method used

By using a monocular acquisition device to collect multiple images of the target scene, each image includes a preset projection image, and the target acquisition light is constructed based on the system calibration parameters, obtain the reference grayscale value and volume density, calculate the rendered grayscale value, and iteratively optimize the preset voxel grid, and finally perform deep recovery.

Benefits of technology

Improve the accuracy and calculation speed of depth estimation, overcome the defects of traditional methods in edges and occlusion areas, and achieve more accurate depth recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014009A_ABST
    Figure CN120014009A_ABST
Patent Text Reader

Abstract

The invention provides a monocular depth recovery method and device based on structured light and electronic equipment. The method comprises the following steps: acquiring a target scene by using a monocular acquisition device to obtain a plurality of target scene images; based on the system calibration parameters, constructing acquisition light rays corresponding to M target pixel positions in the target scene image to obtain M target acquisition light rays; obtaining K reference gray values according to the system calibration parameters, a preset projection image corresponding to the target scene image and the K three-dimensional points on each target acquisition light; obtaining K individual densities according to a preset voxel grid and the K three-dimensional points; obtaining a rendering gray value according to the K reference gray values and the K individual densities corresponding to each target pixel position; and performing iterative optimization on the preset voxel grid according to the M rendering gray values and the M scene image gray values corresponding to the plurality of target scene images, and performing depth recovery on the target scene according to the optimized preset voxel grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and more specifically, to a monocular depth recovery method, device and electronic device based on structured light. Background Art

[0002] Depth information plays an important role in understanding the positional relationship between objects in a 3D scene. For example, a robot or self-driving car in a 3D scene can know the distance of surrounding objects from itself based on depth information, which helps them avoid obstacles and adjust their next behavior in time.

[0003] With the development of various types of structured light in the field of depth recovery, depth recovery systems based on structured light have become a powerful solution for measuring depth information in outdoor scenes. However, traditional structured light depth recovery methods rely on the accuracy of the matching algorithm between the projected pattern and the captured image. Any errors in the matching process, such as blur or occlusion, will cause large errors in the final depth map. Therefore, how to significantly improve the accuracy of existing structured light-based depth recovery algorithms while maintaining a high computing speed is a problem that needs to be solved urgently. Summary of the invention

[0004] In view of this, the present disclosure provides a monocular depth recovery method, device and electronic device based on structured light.

[0005] According to one aspect of the present disclosure, a monocular depth recovery method based on structured light is provided, comprising: using a monocular acquisition device to acquire a target scene to obtain a plurality of target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other; based on the plurality of target scene images, iteratively performing a preset number of the following operations, wherein in the i-th round, for each target scene image, based on system calibration parameters, acquisition light rays corresponding to M target pixel positions in the target scene image are constructed to obtain M target acquisition light rays, wherein M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular acquisition device, and i is a positive integer; according to the system calibration parameters, the preset projection image corresponding to the target scene image and the K three-dimensional points on each target acquisition light ray, K reference grayscale values ​​are obtained, wherein, K is a positive integer; according to the preset voxel grid corresponding to the target scene and the K three-dimensional points, K individual densities are obtained; according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities, the rendering grayscale value corresponding to the target pixel position is obtained; according to the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to each of the multiple target scene images, the preset voxel grid is iteratively optimized, and when i is less than the preset number of rounds, i is incremented, and the operation of constructing the target acquisition light based on the system calibration parameters is returned, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image; when i is greater than or equal to the preset number of rounds, the target scene is deeply restored according to the optimized preset voxel grid.

[0006] For example, the above-mentioned obtaining the rendering grayscale value corresponding to the above-mentioned target pixel position according to the above-mentioned K reference grayscale values ​​corresponding to each target pixel position and the above-mentioned K body densities includes: determining k-1 target three-dimensional points arranged before the k-th three-dimensional point according to the arrangement order of the above-mentioned K three-dimensional points along the light projection direction of the target collection light, wherein k is an integer greater than 1 and less than or equal to K; determining the transmittance corresponding to the k-th three-dimensional point according to the body densities respectively corresponding to the above-mentioned k-1 target three-dimensional points; and obtaining the above-mentioned rendering grayscale value by weighted summation of the transmittances respectively corresponding to the K three-dimensional points and the above-mentioned K reference grayscale values.

[0007] For example, the K reference grayscale values ​​obtained according to the system calibration parameters, the preset projection image corresponding to the target scene image and the K three-dimensional points on each target collection light include: for each three-dimensional point, based on the system calibration parameters, projecting the three-dimensional point to the first projection plane of the projector to obtain the projection pixel position corresponding to the three-dimensional point, wherein the first projection plane is the image plane of the projector where the preset projection image is located; determining a plurality of neighboring pixel positions adjacent to the projection pixel position in the preset projection image; interpolating the initial grayscale value at the projection pixel position according to the grayscale values ​​at the plurality of neighboring pixel positions; and determining the reference grayscale value corresponding to the three-dimensional point according to the plurality of target scene images and the initial grayscale value.

[0008] For example, the above-mentioned determination of the reference grayscale value corresponding to the above-mentioned three-dimensional point based on the above-mentioned multiple target scene images and the above-mentioned initial grayscale value includes: determining the minimum grayscale value and edge contrast at the target pixel position corresponding to the above-mentioned three-dimensional point based on the above-mentioned multiple target scene images; determining the reference grayscale value corresponding to the above-mentioned three-dimensional point based on the above-mentioned initial grayscale value, the above-mentioned minimum grayscale value and the above-mentioned edge contrast.

[0009] For example, the above-mentioned K volume densities are obtained based on the preset voxel grid corresponding to the above-mentioned target scene and the above-mentioned K three-dimensional points, including: for each three-dimensional point, determining multiple grid vertices adjacent to the above-mentioned three-dimensional point in the above-mentioned preset voxel grid; performing trilinear interpolation on the volume density at the above-mentioned multiple grid vertices to obtain the volume density corresponding to the above-mentioned three-dimensional point.

[0010] For example, the above-mentioned depth restoration of the target scene according to the optimized preset voxel grid includes: based on the above-mentioned system calibration parameters, constructing the collection lights corresponding to each pixel position on the second projection plane of the above-mentioned monocular acquisition device, and obtaining a plurality of target reconstruction lights, wherein the second projection collection plane represents the image plane of the monocular acquisition device where the target scene image is located; based on the above-mentioned plurality of target reconstruction lights and the above-mentioned optimized preset voxel grid, obtaining a plurality of reconstruction volume densities; according to the above-mentioned plurality of reconstruction volume densities and the preset volume rendering formula, performing depth restoration of the above-mentioned target scene.

[0011] For example, the iterative optimization of the preset voxel grid according to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to each of the multiple target scene images includes: for each target scene image, performing a difference calculation between the rendering grayscale value and the scene image grayscale value corresponding to each target pixel in the target scene image to obtain a color deviation corresponding to each target pixel; constructing an image loss function according to the M color deviations corresponding one-to-one to the M target pixels in the target scene image; and optimizing the preset voxel grid according to the image loss function.

[0012] For example, each target scene image includes multiple first patterns and multiple second patterns, wherein the multiple first patterns and the multiple second patterns have the same size and are all square, the pixel value at each pixel position of the multiple first patterns is a first grayscale value, and the pixel value at each pixel position of the multiple second patterns is a second grayscale value.

[0013] According to another aspect of the present disclosure, a monocular depth recovery device based on structured light is provided, comprising:

[0014] A first obtaining module is used to collect the target scene using a monocular collection device to obtain a plurality of target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other;

[0015] The second obtaining module is used to iteratively perform the following operations for a preset number of rounds based on the above-mentioned multiple target scene images. In the i-th round, for each target scene image, based on the system calibration parameters, the collection light rays corresponding to the M target pixel positions in the above-mentioned target scene image are constructed to obtain M target collection light rays, wherein M is greater than or equal to 1 and less than or equal to the maximum resolution of the above-mentioned monocular collection device, and i is a positive integer;

[0016] A third obtaining module is used to obtain K reference grayscale values ​​according to the system calibration parameters, the preset projection image corresponding to the target scene image, and K three-dimensional points on each target acquisition light, where K is a positive integer;

[0017] A fourth obtaining module is used to obtain K individual densities according to a preset voxel grid corresponding to the target scene and the K three-dimensional points;

[0018] A fifth obtaining module, configured to obtain a rendering grayscale value corresponding to each target pixel position according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities;

[0019] A sixth obtaining module is used to iteratively optimize the preset voxel grid according to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to each of the multiple target scene images, and when i is less than the preset number of rounds, increment i and return to the operation of constructing the target acquisition light based on the system calibration parameters, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image;

[0020] The reconstruction module is used to perform deep restoration of the target scene according to the optimized preset voxel grid when i is greater than or equal to the preset number of rounds.

[0021] According to another aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0022] According to an embodiment of the present disclosure, after a target scene is captured by a monocular capture device to obtain a plurality of target scene images, each target scene image includes a preset projection image projected by a projector onto the target scene, the following operations are iteratively performed for a preset number of rounds based on the plurality of target scene images. In the i-th round, for each target scene image, based on a system calibration parameter, capture rays corresponding to M target pixel positions in the target scene image are constructed to obtain M target capture rays. K reference grayscale values ​​are obtained according to the system calibration parameter, the preset projection image corresponding to the target scene image, and K three-dimensional points on each target capture ray. K individual densities are obtained according to the preset voxel grid corresponding to the target scene and the K three-dimensional points. The K reference grayscale values ​​corresponding to each target pixel position and the K individual densities are obtained. Rendering grayscale value, according to M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to multiple target scene images, iteratively optimize the preset voxel grid, and when i is less than the preset number of rounds, increase i, and return to the operation of constructing the target acquisition light based on the system calibration parameters. When i is greater than or equal to the preset number of rounds, the technical means of depth recovery of the target scene is performed according to the optimized preset voxel grid, so as to realize the grayscale information in the preset projection image as a known grayscale field to optimize the scene geometry separately, which can accelerate the convergence speed and improve the depth estimation accuracy. Moreover, since the depth recovery process of the volume rendering process does not rely on any image matching algorithm, it fundamentally overcomes the defects of the traditional structured light depth recovery method in processing edges and occluded areas, so it can obtain more accurate depth recovery accuracy while maintaining a high calculation speed, and has extremely high practical potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0024] Figure 1 The flowchart of the monocular depth recovery method based on structured light according to an embodiment of the present disclosure is schematically shown;

[0025] Figure 2 An exemplary system architecture to which a structured light-based monocular depth recovery method can be applied according to an embodiment of the present disclosure is schematically shown;

[0026] Figure 3 The structure block diagram of a monocular depth recovery device based on structured light according to an embodiment of the present disclosure is schematically shown; and

[0027] Figure 4 A block diagram of an electronic device suitable for implementing the above-described structured light-based monocular depth recovery method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0029] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "include", "comprising", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0030] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0032] In the embodiments of the present disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information) are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.

[0033] Typically, a monocular depth recovery system based on structured light consists of a camera and a projector, and the intrinsic and extrinsic parameters of the two devices have been systematically calibrated in advance. The system projects random or artificially designed patterns into three-dimensional space and extracts depth information by analyzing the deformation of these patterns in the captured image. The classic algorithm of structured light aims to establish a robust correspondence between multiple projected patterns.

[0034] The depth recovery process in structured light systems usually requires a trade-off between scanning accuracy and the number of projected patterns or captured images. Increasing the number of captured images can improve the accuracy of correspondence matching. However, more projected patterns also lead to longer acquisition times, which limits the applicability of structured light systems in general scenes, especially those with moving objects or short exposure times. To address this issue, techniques that can embed richer information in a limited set of patterns are used. These methods use complex patterns to encode temporal or spatial features to reduce the uncertainty of matching. By decoding these features in the captured images, the structured light system can determine the correspondences required for depth estimation. Despite these advances, designing features that can generate accurate high-density depth maps while being resilient to environmental influences remains a major challenge.

[0035] Since traditional structured light depth recovery methods rely on the accuracy of the matching algorithm between the projected pattern and the captured image, any errors in the matching process, such as blur or occlusion, will cause large errors in the final depth map. Related technologies provide solutions by using neural networks used in deep learning technology to solve the uncertainty of matching. However, most of these methods directly obtain the predicted depth by training a general model. Although these models can generate dense depth maps, they require a large training data set, and the quality of the data in the training data set will seriously affect the network performance. In structured light-based depth recovery systems, it is challenging to build such a training data set due to the diversity of equipment and pattern configurations. In addition, it takes a lot of time to train a general model based on the collected training data set. As a result, both traditional structured light depth recovery algorithms and deep learning-based depth recovery algorithms face problems such as low computational speed and low depth recovery accuracy due to inaccurate image matching.

[0036] Therefore, how to significantly improve the accuracy of existing structured light-based depth recovery algorithms while maintaining a high computing speed is a problem that needs to be solved urgently.

[0037] In view of this, the embodiments of the present disclosure provide a monocular depth recovery method, device and electronic device based on structured light, which can be applied to the field of computer vision technology.

[0038] The embodiment of the present disclosure provides a monocular depth recovery method based on structured light, comprising: using a monocular acquisition device to acquire a target scene to obtain a plurality of target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other; based on the plurality of target scene images, iteratively performing a preset number of the following operations, wherein in the i-th round, for each target scene image, based on a system calibration parameter, acquisition light rays corresponding to M target pixel positions in the target scene image are constructed respectively to obtain M target acquisition light rays, wherein M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular acquisition device, and i is a positive integer; according to the system calibration parameter, the preset projection image corresponding to the target scene image and the K three-dimensional points on each target acquisition light ray, K reference grayscale values ​​are obtained, wherein, K is a positive integer; according to the preset voxel grid and K three-dimensional points corresponding to the target scene, K individual densities are obtained; according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities, the rendering grayscale value corresponding to the target pixel position is obtained; according to the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to multiple target scene images, the preset voxel grid is iteratively optimized, and when i is less than the preset number of rounds, i is incremented, and the operation of constructing the target acquisition light based on the system calibration parameters is returned, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image; when i is greater than or equal to the preset number of rounds, the target scene is deeply restored according to the optimized preset voxel grid.

[0039] Figure 1 The flowchart of the monocular depth recovery method based on structured light according to an embodiment of the present disclosure is schematically shown.

[0040] like Figure 1 As shown, the monocular depth recovery method based on structured light includes operations S101 to S107.

[0041] In operation S101, a target scene is captured using a monocular capture device to obtain a plurality of target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other.

[0042] For example, before performing operation S101, a monocular depth recovery system based on structured light is first constructed. The monocular depth recovery system based on structured light may include a monocular acquisition device and a projector. The monocular acquisition device may be a monocular structured light camera. The monocular structured light camera and the projector have been system-calibrated and calibrated in advance.

[0043] For example, in the process of capturing an image of a target scene to be measured using a monocular structured light camera, the pattern set of a preset projection image projected by the projector to the target scene may be a black and white coded pattern composed of unit squares of a fixed scale, and the color of each square is randomly set to black or white.

[0044] For example, the sizes of the multiple preset projection images corresponding to the multiple target scene images are the same. Each preset projection image may include a black and white coding pattern composed of unit squares of a fixed scale.

[0045] For example, the number of the plurality of target scene images may be 6. The number of the plurality of preset projection images may be 6. Two groups of patterns may be prepared using squares with unit lengths of 20 pixels, 10 pixels, and 5 pixels, respectively, to obtain six preset projection images.

[0046] According to the embodiments of the present disclosure, the specific design of the preset projection image does not significantly affect the final performance of the monocular depth recovery method based on structured light provided by the embodiments of the present disclosure. Therefore, as long as the projection pattern meets the above paradigm, it can be used in the image capture of the target scene to be measured in the embodiments of the present disclosure.

[0047] In operation S102, based on the plurality of target scene images, the following operations are iterated for a preset number of rounds. In the i-th round, for each target scene image, based on the system calibration parameters, the collection light rays corresponding to the M target pixel positions in the target scene image are constructed to obtain M target collection light rays. Wherein, M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular collection device, and i is a positive integer.

[0048] According to the embodiments of the present disclosure, M can be selected according to actual conditions and is not limited here. For example, M can be 3000, 4096, 4100 or 4500.

[0049] According to the embodiments of the present disclosure, the preset number of rounds can be selected according to actual conditions and is not limited here. For example, the preset number of rounds can be 1000, 2000, 3500 or 4000.

[0050] According to an embodiment of the present disclosure, in the same round, the M target pixel positions corresponding to each of the multiple target scene images may be completely identical, partially identical, or completely different. In different rounds, the M target pixel positions corresponding to the same target scene image may be completely identical, partially identical, or completely different.

[0051] For example, M pixel positions on the second projection plane of the monocular acquisition device can be selected first, and the M pixel positions can be determined as the M target pixel positions in the target scene image. Then, according to the known system calibration parameters of the monocular depth recovery system based on structured light, the light positions and directions in the three-dimensional space corresponding to these target pixel positions are calculated. The light positions and directions in the three-dimensional space corresponding to the target pixel positions are used to represent the target acquisition light. Among them, the second projection acquisition plane represents the image plane of the monocular acquisition device where the target scene image is located.

[0052] In operation S103, K reference grayscale values ​​are obtained according to the system calibration parameters, the preset projection image corresponding to the target scene image, and K three-dimensional points on each target collection light, where K is a positive integer.

[0053] According to the embodiments of the present disclosure, K can be selected according to actual conditions and is not limited here. For example, K can be 100, 128, 150 or 300.

[0054] For example, a series of three-dimensional points in a three-dimensional space may be sampled according to different distances on each target collection ray to participate in a subsequent volume rendering process.

[0055] For example, based on the external parameters characterizing the monocular acquisition device and the projector in the system calibration parameters, K three-dimensional points can be projected onto the image plane of the projector. The preset projection image is queried according to the projection pixel position of the projection, and the grayscale values ​​corresponding to the K three-dimensional points are obtained, and K reference grayscale values ​​are obtained for the subsequent volume rendering process.

[0056] According to an embodiment of the present disclosure, by obtaining K reference grayscale values ​​based on system calibration parameters, a preset projection image corresponding to the target scene image and K three-dimensional points on each target acquisition light, it is possible to interpolate and calculate the grayscale values ​​of these sampled three-dimensional points using the pixel grayscale values ​​of the known preset projection image for use in the subsequent volume rendering process.

[0057] In operation S104, K volume densities are obtained according to a preset voxel grid corresponding to the target scene and K three-dimensional points.

[0058] According to an embodiment of the present disclosure, before performing operation S104, a preset voxel grid for storing volume density may be constructed, wherein the preset voxel grid is a dense three-dimensional grid, and the resolution of the preset voxel grid may be 256×256×256.

[0059] For example, the positions of K three-dimensional points may be input into a preset voxel grid to be optimized, and the corresponding volume density may be interpolated to obtain K individual densities, which may be used in a subsequent volume rendering process.

[0060] In operation S105 , a rendering grayscale value corresponding to the target pixel position is obtained according to K reference grayscale values ​​and K individual densities corresponding to each target pixel position.

[0061] For example, volume rendering is performed on the corresponding target collection light rays according to K reference grayscale values ​​and K volume densities corresponding to each target pixel position, and a rendering grayscale value corresponding to the target pixel position is obtained according to the volume-rendered target collection light rays.

[0062] For example, in the process of obtaining the rendering grayscale value corresponding to the target pixel position, the integral in the rendering equation can be discretized and converted into the sum of the transmittance corresponding to the sampled 3D point and the reference grayscale. The transmittance is calculated by the volume density corresponding to the sampled 3D point and the volume density of all sampled 3D points before the 3D point in the order of the target acquisition light direction.

[0063] In operation S106, the preset voxel grid is iteratively optimized according to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to each of the multiple target scene images, and when i is less than the preset number of rounds, i is incremented, and the operation of constructing the target acquisition light based on the system calibration parameters is returned. The scene image grayscale value represents the grayscale value at the target pixel position in the target scene image.

[0064] For example, in the first round, the preset voxel grid can be optimized using the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to the first target scene image among the multiple target scene images, and then the preset voxel grid optimized based on the first target scene image can be iteratively optimized using the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to the second target scene image among the multiple target scene images, ..., and so on, the preset voxel grid optimized based on the multiple target scene images before the last one can be iteratively optimized using the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to the last target scene image among the multiple target scene images to obtain the preset voxel grid optimized in the first round. In the second round, the preset voxel grid optimized in the first round is iteratively optimized, ..., and so on, the preset voxel grid optimized in the preset number of rounds minus 1 rounds is iteratively optimized in the preset number of rounds to obtain the final optimized preset voxel grid.

[0065] According to an embodiment of the present disclosure, by using an explicit preset voxel grid to represent the target scene geometry, the speed of each iteration in the optimization process can be accelerated.

[0066] According to the embodiments of the present disclosure, through projection transformation, the grayscale information of the preset projection image is used as a known grayscale field to optimize the preset voxel grid, which helps to accelerate the convergence speed and improve the performance of the final target scene geometry.

[0067] According to the embodiments of the present disclosure, once the image matching of the traditional structured light depth recovery method is calculated, the influence of the wrong matching on the depth recovery result cannot be eliminated. The essential difference is that in the monocular depth recovery method based on structured light provided by the embodiments of the present disclosure, the error of the scene geometry can be gradually corrected through the optimization process of the voxel grid. Therefore, the monocular depth recovery method based on structured light provided by the embodiments of the present disclosure overcomes the defects of the matching-based method, and can produce better results at the edge or occluded area.

[0068] In operation S107 , when i is greater than or equal to a preset number of rounds, depth restoration is performed on the target scene according to the optimized preset voxel grid.

[0069] According to an embodiment of the present disclosure, after a target scene is captured by a monocular capture device to obtain a plurality of target scene images, each target scene image includes a preset projection image projected by a projector onto the target scene, the following operations are iteratively performed for a preset number of rounds based on the plurality of target scene images. In the i-th round, for each target scene image, based on a system calibration parameter, capture rays corresponding to M target pixel positions in the target scene image are constructed to obtain M target capture rays. K reference grayscale values ​​are obtained according to the system calibration parameter, the preset projection image corresponding to the target scene image, and K three-dimensional points on each target capture ray. K individual densities are obtained according to the preset voxel grid corresponding to the target scene and the K three-dimensional points. The K reference grayscale values ​​corresponding to each target pixel position and the K individual densities are obtained. Rendering grayscale value, according to M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to multiple target scene images, iteratively optimize the preset voxel grid, and when i is less than the preset number of rounds, increase i, and return to the operation of constructing the target acquisition light based on the system calibration parameters. When i is greater than or equal to the preset number of rounds, the technical means of depth recovery of the target scene is performed according to the optimized preset voxel grid, so as to realize the grayscale information in the preset projection image as a known grayscale field to optimize the scene geometry separately, which can accelerate the convergence speed and improve the depth estimation accuracy. Moreover, since the depth recovery process of the volume rendering process does not rely on any image matching algorithm, it fundamentally overcomes the defects of the traditional structured light depth recovery method in processing edges and occluded areas, so it can obtain more accurate depth recovery accuracy while maintaining a high calculation speed, and has extremely high practical potential.

[0070] According to an embodiment of the present disclosure, each target scene image includes multiple first patterns and multiple second patterns, wherein the multiple first patterns and the multiple second patterns are of the same size and are all square, the pixel value at each pixel position of the multiple first patterns is a first grayscale value, and the pixel value at each pixel position of the multiple second patterns is a second grayscale value.

[0071] For example, the sizes of the plurality of first patterns and the plurality of second patterns may all be 20 pixels, 10 pixels, or 5 pixels. The first grayscale value may be 0, and the second grayscale value may be 255. When the first grayscale value is 0, the first pattern is black. When the second grayscale value is 255, the second pattern is white.

[0072] According to the embodiments of the present disclosure, Figure 1 Operation S103 shown, obtaining K reference grayscale values ​​according to the system calibration parameters, the preset projection image corresponding to the target scene image and the K three-dimensional points on each target collection light, may include the following operations:

[0073] For each 3D point, based on the system calibration parameters, the 3D point is projected onto a first projection plane of the projector to obtain a projection pixel position corresponding to the 3D point, wherein the first projection plane is an image plane of the projector where a preset projection image is located;

[0074] Determine a plurality of adjacent pixel positions adjacent to a projection pixel position in a predetermined projection image;

[0075] According to the grayscale values ​​at the positions of a plurality of adjacent pixels, an initial grayscale value at the position of the projection pixel is obtained by interpolation;

[0076] According to a plurality of target scene images and an initial grayscale value, a reference grayscale value corresponding to a three-dimensional point is determined.

[0077] According to an embodiment of the present disclosure, for each three-dimensional point, based on the system calibration parameters, the three-dimensional point is projected onto the first projection plane of the projector to obtain a projection pixel position corresponding to the three-dimensional point, the first projection plane is an image plane of a preset projection image on the projector, a plurality of neighboring pixel positions of the neighboring projection pixel position in the predetermined image are determined, and an initial grayscale value at the projection pixel position is interpolated according to the grayscale values ​​at the plurality of neighboring pixel positions, and a reference grayscale value corresponding to the three-dimensional point is determined according to a plurality of target scene images and the initial grayscale value, so as to directly calculate the grayscale value of the sampled three-dimensional point through a known projected preset projection image, thereby avoiding the image feature matching process that must be adopted in the traditional structured light depth recovery method, and also avoiding the optimization of the color field in the traditional depth recovery task based on volume rendering.

[0078] According to an embodiment of the present disclosure, determining a reference grayscale value corresponding to a three-dimensional point according to a plurality of target scene images and an initial grayscale value includes:

[0079] Determine the minimum grayscale value and edge contrast at the target pixel position corresponding to the three-dimensional point according to the multiple target scene images;

[0080] The reference grayscale value corresponding to the three-dimensional point is determined according to the initial grayscale value, the minimum grayscale value and the edge contrast.

[0081] According to an embodiment of the present disclosure, the minimum grayscale value represents the background light level of the target collection light.

[0082] For example, the minimum value of the pixel values ​​at the target pixel position r corresponding to the three-dimensional point of the multiple target scene images can be calculated to obtain the minimum grayscale value at the target pixel position corresponding to the three-dimensional point. The maximum value of the pixel values ​​at the target pixel position r corresponding to the three-dimensional point of the multiple target scene images can be calculated to obtain the maximum grayscale value at the target pixel position corresponding to the three-dimensional point. The maximum grayscale value is subtracted from the minimum grayscale value to obtain the edge contrast at the target pixel position r corresponding to the three-dimensional point.

[0083] For example, the operation of determining the reference grayscale value corresponding to the three-dimensional point according to the initial grayscale value, the minimum grayscale value and the edge contrast can be implemented according to formula (1).

[0084] (1)

[0085] Among them, c k is the reference grayscale value of the k-th 3D point, B(r) is the minimum grayscale value at the target pixel position r corresponding to the k-th 3D point determined based on multiple target scene images, F(r) is the edge contrast at the target pixel position r corresponding to the k-th 3D point determined based on multiple target scene images, P j (Π(x k )) is the jth preset projection image P at the kth three-dimensional point x k The corresponding projected pixel position The initial grayscale value at , j is a positive integer.

[0086] According to the embodiments of the present disclosure, the two parameters of minimum grayscale value and edge contrast can process the occluded area. Because in the occluded area, the edge contrast and the minimum grayscale value are close to zero, so the contribution of these areas to the volume rendering process is almost zero.

[0087] According to an embodiment of the present disclosure, the minimum grayscale value and edge contrast at the target pixel position corresponding to the three-dimensional point are determined based on multiple target scene images, and the reference grayscale value corresponding to the three-dimensional point is determined based on the initial grayscale value, the minimum grayscale value and the edge contrast. The initial grayscale value is corrected according to the minimum grayscale value and the edge contrast, and the reference grayscale value of the sampled three-dimensional point participating in the subsequent light volume rendering process is obtained.

[0088] According to the embodiments of the present disclosure, Figure 1 Operation S104 shown, obtaining K individual densities according to a preset voxel grid and K three-dimensional points corresponding to the target scene, may include the following operations:

[0089] For each three-dimensional point, determining a plurality of grid vertices adjacent to the three-dimensional point in a preset voxel grid;

[0090] The volume densities at multiple mesh vertices are trilinearly interpolated to obtain the volume densities corresponding to the 3D points.

[0091] For example, for each 3D point, the position of the 3D point may be input into a preset voxel grid to be optimized, and a plurality of grid vertices adjacent to the 3D point may be obtained by querying the preset voxel grid.

[0092] According to the embodiments of the present disclosure, Figure 1Operation S105 shown, obtaining a rendering gray value corresponding to the target pixel position according to K reference gray values ​​and K individual densities corresponding to each target pixel position, may include the following operations:

[0093] According to the arrangement order of the K three-dimensional points along the light projection direction of the target collection light, k-1 target three-dimensional points arranged before the kth three-dimensional point are determined, where k is an integer greater than 1 and less than or equal to K;

[0094] Determine the transmittance corresponding to the kth three-dimensional point according to the volume densities corresponding to the k-1 target three-dimensional points respectively;

[0095] The rendering gray value is obtained by performing a weighted summation based on the transmittances corresponding to the K three-dimensional points and the K reference gray values.

[0096] For example, the operation of obtaining the rendering grayscale value corresponding to the target pixel position according to the K reference grayscale values ​​and K individual densities corresponding to each target pixel position can be implemented according to formulas (2) and (3).

[0097] (2)

[0098] (3)

[0099] in, is the volume density corresponding to the kth three-dimensional point, is the volume density corresponding to the s-th 3D point, is the transmittance corresponding to the kth three-dimensional point, and is the distance between adjacent 3D points, is the three-dimensional coordinate value of the s+1th three-dimensional point, is the three-dimensional coordinate value of the sth three-dimensional point.

[0100] According to the embodiments of the present disclosure, Figure 1 The operation S106 shown, iteratively optimizing the preset voxel grid according to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to each of the multiple target scene images, may include the following operations:

[0101] For each target scene image, a difference calculation is performed between the rendering grayscale value corresponding to each target pixel in the target scene image and the scene image grayscale value to obtain a color deviation corresponding to each target pixel;

[0102] Constructing an image loss function according to M color deviations corresponding one-to-one to M target pixels in the target scene image;

[0103] According to the image loss function, a preset voxel grid is optimized.

[0104] For example, a volume rendering process similar to the above process can be performed on the volume density corresponding to the sampled three-dimensional point to obtain the intersection points of the light rays under the current optimization and the target scene surface. By projecting these intersection points onto the image plane of the projector to query the grayscale value, a rendered image based on the target scene surface is obtained, which is compared with the corresponding captured target scene image to construct a surface loss function.

[0105] Furthermore, based on the fact that the volume density along each collection ray should be unimodal, the distortion loss can be constructed according to the interval size and volume density of the adjacent sampled 3D points along the collection ray to minimize the weighted distance between point pairs in all intervals and the weighted size of each interval.

[0106] The optimization of the voxel grid to be optimized can be implemented based on an open source deep learning framework for machine learning and deep learning. The optimizer can be Adam (Adaptive Moment Estimation). The resolution of the constructed voxel grid can be set to 256×256×256. During the optimization process, 4096 pixels can be randomly sampled at each iteration to form 4096 collection rays, and each collection ray samples 128 three-dimensional space points. The initial loss function can only use image loss and distortion loss, where the weight of image loss is 1 and the weight of distortion loss is 0.001. After 1000 iterations, surface loss can be added with a weight of 1.

[0107] It should be noted that the above-mentioned parameter values ​​are for illustration only and are not intended to be limiting. In practical applications, the content of the training data set and the parameter values ​​can be adjusted based on existing technologies.

[0108] According to an embodiment of the present disclosure, after the optimization of the voxel grid to be optimized is completed, for the target scene to be measured, the corresponding collection light of each pixel is constructed from the perspective of the structured light camera, the optimized preset voxel grid is queried to obtain the volume density, and a dense depth map is calculated through the volume rendering formula, thereby completing depth recovery.

[0109] According to the embodiments of the present disclosure, Figure 1 Operation S107 shown, performing depth restoration on the target scene according to the optimized preset voxel grid, may include the following operations:

[0110] Based on the system calibration parameters, the collection light rays corresponding to each pixel position on the second projection plane of the monocular collection device are constructed to obtain a plurality of target reconstruction light rays, wherein the second projection collection plane represents the image plane of the monocular collection device where the target scene image is located;

[0111] Based on multiple target reconstruction rays and the optimized preset voxel grid, multiple reconstruction volume densities are obtained;

[0112] According to multiple reconstruction volume densities and preset volume rendering formulas, the target scene is deeply restored.

[0113] Figure 2 An exemplary system architecture to which a structured light-based monocular depth recovery method according to an embodiment of the present disclosure can be applied is schematically shown.

[0114] like Figure 2 As shown, the system architecture 200 includes a monocular structured light camera 210, a projector 220, and a processor.

[0115] In the process of capturing an image of a target scene to be measured using the monocular structured light camera 210, the projector 220 projects a preset projection image 201 onto the target scene. The pattern set of the preset projection image 201 of the preset projection image may be a black and white coded pattern composed of unit squares of a fixed scale, and the color of each square is randomly set to black or white. The target scene surface of the target scene is 230.

[0116] After the target scene is captured by the monocular structured light camera 210 to obtain multiple target scene images 202, the processor can construct the collection light corresponding to the M target pixel positions in the target scene image 202 for each target scene image 202 based on the system calibration parameters to obtain M target collection light rays Lt. According to the system calibration parameters, the preset projection image 201 corresponding to the target scene image 202 and the K three-dimensional points x on each target collection light ray Lt, the processor can construct the collection light corresponding to the M target pixel positions in the target scene image 202 based on the system calibration parameters to obtain M target collection light rays Lt. k , and obtain K reference grayscale values. Among them, I (r j ) is the jth target scene image I in relation to the 3D point x k The grayscale value at the corresponding target pixel position r. k , and obtain K individual densities. According to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities, the rendering grayscale value corresponding to the target pixel position is obtained, and the rendered image 204 is obtained. According to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to each of the multiple target scene images 202, the preset voxel grid 203 is iteratively optimized. The depth of the target scene is restored according to the optimized preset voxel grid.

[0117] according to Figure 1 The monocular depth recovery method based on structured light shown in Figure 2It can be seen from the exemplary system architecture of the monocular depth recovery method based on structured light that the monocular depth recovery method based on structured light provided by the embodiment of the present disclosure has the following advantages over the traditional structured light depth recovery method and the existing learning method:

[0118] 1) By constructing a loss function through the volume rendering process to optimize the scene geometry, sharper object edges can be produced and the overall structure of the object can be better maintained; 2) Using the color information in the preset projection image as a known color field to separately optimize the scene geometry can speed up the convergence speed and improve the depth estimation accuracy; 3) Using an explicit voxel grid to represent the scene geometry can speed up the iteration of each step in the optimization process, such as 3 times faster than the SDF-based structured light depth recovery algorithm; 4) Since the proposed depth recovery process based on volume rendering does not rely on any image matching algorithm, it fundamentally overcomes the defects of traditional methods in dealing with edges and occluded areas, and obtains more accurate depth recovery accuracy, such as 30%-60% less average error in predicted depth on test data than traditional methods such as GC; 5) Since there is no need to pre-train with a data set and there are no high requirements for the design of preset projection images, it has good application capabilities in real scenes.

[0119] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order of multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0120] Based on the above-mentioned monocular depth restoration method based on structured light, an embodiment of the present disclosure provides a monocular depth restoration device based on structured light.

[0121] Figure 3 The structural block diagram of a monocular depth recovery device based on structured light according to an embodiment of the present disclosure is schematically shown.

[0122] like Figure 3 As shown, the monocular depth recovery device 300 based on structured light includes a first obtaining module 310 , a second obtaining module 320 , a third obtaining module 330 , a fourth obtaining module 340 , a fifth obtaining module 350 , a sixth obtaining module 360 ​​and a reconstruction module 370 .

[0123] The first obtaining module 310 is used to collect the target scene using a monocular collection device to obtain multiple target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the multiple target scene images are different from each other. In one embodiment, the first obtaining module 310 can be used to perform the operation S101 described above, which will not be repeated here.

[0124] The second obtaining module 320 is used to iterate the following operations for a preset number of rounds based on multiple target scene images. In the i-th round, for each target scene image, based on the system calibration parameters, the collection light corresponding to the M target pixel positions in the target scene image is constructed to obtain M target collection light, where M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular collection device, and i is a positive integer. In one embodiment, the second obtaining module 320 can be used to perform the operation S102 described above, which will not be repeated here.

[0125] The third obtaining module 330 is used to obtain K reference grayscale values ​​according to the system calibration parameters, the preset projection image corresponding to the target scene image, and the K three-dimensional points on each target acquisition light, where K is a positive integer. In one embodiment, the third obtaining module 330 can be used to perform the operation S103 described above, which will not be repeated here.

[0126] The fourth obtaining module 340 is used to obtain K individual densities according to the preset voxel grid and K three-dimensional points corresponding to the target scene. In one embodiment, the fourth obtaining module 340 can be used to perform the operation S104 described above, which will not be described in detail here.

[0127] The fifth obtaining module 350 is used to obtain the rendering grayscale value corresponding to the target pixel position according to the K reference grayscale values ​​and K individual densities corresponding to each target pixel position. In one embodiment, the fifth obtaining module 350 can be used to perform the operation S105 described above, which will not be repeated here.

[0128] The sixth obtaining module 360 ​​is used to iteratively optimize the preset voxel grid according to the M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to each of the multiple target scene images, and when i is less than the preset number of rounds, increment i and return to the operation of constructing the target acquisition light based on the system calibration parameters, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image. In one embodiment, the fifth obtaining module 360 ​​can be used to perform the operation S106 described above, which will not be repeated here.

[0129] The reconstruction module 370 is used to perform depth restoration of the target scene according to the optimized preset voxel grid when i is greater than or equal to the preset round number. In one embodiment, the reconstruction module 370 can be used to perform the operation S107 described above, which will not be repeated here.

[0130] According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits, or at least part of the functions of any one of them can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as hardware circuits, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems on chips, systems on substrates, systems on packages, application specific integrated circuits (ASICs), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, submodules, units, and subunits can be at least partially implemented as computer program modules, and when the computer program modules are run, the corresponding functions can be performed.

[0131] For example, any multiple of the first obtaining module 310, the second obtaining module 320, the third obtaining module 330, the fourth obtaining module 340, the fifth obtaining module 350, the sixth obtaining module 360 ​​and the reconstruction module 370 can be combined in one module / unit / subunit for implementation, or any one of the modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functions of one or more of the modules / units / subunits can be combined with at least part of the functions of other modules / units / subunits and implemented in one module / unit / subunit. According to an embodiment of the present disclosure, at least one of the first obtaining module 310, the second obtaining module 320, the third obtaining module 330, the fourth obtaining module 340, the fifth obtaining module 350, the sixth obtaining module 360 ​​and the reconstruction module 370 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the first obtaining module 310, the second obtaining module 320, the third obtaining module 330, the fourth obtaining module 340, the fifth obtaining module 350, the sixth obtaining module 360 ​​and the reconstruction module 370 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function can be executed.

[0132] It should be noted that the part of the monocular depth recovery device based on structured light in the embodiments of the present disclosure corresponds to the part of the monocular depth recovery method based on structured light in the embodiments of the present disclosure. For the specific description corresponding to the part of the monocular depth recovery device based on structured light, please refer to the part of the monocular depth recovery method based on structured light, which will not be repeated here.

[0133] Figure 4 A block diagram of an electronic device suitable for implementing the above-described structured light-based monocular depth recovery method according to an embodiment of the present disclosure is schematically shown. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0134] like Figure 4 As shown, the electronic device 400 according to an embodiment of the present disclosure includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 to a random access memory (RAM) 403. The processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 may also include an onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0135] In RAM 403, various programs and data required for the operation of electronic device 400 are stored. Processor 401, ROM 402 and RAM 403 are connected to each other via bus 404. Processor 401 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 402 and / or RAM 403. It should be noted that the program can also be stored in one or more memories other than ROM 402 and RAM 403. Processor 401 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.

[0136] According to an embodiment of the present disclosure, the electronic device 400 may further include an input / output (I / O) interface 405, which is also connected to the bus 404. The electronic device 400 may further include one or more of the following components connected to the input / output (I / O) interface 405: an input portion 406 including a keyboard, a mouse, etc.; an output portion 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 408 including a hard disk, etc.; and a communication portion 409 including a network interface card such as a LAN card, a modem, etc. The communication portion 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output (I / O) interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage portion 408 as needed.

[0137] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the processor 401, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0138] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0139] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.

[0140] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 402 and / or the RAM 403 described above and / or one or more memories other than the ROM 402 and the RAM 403 .

[0141] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the monocular depth recovery method based on structured light provided by the embodiment of the present disclosure.

[0142] When the computer program is executed by the processor 401, the above functions defined in the system / device of the embodiment of the present disclosure are executed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0143] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 409, and / or installed from the removable medium 411. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0144] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It can be understood by those skilled in the art that the features recorded in the various embodiments and / or claims of the present disclosure can be combined and / or combined in a variety of ways, even if such a combination or combination is not explicitly recorded in the present disclosure. In particular, without departing from the spirit and teaching of the present disclosure, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0146] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A monocular depth recovery method based on structured light, comprising: Using a monocular acquisition device to acquire a target scene, a plurality of target scene images are obtained, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other; Based on the multiple target scene images, iteratively perform the following operations for a preset number of rounds, wherein in the i-th round, for each target scene image, based on the system calibration parameters, the collection light rays corresponding to the M target pixel positions in the target scene image are constructed to obtain M target collection light rays, wherein M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular collection device, and i is a positive integer; Obtain K reference grayscale values ​​according to the system calibration parameters, a preset projection image corresponding to the target scene image, and K three-dimensional points on each target acquisition light, where K is a positive integer; Obtaining K individual densities according to a preset voxel grid corresponding to the target scene and the K three-dimensional points; Obtaining a rendering grayscale value corresponding to each target pixel position according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities; Iteratively optimize the preset voxel grid according to M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to each of the multiple target scene images, and when i is less than a preset number of rounds, increment i and return to the operation of constructing the target acquisition light based on the system calibration parameters, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image; When i is greater than or equal to the preset number of rounds, depth restoration is performed on the target scene according to the optimized preset voxel grid.

2. The monocular depth restoration method according to claim 1, wherein: The obtaining, according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities, a rendering grayscale value corresponding to the target pixel position comprises: Determine k-1 target three-dimensional points arranged before the kth three-dimensional point according to the arrangement order of the K three-dimensional points along the light projection direction of the target collection light, where k is an integer greater than 1 and less than or equal to K; Determining the transmittance corresponding to the kth three-dimensional point according to the volume densities respectively corresponding to the k-1 target three-dimensional points; The rendering grayscale value is obtained by performing weighted summation according to the transmittances respectively corresponding to the K three-dimensional points and the K reference grayscale values.

3. The monocular depth restoration method according to claim 1 or 2, wherein: The obtaining of K reference grayscale values ​​according to the system calibration parameters, the preset projection image corresponding to the target scene image, and the K three-dimensional points on each target acquisition light line comprises: For each three-dimensional point, based on the system calibration parameters, project the three-dimensional point onto a first projection plane of the projector to obtain a projection pixel position corresponding to the three-dimensional point, wherein the first projection plane is an image plane of the projector where the preset projection image is located; Determine a plurality of adjacent pixel positions adjacent to the projection pixel position in the preset projection image; Interpolating an initial grayscale value at the projection pixel position according to the grayscale values ​​at the plurality of adjacent pixel positions; A reference grayscale value corresponding to the three-dimensional point is determined according to the multiple target scene images and the initial grayscale value.

4. The monocular depth restoration method according to claim 3, wherein: The determining, according to the plurality of target scene images and the initial grayscale value, a reference grayscale value corresponding to the three-dimensional point comprises: Determining, based on the plurality of target scene images, a minimum grayscale value and an edge contrast at a target pixel position corresponding to the three-dimensional point; A reference grayscale value corresponding to the three-dimensional point is determined according to the initial grayscale value, the minimum grayscale value and the edge contrast.

5. The monocular depth restoration method according to claim 1 or 2, wherein: The obtaining K individual densities according to the preset voxel grid corresponding to the target scene and the K three-dimensional points includes: For each three-dimensional point, determining a plurality of grid vertices in the preset voxel grid adjacent to the three-dimensional point; Perform trilinear interpolation on the volume densities at the plurality of mesh vertices to obtain the volume density corresponding to the three-dimensional point.

6. The monocular depth restoration method according to claim 1 or 2, wherein: The performing depth restoration of the target scene according to the optimized preset voxel grid comprises: Based on the system calibration parameters, collection rays corresponding to respective pixel positions on a second projection collection plane of the monocular collection device are constructed to obtain a plurality of target reconstruction rays, wherein the second projection collection plane represents an image plane of the monocular collection device where the target scene image is located; Obtaining a plurality of reconstructed volume densities based on the plurality of target reconstruction rays and the optimized preset voxel grid; The target scene is depth-restored according to the plurality of reconstruction volume densities and a preset volume rendering formula.

7. The monocular depth restoration method according to claim 1 or 2, wherein: The iterative optimization of the preset voxel grid according to the M rendering grayscale values ​​and the M scene image grayscale values ​​corresponding to the multiple target scene images respectively includes: For each target scene image, performing a difference calculation between a rendering grayscale value corresponding to each target pixel in the target scene image and a scene image grayscale value to obtain a color deviation corresponding to each target pixel; Constructing an image loss function according to M color deviations corresponding one-to-one to the M target pixels in the target scene image; The preset voxel grid is optimized according to the image loss function.

8. The monocular depth restoration method according to claim 1 or 2, wherein: Each target scene image includes multiple first patterns and multiple second patterns, wherein the multiple first patterns and the multiple second patterns are of the same size and are all square, the pixel value at each pixel position of the multiple first patterns is a first grayscale value, and the pixel value at each pixel position of the multiple second patterns is a second grayscale value.

9. A monocular depth recovery device based on structured light, comprising: A first obtaining module is used to collect the target scene using a monocular collection device to obtain a plurality of target scene images, wherein each target scene image includes a preset projection image projected by a projector onto the target scene, and the preset projection images included in the plurality of target scene images are different from each other; A second obtaining module is used to iteratively perform a preset number of rounds of the following operations based on the multiple target scene images, wherein in the i-th round, for each target scene image, based on the system calibration parameters, the collection light rays corresponding to the M target pixel positions in the target scene image are constructed to obtain M target collection light rays, wherein M is greater than or equal to 1 and less than or equal to the maximum resolution of the monocular collection device, and i is a positive integer; A third obtaining module is used to obtain K reference grayscale values ​​according to the system calibration parameters, a preset projection image corresponding to the target scene image, and K three-dimensional points on each target acquisition light, wherein K is a positive integer; A fourth obtaining module, used for obtaining K individual densities according to a preset voxel grid corresponding to the target scene and the K three-dimensional points; A fifth obtaining module, configured to obtain a rendering grayscale value corresponding to each target pixel position according to the K reference grayscale values ​​corresponding to each target pixel position and the K individual densities; A sixth obtaining module is used to iteratively optimize the preset voxel grid according to M rendering grayscale values ​​and M scene image grayscale values ​​corresponding to each of the multiple target scene images, and when i is less than a preset number of rounds, increment i and return to the operation of constructing the target acquisition light based on the system calibration parameters, wherein the scene image grayscale value represents the grayscale value at the target pixel position in the target scene image; The reconstruction module is used to perform depth restoration of the target scene according to the optimized preset voxel grid when i is greater than or equal to the preset number of rounds.

10. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Depth image processing method based on light field

    CN106228507A

  • Snapshot spectrum depth joint imaging method and system based on deep learning

    CN110880162A

  • Image restoration method and apparatus

    EP3992904A1

  • Dark channel based image defogging method for linear self-adaptive improvement of global atmospheric light

    WO2019205707A1