Depth completion method, system and device based on multi-modal depth camera and storage medium

By using confidence weighting and spatial consistency fusion of multimodal depth cameras, adaptive depth gradient filtering, and deep learning networks, the incompleteness problem in multimodal depth map fusion is solved, generating high-quality full-range depth maps and improving the accuracy and continuity of depth measurement.

CN120852237APending Publication Date: 2025-10-28SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510828514.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing depth completion methods struggle to effectively integrate depth maps acquired by multimodal depth cameras at different distances, resulting in incomplete and inaccurate depth information, particularly in edge and discontinuous regions.

Method used

By acquiring dense, sparse, and sparse depth maps at extremely close range, near-mid range, and far range using a multimodal depth camera, confidence-weighted and spatial consistency fusion is performed. Combined with depth gradient adaptive filtering, viewpoint reprojection, and depth continuity edge fusion, a deep learning network is constructed for depth completion.

Benefits of technology

It generates high-quality depth maps across the entire measurement range, improving the accuracy and reliability of depth measurements, ensuring the integrity and continuity of depth maps, and enhancing the efficiency of depth information utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852237A_ABST
    Figure CN120852237A_ABST
Patent Text Reader

Abstract

A depth completion method, system and device based on a multi-modal depth camera, and a storage medium, the method comprising: S1, obtaining three depth maps through a multi-modal modulation depth camera to perform confidence weighting and airspace consistency fusion, and generating a full-scale high-quality depth map and an auxiliary depth map; s2, performing re-projection on the full-scale high-quality depth map and the auxiliary depth map, and performing pixel-level alignment on the full-scale high-quality depth map and the auxiliary depth map and an RGB image; s3, performing adaptive filtering on the auxiliary depth map, fusing the auxiliary depth map with the full-scale high-quality depth map, and segmenting the fused depth map into a first short-distance sub-map and a first long-distance sub-map according to statistical characteristics of depth distribution; s4, constructing a deep completion network based on deep learning, inputting the second short-distance sub-image and the RGB image into the deep completion network for deep completion, and generating a dense depth image; and S5, edge fusion based on depth continuity is carried out on the dense depth map and the second long-distance sub-map, and a final depth map is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of depth completion technology, and more specifically, to a depth completion method, system, device, and storage medium based on a multimodal depth camera. Background Art

[0002] In fields such as computer vision and 3D reconstruction, the acquisition and processing of depth maps are crucial. Depth maps provide distance information between objects in a scene and the camera, laying the foundation for subsequent tasks such as target recognition, scene understanding, and robot navigation.

[0003] Traditional depth cameras, such as LiDAR and structured light cameras, have certain limitations in acquiring depth information. For example, while LiDAR can provide high-precision depth information, it is expensive and the data volume is relatively sparse, making it difficult to meet the needs of detailed scene description. Structured light cameras can acquire relatively dense depth maps at close range, but their accuracy decreases significantly with increasing distance and are easily affected by ambient light.

[0004] In recent years, multimodal depth cameras have emerged, combining various depth measurement technologies to acquire depth information with varying degrees of accuracy across different distances. For example, by combining time-of-flight (ToF) cameras and stereo vision cameras, dense and sparse depth maps can be acquired at extremely close, near-mid, and long distances, respectively. However, these depth maps at different distances exhibit inconsistencies in accuracy and density, leading to incomplete and inaccurate depth information when used directly.

[0005] Currently, existing depth completion methods mainly process single-type depth maps, making it difficult to fully utilize multi-scale depth information acquired by multimodal depth cameras at different distances. Furthermore, existing methods often fail to guarantee depth continuity and consistency when processing edges and discontinuous regions of the depth map, resulting in noticeable flaws in the completed depth map.

[0006] To achieve high-precision, full-range depth map acquisition and meet the higher requirements for depth information in fields such as computer vision and 3D reconstruction, a method is needed that can effectively fuse depth maps acquired by multimodal depth cameras at different distances and complete and optimize the depth maps.

[0007] The above background information is provided only to aid in understanding the inventive concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0008] To address this, the proposed depth completion method based on a multimodal depth camera generates a high-quality depth map across the entire range by performing confidence-weighted and spatial consistency fusion on dense depth maps at close range, sparse depth maps at near-mid range, and sparse depth maps at far range. Furthermore, by combining techniques such as adaptive filtering based on depth gradients, viewpoint reprojection, depth completion, and edge fusion based on depth continuity, the method effectively solves the problems existing in current methods and improves the quality and completeness of the depth map.

[0009] In a first aspect, the present invention provides a depth completion method based on a multimodal depth camera, characterized in that it includes:

[0010] Step S1: Acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulation depth camera. Perform confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, extract the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map.

[0011] Step S2: Reproject the full-range high-quality depth map and the auxiliary depth map, and align them pixel-level with the RGB image;

[0012] Step S3: Based on RGB information and depth gradient, perform adaptive filtering on the auxiliary depth map and fuse it with the full-range high-quality depth map to obtain a fused depth map. Then, according to the statistical characteristics of the depth distribution, divide the fused depth map into a first near-range sub-map and a first far-range sub-map.

[0013] Step S4: Construct a deep learning-based depth completion network, input the second near sub-image and the RGB image into the depth completion network for depth completion, and generate a dense depth map;

[0014] Step S5: Perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

[0015] Optionally, the depth completion method based on a multimodal depth camera is characterized in that the confidence weighting and spatial consistency fusion in step S1 specifically involves: weighting the depth values ​​according to the confidence of the three depth maps, while considering the neighborhood relationship of pixels in the spatial domain to ensure that the fused depth map has spatial consistency, and filtering and adjusting the depth values ​​by setting a consistency threshold to generate a full-range high-quality depth map.

[0016] Optionally, the depth completion method based on a multimodal depth camera is characterized in that, in step S3, the adaptive filtering of the auxiliary depth map specifically involves: combining the gradient of the dense depth at close range and the gradient change of the RGB image to generate a hybrid gradient weight, adaptively adjusting the filter parameters, performing smoothing filtering on the region with a smaller gradient, and performing detail-preserving filtering on the region with a larger gradient, thereby preprocessing the auxiliary depth map.

[0017] Optionally, the depth completion method based on a multimodal depth camera is characterized in that, in step S3, the fused depth map is divided into a first near-range sub-map and a first far-range sub-map, specifically including:

[0018] Step S21: Calculate the mean and standard deviation of the depth values ​​in the fused depth map;

[0019] Step S22: Based on the preset near and far depth range thresholds, and combined with the mean and the standard deviation, divide the pixels in the fused depth map into near and far regions.

[0020] Step S23: Extract the pixels of the near region and the far region respectively to generate a first near sub-image and a first far sub-image.

[0021] Optionally, the depth completion method based on a multimodal depth camera is characterized in that, in step S3, the joint calibration parameters of the depth camera and the RGB camera include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients. By reprojecting the first near-field sub-image and the first far-field sub-image from the viewpoint, they are aligned with the pixel coordinate system of the RGB image to generate a second near-field sub-image and a second far-field sub-image.

[0022] Optionally, the depth completion method based on a multimodal depth camera is characterized in that, in step S4, when constructing the depth completion network based on deep learning, the network structure adopts an encoder-decoder structure. The encoder part uses a convolutional neural network to extract features from the RGB image and the second nearest sub-image, and the decoder part uses a deconvolutional neural network to map the features back to the depth map space. Furthermore, skip connections are introduced into the network to fuse features from different levels of the encoder with features from the corresponding levels of the decoder.

[0023] Optionally, the depth completion method based on a multimodal depth camera is characterized in that, in step S5, edge fusion based on depth continuity is performed on the dense depth map and the second distant sub-map, specifically including:

[0024] Step S51: Calculate the depth difference between adjacent pixels in the dense depth map and the second distant sub-map;

[0025] Step S52: Determine the edge position of the fusion region based on the magnitude of the depth difference;

[0026] Step S53: At the edge location, the depth values ​​of the dense depth map and the second distant sub-map are weighted and fused according to the principle of depth continuity to eliminate depth discontinuities at the edge.

[0027] Secondly, the present invention provides a depth completion system based on a multimodal depth camera, used to implement the depth completion method based on a multimodal depth camera as described in any of the preceding claims, characterized in that it includes:

[0028] The initial generation module is used to acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulated depth camera. It performs confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, it extracts the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map.

[0029] The reprojection module is used to reproject the full-range high-quality depth map and the auxiliary depth map, and align them with the RGB image at the pixel level.

[0030] The sub-image generation module is used to adaptively filter the auxiliary depth map based on RGB information and depth gradient, and fuse it with the full-range high-quality depth map to obtain a fused depth map. The fused depth map is then divided into a first near sub-image and a first far sub-image according to the statistical characteristics of the depth distribution.

[0031] A dense module is used to construct a deep learning-based depth completion network. The second near sub-image and the RGB image are input into the depth completion network for depth completion to generate a dense depth map.

[0032] The fusion module is used to perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

[0033] Thirdly, the present invention provides a depth completion device based on a multimodal depth camera, characterized in that it includes:

[0034] processor;

[0035] A memory in which executable instructions of the processor are stored;

[0036] The processor is configured to execute the steps of any of the preceding depth completion methods based on a multimodal depth camera by executing the executable instructions.

[0037] Fourthly, the present invention provides a computer-readable storage medium for storing a program, characterized in that, when the program is executed, it implements the steps of the depth completion method based on a multimodal depth camera as described in any of the preceding claims.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] This invention acquires depth maps at extremely close, near-mid, and long distances using a multimodal modulated depth camera, comprehensively covering scenes across different distance ranges and ensuring the integrity of depth information. By performing confidence-weighted and spatially consistent fusion on the depth maps from different distance ranges, the advantages of each modality can be fully utilized to generate high-quality depth maps across the entire measurement range, improving the accuracy and reliability of depth measurements.

[0040] In this invention, during the fusion process, depth point sets with confidence levels below a first threshold but above a second threshold are extracted to generate auxiliary depth maps. This method fully utilizes depth information that, while having slightly lower confidence levels, still possesses some value, avoiding information waste. The auxiliary depth map, used in conjunction with the full-range high-quality depth map, provides more information support for subsequent processing, helping to improve the depth completion effect.

[0041] This invention reprojects the full-range high-quality depth map and auxiliary depth map, and aligns them with the RGB image at the pixel level. This ensures accurate spatial matching between the depth map and the RGB image, providing an accurate foundation for subsequent depth completion and fusion.

[0042] This invention employs adaptive filtering based on RGB information and depth gradients: Based on the texture features and depth gradient variations of the RGB image, an auxiliary depth map is adaptively filtered to retain effective depth information while removing noise and outliers. The fused depth map is segmented into near-range and far-range sub-maps based on the statistical characteristics of the depth distribution, facilitating more refined processing of depth information at different distance ranges.

[0043] This invention constructs a deep learning-based deep completion network that uses RGB images and nearest-range sub-images for depth completion, generating high-quality dense depth maps. The deep learning model can learn complex patterns and relationships in depth information, improving the accuracy and robustness of depth completion.

[0044] This invention employs depth continuity-based edge fusion of dense depth maps and distant sub-maps, ensuring the continuity and smoothness of the final depth map at edges and generating a final depth map containing near-field high-density and far-field low-density depth information. This method fully leverages the advantages of dense depth maps and distant sub-maps to generate high-quality depth maps. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0046] Figure 1 This is a flowchart illustrating the steps of a depth completion method based on a multimodal depth camera in an embodiment of the present invention.

[0047] Figure 2 This is a flowchart illustrating a step in an embodiment of the present invention to segment the image into a first near-range sub-image and a first far-range sub-image.

[0048] Figure 3 This is a flowchart illustrating the steps of edge blending in an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of a depth completion system based on a multimodal depth camera according to an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of the structure of a depth completion device based on a multimodal depth camera according to an embodiment of the present invention; and

[0051] Figure 6 This is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Detailed Implementation

[0052] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.

[0053] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0054] The present invention provides a depth completion method based on a multimodal depth camera, which aims to solve the problems existing in the prior art.

[0055] The technical solutions of the present invention and how they solve the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0056] The depth completion method based on a multimodal depth camera proposed in this invention generates a high-quality depth map across the entire range by performing confidence weighting and spatial consistency fusion on dense depth maps at close range, sparse depth maps at near-mid range, and sparse depth maps at far range. Combined with techniques such as adaptive filtering based on depth gradient, viewpoint reprojection, depth completion, and edge fusion based on depth continuity, this method can effectively solve the problems existing in the current method and improve the quality and completeness of the depth map.

[0057] Figure 1 This is a flowchart illustrating the steps of a depth completion method based on a multimodal depth camera in an embodiment of the present invention.

[0058] like Figure 1 As shown, the steps of a depth completion method based on a multimodal depth camera in an embodiment of the present invention include:

[0059] Step S1: Acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulation depth camera. Perform confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, extract the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map.

[0060] In this step, the obtained depth values ​​and pixels are segmented according to the range of depth values ​​to obtain three different depth maps.

[0061] Extremely close-range dense depth maps: Multimodal modulation depth cameras acquire dense depth information within an extremely close range using specific modulation modes (such as time-of-flight combined with high-precision modulation signals). Because objects reflect signals more strongly at extremely close range, the camera can capture more details, thus generating dense depth maps and providing rich depth data for close-range scenes. The depth values ​​in extremely close-range dense depth maps are all within the extremely close-range range.

[0062] Near-to-mid-range sparse depth maps: Within the near-to-mid-range, the camera employs a different modulation mode (such as structured light coded modulation). Due to factors such as signal attenuation with increasing distance, the acquired depth points are relatively sparse. However, this distance range is a critical area in many application scenarios (such as indoor navigation and object grasping), and the sparse depth map can still provide important spatial structure information. The depth values ​​in the near-to-mid-range sparse depth map are all within the near-to-mid-range.

[0063] Long-range sparse depth maps: At long distances, cameras employ long-range modulation modes (such as low-frequency time-of-flight), further weakening the signal strength and resulting in sparser depth points. However, long-range depth maps are still crucial for perceiving large scenes and understanding overall spatial layout. The depth values ​​in long-range sparse depth maps are all located at long distances.

[0064] Confidence weighting: A confidence score is calculated for each depth point in each depth map. The confidence score reflects the reliability of the measurement at that depth point. Different weighting coefficients are assigned to the three depth maps based on the measurement accuracy of different modes within their respective distance ranges. For example, a dense depth map with high accuracy at close range has a larger weighting coefficient; a sparse depth map with some reference value at long range has relatively low accuracy, so its weighting coefficient is smaller. The three depth maps are then fused according to the weighting coefficients to obtain a preliminary fused depth map.

[0065] Spatial Consistency Fusion: Considering the spatial continuity of adjacent pixels, a local window-based depth value smoothing algorithm is used to further process the preliminary fused depth map. The window size is adaptively adjusted according to the spatial resolution of the depth map; a smaller window is used in high-resolution areas to retain more details, while a larger window is used in low-resolution areas to improve the smoothing effect. In this way, local noise and discontinuities that may occur during the fusion process are eliminated, generating a high-quality depth map across the entire range.

[0066] Auxiliary depth map generation: During the fusion process, some depth points with confidence levels below the first threshold are filtered out. However, some of these depth points may have confidence levels above the second threshold (first threshold > second threshold), and these points still have some reference value. These depth point sets are extracted to generate an auxiliary depth map, providing additional depth information for subsequent processing.

[0067] Step S2: Reproject the full-range high-quality depth map and the auxiliary depth map, and align them pixel-level with the RGB image.

[0068] In this step, the full-range high-quality depth map and auxiliary depth map are acquired from different viewpoints or coordinate systems. To achieve pixel-level alignment with the RGB image, they need to be reprojected onto the coordinate system of the RGB image. The reprojection process uses a geometric transformation method based on camera intrinsic and extrinsic parameters. The camera intrinsic parameters include focal length, principal point coordinates, etc., while the extrinsic parameters include rotation matrices and translation vectors. These parameters are used to convert the pixel coordinates in the depth map into pixel coordinates in the RGB image.

[0069] After reprojection, pixel-level alignment is performed between the depth map and the RGB image. Due to potential interpolation errors during reprojection, sub-pixel-level interpolation is applied to both the RGB image and the depth map to improve alignment accuracy. A bicubic interpolation algorithm is employed, which can more accurately calculate the pixel value at the sub-pixel position based on the values ​​of surrounding pixels, ensuring precise pixel-level alignment between the depth map and the RGB image. This provides an accurate foundation for subsequent depth processing based on RGB information.

[0070] Step S3: Based on RGB information and depth gradient, perform adaptive filtering on the auxiliary depth map and fuse it with the full-range high-quality depth map to obtain a fused depth map. Then, according to the statistical characteristics of the depth distribution, divide the fused depth map into a first near-range sub-map and a first far-range sub-map.

[0071] In this step, although the auxiliary depth map contains some depth information, it may contain noise and inaccuracies. To optimize the auxiliary depth map, adaptive filtering is applied based on the texture features and depth gradient changes of the RGB image. The texture features of the RGB image reflect the details and structure of the object's surface, while the depth gradient reflects the changes in depth values. When the RGB image has rich texture and a large depth gradient change, it indicates that the area may have a complex object surface; in this case, the filtering intensity is reduced to retain more depth details. Conversely, when the RGB image has simple texture and a small depth gradient change, it indicates that the area may be relatively smooth; in this case, the filtering intensity is increased to remove noise and outliers.

[0072] The adaptively filtered auxiliary depth map is fused with a full-range high-quality depth map using a weighted average fusion method. Weights are determined based on the confidence and reliability of the two depth maps at corresponding locations to generate the fused depth map. The fused depth map combines the advantages of both maps, further improving the quality of the depth information.

[0073] Statistical feature analysis of the depth distribution is performed on the fused depth map, including calculating the mean, variance, and distribution range of depth values. Based on these statistical features, the fused depth map is segmented into a first near-range sub-map and a first far-range sub-map. For example, the mean of the depth distribution plus k times the standard deviation can be used as the segmentation threshold (k ranges from [1,3]). Regions with depth values ​​less than this threshold are classified as the first near-range sub-map, and regions with depth values ​​greater than this threshold are classified as the first far-range sub-map. This segmentation method facilitates more refined processing of depth information across different distance ranges.

[0074] Step S4: Construct a deep learning-based depth completion network, input the second near sub-image and the RGB image into the depth completion network for depth completion, and generate a dense depth map.

[0075] In this step, a deep learning-based deep completion network is constructed, employing an encoder-decoder structure. The encoder uses a convolutional neural network, extracting features from the RGB image and the second nearest sub-image (the first nearest sub-image undergoes further processing or is directly used as the second nearest sub-image, depending on the actual processing flow) through multiple convolutional and pooling operations. The decoder uses a deconvolutional neural network, mapping the features extracted by the encoder back to the depth map space, gradually restoring the resolution of the depth map. Simultaneously, skip connections are introduced into the network to fuse features from different levels of the encoder with corresponding features from the decoder, preserving more detailed information.

[0076] A supervised learning method was used to train the deep incomplete network. The training dataset included a large number of labeled RGB images and their corresponding dense depth maps. The loss function was a weighted combination of mean squared error (MSE) and structural similarity loss functions. The MSE loss function measured the pixel-level difference between the predicted and ground truth depth maps, while the structural similarity loss function measured their structural similarity. The MSE loss function weights were set to [0.7, 0.9], and the structural similarity loss function weights were set to [0.1, 0.3]. By optimizing the loss function, the network was able to learn the mapping relationship from RGB images and near sub-images to the dense depth map.

[0077] The second near-range sub-image and the RGB image are input into the trained depth completion network, which outputs a dense depth map. This dense depth map has high density and accuracy in the near-range range, providing high-quality near-range depth information for subsequent fusion with the far-range sub-image.

[0078] Step S5: Perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

[0079] In this step, edge fusion based on depth continuity is performed on the generated dense depth map and the second distant sub-map. During the fusion process, depth continuity is considered, meaning the depth values ​​of adjacent pixels should change continuously. For the edge regions of the dense depth map and the second distant sub-map, the changes in depth values ​​are analyzed, and a suitable fusion algorithm (such as weighted averaging, depth gradient-based fusion, etc.) is used to fuse them together to generate the final depth map. The final depth map contains near-range high-density depth information (from the dense depth map) and far-range low-density depth information (from the second distant sub-map), which can more accurately reflect the depth distribution of objects in the scene.

[0080] In some embodiments, the confidence-weighted and spatial consistency fusion in step S1 specifically involves: weighting the depth values ​​based on the confidence levels of the three depth maps, while considering the neighborhood relationships of pixels in the spatial domain to ensure spatial consistency of the fused depth map. A consistency threshold is set to filter and adjust depth values ​​to generate a high-quality depth map across the entire depth range. Each corresponding pixel location (pixels in the same spatial location) in the three depth maps—the extremely close-range dense depth map, the near-mid-range sparse depth map, and the far-range sparse depth map—has its own depth value and confidence level. Confidence level is an indicator of the reliability of the depth value, which can be determined by various factors such as the depth camera's measurement mechanism, signal strength, and the number of measurements. For example, the extremely close-range dense depth map, due to its closer distance, is relatively accurate, and its pixel confidence levels may generally be higher; while the far-range sparse depth map, due to its greater distance and higher measurement difficulty, may have lower confidence levels.

[0081] The weighted averaging process involves calculating the depth values ​​at the same pixel location from the three depth maps by weighting them according to their respective confidence levels. Depth values ​​with higher confidence levels have a greater weight in the fusion result, making the fused depth map more likely to use reliable depth information.

[0082] In images, the depth values ​​of adjacent pixels typically exhibit a certain correlation, meaning they should be spatially coherent. For example, on a planar object, the depth values ​​of adjacent pixels should be similar. Therefore, during the fusion process, it is necessary to consider not only the depth value and confidence level of individual pixels but also the situation of their surrounding neighboring pixels.

[0083] Typically, a neighborhood is defined, such as a 3×3 or 5×5 pixel area centered on the current pixel. For all pixels within this area, their depth value distribution is analyzed. If the depth value of a pixel differs significantly from the depth values ​​of other pixels in the neighborhood, then the depth value of this pixel may be inaccurate and needs adjustment.

[0084] To determine whether the difference between a pixel's depth value and its neighboring pixel depth values ​​is too large, a consistency threshold is set. This threshold is determined based on the specific application scenario and the characteristics of the depth map.

[0085] When the difference between the depth value of a pixel and the depth values ​​of other pixels in its neighborhood exceeds a consistency threshold, it indicates that the depth value of that pixel may be problematic. In this case, the depth value of that pixel can be adjusted, for example, by recalculating its weighted average depth value, or by using a statistical measure of the depth values ​​of neighboring pixels (such as the mean or median) instead. This method filters out unreasonable depth values ​​and adjusts them, ensuring that the fused depth map has spatial consistency and avoiding abrupt depth changes.

[0086] Through a series of operations, including confidence-weighted averaging, considering neighborhood relationships, and adjusting depth values ​​based on consistency thresholds, each pixel of the three depth maps is processed to generate a high-quality depth map covering the entire range (from very close to very far distances). This depth map combines the advantages of depth maps at different distance ranges, utilizing dense depth information from very close distances, as well as depth information from near, mid, and far distances, and exhibits good spatial consistency, thus more accurately reflecting the depth distribution of objects in the scene.

[0087] In some embodiments, in step S3, the adaptive filtering of the auxiliary depth map specifically involves: combining the gradient of the dense depth at close range with the gradient changes of the RGB image to generate a hybrid gradient weight, adaptively adjusting the filter parameters, performing smoothing filtering on the regions with smaller gradients, and performing detail-preserving filtering on the regions with larger gradients, thereby preprocessing the auxiliary depth map.

[0088] In this step, the core purpose of adaptive filtering on the auxiliary depth map is to flexibly adjust the filtering strategy based on the depth and gradient features of the RGB image, so as to preserve key details while smoothing noise.

[0089] Depth gradient calculation:

[0090] For each pixel in the auxiliary depth map, its depth gradient in the horizontal and vertical directions is calculated. Classic edge detection operators such as the Sobel operator and the Prewitt operator are typically used for this calculation. Taking the Sobel operator as an example, convolution operations are performed between the corresponding convolution kernels and the depth map in the horizontal (x-direction) and vertical (y-direction) directions, respectively, to obtain the gradient components Gx and Gy in the two directions.

[0091] Then, the depth gradient magnitude Gd for each pixel is calculated based on these two components, using the formula Gd = Gx² + Gy². This magnitude reflects the degree of change in depth value near that pixel; a larger value indicates a more drastic change in depth, which may correspond to the edge or surface details of an object.

[0092] RGB image gradient calculation:

[0093] Similarly, the RGB image is processed. Since the RGB image has three channels (red, green, and blue), the gradient of each channel in the horizontal and vertical directions is calculated to obtain the gradient components Gx_R, Gy_R, Gx_G, Gy_G, Gx_B, and Gy_B for each channel.

[0094] Calculate the gradient magnitudes GR, GG, and GB for each channel, i.e., GR = Gx_R2 + Gy_R2, GG = Gx_G2 + Gy_G2, GB = Gx_B2 + Gy_B2.

[0095] To comprehensively consider the gradient information of RGB images, a weighted average method can be used to calculate the overall gradient magnitude (Grgb) of the RGB images. For example, based on the human eye's sensitivity to different colors, a higher weight (e.g., 0.6) can be given to the green channel, while lower weights (e.g., 0.2) can be given to the red and blue channels. Then, Grgb = 0.2GR + 0.6GG + 0.2GB.

[0096] Depth and RGB gradient blending:

[0097] The calculated depth gradient magnitude Gd and the combined gradient magnitude Grgb from the RGB image are fused to generate a mixed gradient magnitude Gmix. A simple fusion method is to directly add them, i.e., Gmix = Gd + Grgb. This method combines the gradient information from both the depth and RGB images, and can more comprehensively reflect the spatial changes around each pixel.

[0098] Weight calculation:

[0099] The blended gradient weights W are calculated based on the blended gradient magnitude Gmix. To ensure that the weights are smaller in regions with large gradient changes (corresponding to object edges or details) and larger in regions with small gradient changes (corresponding to smooth regions), a combination of normalization and an inverse proportional function can be used. For example, Gmix is ​​first normalized to obtain the normalized gradient magnitude Gnorm, which ranges from [0,1]. Then, the weights W = 1 - Gnormk are calculated.

[0100] Here, k is an adjustable parameter used to control the sensitivity of the weights to changes in the gradient magnitude. When k is large, the weights are more sensitive to changes in the gradient magnitude; when k is small, the weight changes are relatively gradual.

[0101] Filter selection:

[0102] A Gaussian filter is chosen as the base filter because it has good smoothing properties, and its filtering effect can be controlled by adjusting the standard deviation σ. The larger the standard deviation σ, the stronger the smoothing effect of the filter; the smaller the standard deviation σ, the better the filter's ability to preserve details.

[0103] Parameter adjustment:

[0104] The standard deviation σ of the Gaussian filter is adaptively adjusted based on the calculated mixed gradient weights W. For example, a mapping relationship can be established between the weights W and the standard deviation σ, such as σ = σmax - (σmax - σmin)W, where σmax and σmin are the maximum and minimum values ​​of the standard deviation, respectively, and are set according to actual needs. When W is small, it indicates a large gradient in that region, so σ is set to a smaller value, and the filter retains more detail; when W is large, it indicates a small gradient in that region, so σ is set to a larger value, and the filter performs stronger smoothing.

[0105] Filtering operation:

[0106] A Gaussian filter with adjusted parameters is used to filter the auxiliary depth map. For each pixel in the auxiliary depth map, a suitable neighborhood window (such as 3×3, 5×5, etc.) is selected centered on that pixel. The depth values ​​within the neighborhood are then weighted and averaged according to the weight matrix of the Gaussian filter to obtain the filtered depth value. In this way, the auxiliary depth map is preprocessed, smoothing noise while preserving key details, providing better depth information for subsequent fusion with the full-range high-quality depth map.

[0107] Figure 2 This is a flowchart illustrating a step in an embodiment of the present invention to divide the image into a first near-range sub-image and a first far-range sub-image.

[0108] like Figure 2As shown, in an embodiment of the present invention, a step of segmenting into a first near-range subgraph and a first far-range subgraph includes:

[0109] Step S21: Calculate the mean and standard deviation of the depth values ​​in the fused depth map.

[0110] In this step, the mean is the average level of all depth values ​​in the fused depth map. The mean reflects an average level of depth values ​​across the entire fused depth map and is an important reference indicator for subsequent analysis and segmentation. The standard deviation measures the dispersion of depth values ​​relative to the mean. It reflects the fluctuation of depth values ​​in the fused depth map. A larger standard deviation indicates that the distribution of depth values ​​around the mean is more dispersed, meaning that the differences in depth values ​​are greater; a smaller standard deviation indicates that the depth values ​​are relatively concentrated around the mean. By calculating the standard deviation, we can understand the changes in depth values ​​in the fused depth map, providing more comprehensive information for subsequent region segmentation based on depth range.

[0111] Step S22: Based on the preset near and far depth range thresholds, and combined with the mean and the standard deviation, divide the pixels in the fused depth map into near and far regions.

[0112] In this step, before segmentation, thresholds for near and far depth ranges need to be pre-defined. These thresholds are determined based on the specific application scenario and the definition of near and far. For example, in some scenarios, areas with a depth value less than 1 meter might be defined as near regions, and areas with a depth value greater than 5 meters might be defined as far regions. However, in practical applications, these thresholds may be adjusted based on factors such as the camera's measurement range and the distribution of objects.

[0113] The pixels are divided by combining the mean and standard deviation calculated in step S21 with a preset depth range threshold. There are various specific division methods, which will not be elaborated here.

[0114] Step S23: Extract the pixels of the near region and the far region respectively to generate a first near sub-image and a first far sub-image.

[0115] In this step, after dividing the pixels into regions, all pixels in the fused depth map are traversed to find those classified as near-field regions. These near-field pixels are then extracted according to their positional relationship in the fused depth map to form a new image, namely the first near-field sub-image. Thus, the first near-field sub-image only contains the pixels belonging to the near-field regions in the fused depth map and their corresponding depth information, facilitating subsequent separate processing and analysis of the depth information of the near-field regions.

[0116] Similarly, for pixels classified as distant regions, they are extracted according to their positional relationship in the fused depth map to form the first distant sub-map. This first distant sub-map contains the pixels belonging to the distant regions in the fused depth map and their depth information, providing an independent dataset for subsequent depth processing of the distant regions. This segmentation operation effectively divides the fused depth map into two sub-maps based on depth range, facilitating more targeted processing based on the different characteristics of the near and distant regions, such as viewpoint reprojection and depth completion.

[0117] In some embodiments, in step S3, the joint calibration parameters of the depth camera and the RGB camera include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients. By reprojecting the first near-field sub-image and the first far-field sub-image from the viewpoint, aligning them with the pixel coordinate system of the RGB image, a second near-field sub-image and a second far-field sub-image are generated. Both the depth camera and the RGB camera have intrinsic parameter matrices. These intrinsic parameter matrices describe the inherent properties of the camera, including information such as focal length and optical center coordinates. These parameters determine the geometric relationship between pixels on the image plane and points in the camera coordinate system. For example, during camera imaging, points in three-dimensional space are projected onto a two-dimensional image plane through the transformation of the intrinsic parameter matrix. For the depth camera, the intrinsic parameter matrix is ​​used to transform the three-dimensional points corresponding to its measured depth information onto the camera's image plane; for the RGB camera, the intrinsic parameter matrix is ​​used to convert three-dimensional points in the scene into pixels on the RGB image.

[0118] The intrinsic parameter matrix can usually be accurately obtained through camera calibration experiments, such as using a checkerboard or similar calibration board. By capturing images of the calibration board at different angles and positions, specific algorithms are used to calculate the values ​​of each parameter in the intrinsic parameter matrix. The extrinsic parameter matrix describes the relative position and pose relationship between the depth camera and the RGB camera. It consists of a rotation matrix R and a translation vector t. The rotation matrix R represents the rotation angle between the two cameras, describing the pose difference between them through three rotation axes (usually rotations around the x, y, and z axes); the translation vector t represents the relative translation distance between the two cameras.

[0119] Due to the optical characteristics of camera lenses, actual captured images may contain distortions, such as radial and tangential distortion. Radial distortion causes straight lines in the image to appear curved, while tangential distortion causes the image to appear tilted or distorted. Distortion coefficients are used to describe and correct these distortions.

[0120] For each pixel in the first near-field sub-image, its 3D coordinates in the depth camera coordinate system can be determined based on its depth value (assuming the origin of the depth camera coordinate system is the camera's optical center). Then, the intrinsic parameter matrix of the depth camera is used to transform the 3D point to coordinates on the image plane of the depth camera (this step may involve some geometric transformations and projection calculations).

[0121] Next, the coordinates on the depth camera image plane are transformed to the RGB camera coordinate system using an extrinsic parameter matrix. In this process, the effects of rotation and translation need to be considered to accurately map points from the depth camera coordinate system to the RGB camera coordinate system.

[0122] Finally, using the intrinsic parameter matrix and distortion coefficients of the RGB camera, the points in the RGB camera coordinate system are transformed to the pixel coordinate system of the RGB image, and distortion correction is performed. After these steps, each pixel in the first near-field sub-image is reprojected to a position aligned with the pixel coordinate system of the RGB image, thus generating the second near-field sub-image.

[0123] The reprojection process of the first distant sub-image is similar to that of the first near sub-image. Similarly, the 3D coordinates of each pixel in the depth camera coordinate system are first determined based on the depth value of the pixel in the first distant sub-image. Then, each pixel is reprojected into the pixel coordinate system of the RGB image through the transformation of the depth camera intrinsic matrix, the transformation of the extrinsic matrix, the transformation of the RGB camera intrinsic matrix, and distortion correction, thereby generating the second distant sub-image.

[0124] The second near-field sub-image and the second far-field sub-image generated by the viewpoint reprojection operation are aligned with the pixel coordinate system of the RGB image. This allows the depth information and the texture information of the RGB image to correspond accurately in space, facilitating subsequent operations such as depth completion and fusion.

[0125] Figure 3 This is a flowchart illustrating the steps of edge blending in an embodiment of the present invention. Figure 3 As shown, the step S5, which involves performing depth-continuity-based edge fusion on the dense depth map and the second distant sub-map, includes:

[0126] Step S51: Calculate the depth difference between adjacent pixels in the dense depth map and the second distant sub-map.

[0127] In this step, for each pixel in the dense depth map and the second distant sub-map, it is necessary to consider the situation of its neighboring pixels. Typically, a neighborhood is defined, such as a 3×3 or 5×5 pixel region centered on the current pixel (of course, the specific neighborhood size can be adjusted according to the actual situation).

[0128] For each neighboring pixel in the neighborhood, calculate the depth difference between the current pixel and that neighboring pixel. For example, for pixel (i,j) in the dense depth map and its neighboring pixel (i+m,j+n) (m and n are offsets determined according to the neighborhood range, such as -1 to 1 in a 3×3 neighborhood), calculate their depth difference Δd=|d(i,j)-d(i+m,j+n)|, where d(i,j) and d(i+m,j+n) are the depth values ​​of pixels (i,j) and (i+m,j+n), respectively.

[0129] The same method is applied to the second distant sub-image, calculating the depth difference between each pixel and its neighboring pixels. By calculating these depth differences, we can understand the changes in depth around each pixel, providing a basis for subsequent edge location determination.

[0130] Step S52: Determine the edge position of the fusion region based on the magnitude of the depth difference.

[0131] In this step, a depth difference threshold T is set. This threshold is determined based on the specific application scenario and the characteristics of the depth map. When the calculated depth difference between adjacent pixels is greater than the threshold T, it indicates that a significant change in depth has occurred at that location, and an edge is likely to exist.

[0132] Traverse all pixels and their neighborhoods in the dense depth map and the second distant sub-map. Based on the comparison between the depth difference and the threshold T, mark the positions of pixels whose depth difference is greater than the threshold. These marked positions constitute the edge positions of the fusion region.

[0133] For example, in a scene, the depth values ​​of adjacent pixels at the edge of an object may differ significantly. This method can accurately locate these edge positions so that they can be specially processed in the subsequent fusion process to ensure the continuity of the fused depth map at the edges.

[0134] Step S53: At the edge location, the depth values ​​of the dense depth map and the second distant sub-map are weighted and fused according to the principle of depth continuity to eliminate depth discontinuities at the edge.

[0135] In this step, the depth continuity principle means that in a real-world scene, the depth values ​​of adjacent objects or adjacent parts of the same object should change continuously without sudden jumps. This principle must be followed when blending at edges to ensure smoother depth changes in the blended depth map at the edges.

[0136] Weighted fusion: For a pixel at the edge position determined in step S52, a weighting coefficient is determined based on its depth value in the dense depth map and the second distant sub-map, as well as some related factors (such as the confidence level of the depth, the relationship with surrounding pixels, etc.).

[0137] There are several ways to determine the weighting coefficients. For example, they can be determined based on the confidence level of the depth values, with higher confidence levels corresponding to larger weighting coefficients. Alternatively, they can be determined based on the depth relationship between a pixel and its surrounding pixels, ensuring that the fused depth value better reflects the depth variation trend of the surrounding pixels. This weighted fusion method effectively eliminates depth discontinuities at edges, generating a final depth map that contains near-field high-density and far-field low-density depth information, making the depth map more accurately reflect the true depth distribution of the scene.

[0138] Figure 4 This is a schematic diagram of a depth completion system based on a multimodal depth camera according to an embodiment of the present invention.

[0139] like Figure 4 As shown, an embodiment of the present invention provides a depth completion system based on a multimodal depth camera, comprising:

[0140] The initial generation module is used to acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulated depth camera. It performs confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, it extracts the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map.

[0141] The reprojection module is used to reproject the full-range high-quality depth map and the auxiliary depth map, and align them with the RGB image at the pixel level.

[0142] The sub-image generation module is used to adaptively filter the auxiliary depth map based on RGB information and depth gradient, and fuse it with the full-range high-quality depth map to obtain a fused depth map. The fused depth map is then divided into a first near sub-image and a first far sub-image according to the statistical characteristics of the depth distribution.

[0143] A dense module is used to construct a deep learning-based depth completion network. The second near sub-image and the RGB image are input into the depth completion network for depth completion to generate a dense depth map.

[0144] The fusion module is used to perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

[0145] Specifically, the initial generation module uses a multimodal modulated depth camera to acquire depth maps at different distance ranges, including a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range. Through confidence-weighted and spatial consistency fusion techniques, these three depth maps are integrated into a high-quality, full-range depth map, which provides good depth information coverage across the overall scene. Simultaneously, depth point sets filtered out during the fusion process due to confidence levels falling within a specific range (below a first threshold but above a second threshold) are extracted to generate an auxiliary depth map, providing additional depth information for subsequent processing.

[0146] The reprojection module reprojects the full-scale high-quality depth map and auxiliary depth map, along with the RGB image, onto the coordinate system of the RGB image and performs pixel-level alignment. This operation ensures that the depth information and the RGB image correspond accurately in space, providing a foundation for subsequent depth processing based on RGB information.

[0147] The sub-image generation module performs adaptive filtering on the auxiliary depth map based on the texture features and depth gradient information of the RGB image, optimizing its quality. Then, it fuses the filtered auxiliary depth map with the full-range high-quality depth map to obtain a fused depth map. On the other hand, based on the depth distribution statistical characteristics (such as depth mean and variance) of the fused depth map, it divides the fused depth map into a first near-range sub-image and a first far-range sub-image to allow for more refined processing of depth information at different distance ranges.

[0148] A dense module constructs a deep learning-based deep completion network. This network is trained on a large amount of labeled data and learns the mapping relationship from RGB images and near sub-images to a dense depth map. The second near sub-image (either a further processed first near sub-image or used directly as the second near sub-image) and the RGB image are input into the trained network, which outputs a dense depth map with high depth density and accuracy within the near range.

[0149] The fusion module performs edge fusion based on depth continuity between the dense depth map and the second distant sub-map (the first distant sub-map is either further processed or used directly as the second distant sub-map). By analyzing the continuous changes in depth values ​​near the edges, the fusion weights are adaptively adjusted to ensure that the fused depth map has a smooth transition at the edges, ultimately generating a final depth map containing near-field high-density and far-field low-density depth information.

[0150] The high-quality full-range depth map and auxiliary depth map generated by the initial generation module are the inputs to the reprojection module. The reprojection module processes these depth maps to spatially align them with the RGB image, providing accurate depth data for subsequent depth processing based on RGB information.

[0151] The full-range high-quality depth map and auxiliary depth map processed by the reprojection module are input into the sub-map generation module. The sub-map generation module uses these aligned depth maps to perform adaptive filtering, fusion, and segmentation operations to generate the first near-range sub-map and the first far-range sub-map, providing depth information at different distance ranges for subsequent depth completion and fusion.

[0152] The first near sub-image generated by the sub-image generation module is one of the inputs to the dense module. The dense module inputs the first near sub-image and the RGB image into the depth completion network to generate a dense depth map, which further improves the density and accuracy of near depth information.

[0153] The dense depth map generated by the dense module and the first distant sub-map (which serves as the second distant sub-map) generated by the sub-map generation module are the inputs to the fusion module. The fusion module performs edge fusion based on depth continuity on these two depth maps to generate the final depth map, achieving effective integration of near-range high-density and far-range low-density depth information.

[0154] This embodiment generates a high-quality depth map across the entire range by performing confidence-weighted and spatial consistency fusion on dense depth maps at near distances, sparse depth maps at near-mid distances, and sparse depth maps at far distances. Combined with techniques such as adaptive filtering based on depth gradients, viewpoint reprojection, depth completion based on depth, and edge fusion based on depth continuity, it can effectively solve the problems existing in the current methods and improve the quality and completeness of the depth map.

[0155] This invention also provides a depth completion device based on a multimodal depth camera, including a processor and a memory storing executable instructions for the processor. The processor is configured to execute steps of a depth completion method based on a multimodal depth camera by executing the executable instructions.

[0156] As described above, this embodiment generates a high-quality depth map across the entire range by performing confidence weighting and spatial consistency fusion on dense depth maps at close range, sparse depth maps at near-mid range, and sparse depth maps at far range. Combined with techniques such as adaptive filtering based on depth gradient, viewpoint reprojection, depth completion based on depth, and edge fusion based on depth continuity, it can effectively solve the problems existing in the current methods and improve the quality and integrity of the depth map.

[0157] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "platform."

[0158] Figure 5 This is a schematic diagram of a depth completion device based on a multimodal depth camera according to an embodiment of the present invention. The following refers to... Figure 5 To describe an electronic device 600 according to this embodiment of the present invention. Figure 5 The electronic device 600 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0159] like Figure 5 As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.

[0160] The storage unit stores program code, which can be executed by the processing unit 610, causing the processing unit 610 to perform the steps described in the section on a depth completion method based on a multimodal depth camera described in this specification, according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform, as follows: Figure 1 The steps are shown in the figure.

[0161] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.

[0162] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a grid environment.

[0163] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0164] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 650. Furthermore, electronic device 600 can also communicate with one or more meshes (e.g., local area network (LAN), wide area network (WAN), and / or public meshes, such as the Internet) via mesh adapter 660. Mesh adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although... Figure 5 As not shown in the diagram, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0165] This invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a depth completion method based on a multimodal depth camera. In some possible implementations, various aspects of the invention can also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the above-described section of this specification regarding a depth completion method based on a multimodal depth camera, according to various exemplary embodiments of the invention.

[0166] As shown above, this embodiment generates a high-quality depth map across the entire range by performing confidence weighting and spatial consistency fusion on dense depth maps at close range, sparse depth maps at near-mid range, and sparse depth maps at far range. Combined with techniques such as adaptive filtering based on depth gradient, viewpoint reprojection, depth completion based on depth, and edge fusion based on depth continuity, it can effectively solve the problems existing in the current methods and improve the quality and integrity of the depth map.

[0167] Figure 6 This is a schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present invention. (Reference) Figure 6 As shown, a program product 800 for implementing the above-described method according to an embodiment of the present invention is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0168] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0169] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0170] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of mesh, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0171] This embodiment generates a high-quality depth map across the entire range by performing confidence-weighted and spatial consistency fusion on dense depth maps at near distances, sparse depth maps at near-mid distances, and sparse depth maps at far distances. Combined with techniques such as adaptive filtering based on depth gradients, viewpoint reprojection, depth completion based on depth, and edge fusion based on depth continuity, it can effectively solve the problems existing in the current methods and improve the quality and completeness of the depth map.

[0172] The various embodiments described in this specification are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0173] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A depth completion method based on a multimodal depth camera, characterized in that, include: Step S1: Acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulation depth camera. Perform confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, extract the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map. Step S2: Reproject the full-range high-quality depth map and the auxiliary depth map, and align them pixel-level with the RGB image; Step S3: Based on RGB information and depth gradient, perform adaptive filtering on the auxiliary depth map and fuse it with the full-range high-quality depth map to obtain a fused depth map. Then, according to the statistical characteristics of the depth distribution, divide the fused depth map into a first near-range sub-map and a first far-range sub-map. Step S4: Construct a deep learning-based depth completion network, input the second near sub-image and the RGB image into the depth completion network for depth completion, and generate a dense depth map; Step S5: Perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

2. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, The confidence weighting and spatial consistency fusion in step S1 specifically involves: weighting the depth values ​​based on the confidence of the three depth maps, while considering the neighborhood relationship of pixels in the spatial domain to ensure that the fused depth map has spatial consistency; and filtering and adjusting the depth values ​​by setting a consistency threshold to generate a full-range high-quality depth map.

3. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, In step S3, the adaptive filtering of the auxiliary depth map specifically involves: combining the gradient of the dense depth at close range with the gradient changes of the RGB image to generate a hybrid gradient weight, adaptively adjusting the filter parameters, performing smoothing filtering on the regions with smaller gradients, and performing detail-preserving filtering on the regions with larger gradients, thereby preprocessing the auxiliary depth map.

4. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, Step S3, which involves dividing the fused depth map into a first near-range sub-map and a first far-range sub-map, specifically includes: Step S21: Calculate the mean and standard deviation of the depth values ​​in the fused depth map; Step S22: Based on the preset near and far depth range thresholds, and combined with the mean and the standard deviation, divide the pixels in the fused depth map into near and far regions. Step S23: Extract the pixels of the near region and the far region respectively to generate a first near sub-image and a first far sub-image.

5. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, In step S3, the joint calibration parameters of the depth camera and the RGB camera include the intrinsic parameter matrix, the extrinsic parameter matrix, and the distortion coefficient. By reprojecting the first near-field sub-image and the first far-field sub-image from the viewpoint, they are aligned with the pixel coordinate system of the RGB image to generate the second near-field sub-image and the second far-field sub-image.

6. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, In step S4, when constructing the deep learning-based deep completion network, the network structure adopts an encoder-decoder structure. The encoder part uses a convolutional neural network to extract features from the RGB image and the second nearest sub-image. The decoder part uses a deconvolutional neural network to map the features back to the depth map space. Skip connections are introduced into the network to fuse features from different levels of the encoder with features from the corresponding levels of the decoder.

7. The depth completion method based on a multimodal depth camera according to claim 1, characterized in that, In step S5, edge fusion based on depth continuity is performed on the dense depth map and the second distant sub-map, specifically including: Step S51: Calculate the depth difference between adjacent pixels in the dense depth map and the second distant sub-map; Step S52: Determine the edge position of the fusion region based on the magnitude of the depth difference; Step S53: At the edge location, the depth values ​​of the dense depth map and the second distant sub-map are weighted and fused according to the principle of depth continuity to eliminate depth discontinuities at the edge.

8. A depth completion system based on a multimodal depth camera, used to implement the depth completion method based on a multimodal depth camera as described in any one of claims 1 to 7, characterized in that, include: The initial generation module is used to acquire a dense depth map at very close range, a sparse depth map at near-mid range, and a sparse depth map at far range using a multimodal modulated depth camera. It performs confidence weighting and spatial consistency fusion on the three depth maps to generate a high-quality depth map across the entire range. At the same time, it extracts the depth point set that was filtered out during the fusion process because its confidence was lower than the first threshold, but whose confidence was higher than the second threshold, to generate an auxiliary depth map. The reprojection module is used to reproject the full-range high-quality depth map and the auxiliary depth map, and align them with the RGB image at the pixel level. The sub-image generation module is used to adaptively filter the auxiliary depth map based on RGB information and depth gradient, and fuse it with the full-range high-quality depth map to obtain a fused depth map. The fused depth map is then divided into a first near sub-image and a first far sub-image according to the statistical characteristics of the depth distribution. A dense module is used to construct a deep learning-based depth completion network. The second near sub-image and the RGB image are input into the depth completion network for depth completion to generate a dense depth map. The fusion module is used to perform edge fusion based on depth continuity on the dense depth map and the second distant sub-map to generate a final depth map containing near-distance high-density and far-distance low-density depth information.

9. A depth completion device based on a multimodal depth camera, characterized in that, include: processor; A memory in which executable instructions of the processor are stored; The processor is configured to perform the steps of the depth completion method based on a multimodal depth camera according to any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that, When the program is executed, it implements the steps of the depth completion method based on a multimodal depth camera as described in any one of claims 1 to 7.