RGBD camera depth reconstruction method, system and device and storage medium

By using deep learning networks in RGBD cameras to generate dense depth maps and smoothly correct the ground through ground plane coefficients, the problems of insufficient depth reconstruction accuracy and high computing power requirements in the prior art are solved, and efficient and low-cost deep reconstruction effect is achieved.

CN120047518APending Publication Date: 2025-05-27SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411962046.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing depth reconstruction technology based on depth cameras has insufficient accuracy when dealing with ground scenes, especially when highlights or specular reflections exist, which leads to ground bending, affecting data accuracy and system performance. At the same time, existing algorithms have high requirements for computing power, resulting in increased hardware costs and energy consumption, limiting their application in scenarios such as mobile devices.

Method used

Based on the RGBD camera, a dense depth map with high resolution is generated using a deep learning network through sparse depth maps, and the ground plane is smoothly corrected by the ground plane coefficient in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, while reducing calculation costs.

Benefits of technology

It realizes efficient deep reconstruction at a low computing cost, ensures the flatness and stability of the ground, avoids ground bending, improves data accuracy and system performance, and is suitable for scenarios such as mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005217412970000093
    Figure BDA0005217412970000093
  • Figure BDA0005217412970000094
    Figure BDA0005217412970000094
  • Figure BDA0005217412970000101
    Figure BDA0005217412970000101
Patent Text Reader

Abstract

The invention discloses a depth reconstruction method, system and device for an RGBD camera, and a storage medium. The method comprises the steps: S1, obtaining an RGB image and a sparse depth image; s2, fusing the RGB image and the sparse depth image through a feature encoder to obtain fused features, and decoding the fused features through a feature decoder to obtain a super-divided dense depth image and a ground plane space coefficient; s3, converting the dense depth map into three-dimensional point cloud data according to internal reference; s4, fitting a ground plane according to the ground plane space coefficient and the three-dimensional point cloud data; s5, according to the ground plane, error points lower than the ground plane are obtained in the three-dimensional point cloud; and S6, finding two sites (x, y) of the corresponding image coordinates according to the error points, taking the depth z as an unknown parameter, calculating expressions of X and Y according to the internal reference, substituting the expressions into a plane formula, solving the depth z ', and obtaining correct three-dimensional coordinates (X, Y, Z).
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In recent years, with the rapid development of science and technology and the continuous growth of application needs in many fields, the demand for image-based depth reconstruction has shown an increasingly high trend. In many fields such as autonomous driving, virtual reality, augmented reality, and intelligent security monitoring, accurate and efficient depth reconstruction technology plays an extremely critical role. It can provide rich and accurate spatial information for these fields, helping the system to make more intelligent and accurate decisions and feedback. However, to successfully achieve deep reconstruction in the target scene, strictly ensure the flatness of the ground, and maintain its high stability between time series, this is undoubtedly a very challenging and arduous task at the technical level.

[0003] As far as the current state of the art is concerned, although the typical depth camera-based depth prediction model has achieved a certain degree of improvement in overall accuracy, its prediction effect is far from satisfactory when dealing with some ground scenes that are extremely common in daily life. Taking CRE as an example, although the model can achieve relatively good results in detail presentation of depth prediction, once it encounters special scenes such as highlights or mirror reflections on the ground, it will expose serious defects and the ground will easily bend. This ground bending will greatly interfere with various subsequent analyses and applications based on depth information, resulting in a significant reduction in the accuracy and reliability of the data, which in turn affects the performance of the entire system and the correctness of decision-making.

[0004] In addition, if the goal of real-time depth reconstruction in a variety of scenarios is to be achieved, there are currently many three-dimensional depth reconstruction algorithms based on depth cameras on the market. These algorithms theoretically provide a variety of ways and methods to achieve real-time depth reconstruction, but they all have a significant problem, which is that they have extremely high requirements for computing power. In actual application scenarios, this means that high-performance computing devices are required, such as powerful GPUs or professional computing clusters. This not only increases the hardware cost and energy consumption of the system, but also in some scenarios that require device size, weight and portability, such as mobile robots, head-mounted virtual reality devices, etc., the application of these algorithms will be greatly restricted and difficult to meet actual usage needs.

[0005] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0006] To this end, based on an RGBD camera, according to a sparse depth map, a high-resolution dense depth map is obtained using a deep learning network, and finally, through the smoothing of the ground plane coefficients in three-dimensional space, the ground plane is corrected, ensuring the flatness and stability of the ground during depth reconstruction while achieving the binocular depth reconstruction task at a relatively low computational cost.

[0007] In a first aspect, the present invention provides a depth reconstruction method for an RGBD camera, characterized by comprising:

[0008] Step S1: Obtain an RGB image and a sparse depth map;

[0009] Step S2: Fuse the RGB image and the sparse depth map through a feature encoder to obtain a fused feature, and decode the fused feature through a feature decoder to obtain a dense depth map and ground plane spatial coefficients; the depth value density of the dense depth map is higher than that of the sparse depth map;

[0010] Step S3: Convert the dense depth map into three-dimensional point cloud data according to the internal parameters;

[0011] Step S4: Fit a ground plane according to the ground plane spatial coefficients and the three-dimensional point cloud data;

[0012] Step S5: Obtain incorrect points below the ground plane in the three-dimensional point cloud according to the ground plane;

[0013] Step S6: Find the two-dimensional points (x, y) of the corresponding image coordinates according to the incorrect points, take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and find the depth z' to obtain the correct three-dimensional coordinates (X, Y, Z).

[0014] Optionally, the depth reconstruction method for an RGBD camera further comprises:

[0015] Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through the internal parameters to obtain a final depth prediction result.

[0016] Optionally, in step S2 of the depth reconstruction method for an RGBD camera, both the feature encoder and the feature decoder are convolutional neural networks.

[0017] Optionally, step S4 of the depth reconstruction method for an RGBD camera comprises:

[0018] Step S41: Obtain a first ground plane according to the ground plane spatial coefficients;

[0019] Step S42: Fit a second ground plane based on the three-dimensional point cloud data, and calculate the average deviation value between the point cloud and the second ground plane;

[0020] Step S43: Perform weighted averaging on the first ground plane and the second ground plane to obtain a third ground plane; wherein, the weight of the second ground plane is inversely proportional to the average deviation value.

[0021] Optionally, in the depth reconstruction method of an RGBD camera, the weighted sum of the third ground plane and the ground plane coefficients of the previous frame is used as the ground plane coefficient of the current frame.

[0022] Optionally, in the depth reconstruction method of an RGBD camera, the weights of the third ground plane and the ground plane coefficients of the previous frame are proportional to the corresponding point cloud numbers.

[0023] Optionally, in the depth reconstruction method of an RGBD camera, in step S6, according to the formula obtain the depth z', and further obtain

[0024] In a second aspect, the present invention provides a depth reconstruction system for an RGBD camera, which is used to implement the depth reconstruction method of an RGBD camera described in any one of the foregoing items, and is characterized in that it includes:

[0025] An acquisition module, which is used to acquire an RGB image and a sparse depth map;

[0026] A super-resolution module, which is used to fuse the RGB image and the sparse depth map through a feature encoder to obtain fused features, and decode the fused features through a feature decoder to obtain a super-resolved dense depth map and ground plane spatial coefficients; the depth value density of the dense depth map is higher than that of the sparse depth map;

[0027] A reconstruction module, which is used to convert the dense depth map into three-dimensional point cloud data according to the internal parameters;

[0028] A fitting module, which is used to fit a ground plane according to the ground plane spatial coefficients;

[0029] A screening module, which is used to obtain incorrect points below the ground plane in the three-dimensional point cloud according to the ground plane;

[0030] A correction module, which is used to find the two-dimensional points (x, y) of the corresponding image coordinates according to the incorrect points, take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and find the depth z' to obtain the correct three-dimensional coordinates (X, Y, Z).

[0031] In a third aspect, the present invention provides a depth reconstruction device for an RGBD camera, characterized by comprising:

[0032] a processor;

[0033] a memory storing executable instructions of the processor;

[0034] wherein the processor is configured to execute the steps of the depth reconstruction method of the RGBD camera described in any one of the foregoing by executing the executable instructions.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of the depth reconstruction method of the RGBD camera described in any one of the foregoing are implemented.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] Based on the RGBD camera, the present invention uses a deep learning network to obtain a high-resolution dense depth map according to the sparse depth map, and finally corrects the ground plane through the ground plane coefficient smoothing in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and realizing the binocular depth reconstruction task at a low computational cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more obvious:

[0039] Figure 1 is a flowchart of the steps of a depth reconstruction method for an RGBD camera in an embodiment of the present invention;

[0040] Figure 2 is a flowchart of the steps of another depth reconstruction method for an RGBD camera in an embodiment of the present invention;

[0041] Figure 3 is a flowchart of the steps of fitting a ground plane in an embodiment of the present invention;

[0042] Figure 4 is a network framework diagram combining time series in an embodiment of the present invention;

[0043] Figure 5Schematic structural diagram of a depth reconstruction system for an RGBD camera in an embodiment of the present invention;

[0044] Figure 6 Schematic structural diagram of a depth reconstruction device for an RGBD camera in an embodiment of the present invention; and

[0045] Figure 7 Schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. Detailed implementation manners

[0046] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0047] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0048] A depth reconstruction method for an RGBD camera provided by an embodiment of the present invention aims to solve the problems existing in the prior art.

[0049] The technical solutions of the present invention and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.

[0050] Based on the RGBD camera, according to the sparse depth map, the present invention uses a deep learning network to obtain a high-resolution dense depth map, and finally corrects the ground plane through the ground plane coefficient smoothing in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and realizing the binocular depth reconstruction task at a low computational cost.

[0051] Figure 1This is a flowchart of the steps of a depth reconstruction method for an RGBD camera in an embodiment of the present invention. As Figure 1 shown, the steps of a depth reconstruction method for an RGBD camera in an embodiment of the present invention include:

[0052] Step S1: Obtain an RGB image and a sparse depth map.

[0053] In this step, the color image information of the scene is collected by the image sensor in the RGBD camera. This image contains the appearance features such as the color and texture of the objects in the scene, and represents the color conditions of each pixel point in the scene with the data of three color channels: Red, Green, and Blue. It can visually present the visual appearance of the scene and provide rich surface feature information for subsequent processing. The RGBD camera uses its own depth measurement principle (such as structured light, time-of-flight, etc.) to obtain the depth information of some pixel points in the scene. Due to the characteristics of the object surface, these depth information are relatively discrete and discontinuous, that is, in a so-called sparse state. The sparse depth map reflects the distance relationship between some key positions in the scene and the camera, but the overall covered depth data is not dense enough in spatial distribution, and there are many areas where the depth value or effective depth value is not obtained.

[0054] Step S2: Fuse the RGB image and the sparse depth map through a feature encoder to obtain a fused feature, and decode the fused feature through a feature decoder to obtain a super-resolved dense depth map and a ground plane space coefficient; the depth value density of the dense depth map is higher than that of the sparse depth map.

[0055] In this step, the feature encoder is a neural network structure or a specific algorithm module, which receives the RGB image and the sparse depth map as inputs. For the RGB image, it extracts different levels of visual features such as object edges and texture patterns; for the sparse depth map, it mines the spatial layout features contained in the existing depth values, etc. Then, through a specific fusion mechanism (such as pixel-by-pixel feature splicing, feature weighted fusion based on the attention mechanism, etc.), the appearance features from the RGB image and the depth-related features of the sparse depth map are organically integrated to form a fused feature. Such a fused feature contains both the visual representation information of the scene and the existing depth clues, preparing for generating more complete depth information later.

[0056] The feature decoder is also part of the neural network structure. It decodes the fused features output by the feature encoder as input. Its purpose is to restore more complete depth information of the scene and the ground plane spatial coefficients through a series of operations such as transformation and mapping on the fused features. Specifically, compared with the input sparse depth map, the dense depth map obtained after decoding has a significantly increased depth value density, which means that more pixel points are assigned depth values and can more delicately reflect the depth distribution of objects in the scene in space. The ground plane spatial coefficients are a set of parameters related to the ground plane in the scene (such as the coefficients of the ground plane equation, etc.), which characterize the attitude, position, etc. of the ground plane in the entire three-dimensional space. Subsequently, accurate fitting and other operations can be performed on the ground plane based on these coefficients.

[0057] The structures of the feature encoder and the feature decoder can be any structure that meets the requirements of this embodiment. For illustrative purposes only, an exemplary structure of the feature encoder and the feature decoder is provided here.

[0058] Feature Encoder Structure

[0059] Convolution Layers

[0060] The convolution layer is one of the core components of the feature encoder. When processing RGB images and sparse depth maps, it performs convolution operations by sliding a convolution kernel (a small filter matrix) over the image. For example, for an RGB image, a typical convolution kernel size may be 3×3 or 5×5, and the number of channels can be set as needed. Assuming the input RGB image size is H×W×3 (H is the image height, W is the image width, and 3 is the three RGB channels), the convolution kernel size is 3×3, and the number of output channels is 16. After the convolution operation, the size of the output feature map will change according to parameters such as the convolution stride, and at the same time, the depth (number of channels) becomes 16.

[0061] For the sparse depth map, the convolution layer can also be used to extract features. Since the depth map may have only a single channel (representing depth values), the convolution kernel can extract features of depth changes, such as depth edges, on this single-channel depth map. The convolution layer can automatically extract local features in the image, such as textures and edges. By stacking multiple convolution layers, more abstract and higher-level features can be gradually extracted.

[0062] Pooling Layers

[0063] Pooling layers are mainly used to reduce the data dimension while retaining important feature information. Common pooling methods include max pooling and average pooling. For example, max pooling selects the maximum value within a small region (such as a 2×2 window) as the output. If the size of the input feature map is H×W×C (C is the number of channels), after passing through a 2×2 max pooling layer, the size of the output feature map becomes H / 2×W / 2×C, which can reduce the amount of data to a certain extent and has a certain invariance to transformations such as translation and rotation. When processing the features after fusing RGB images and depth maps, the pooling layer helps to extract more robust features and remove some minor variations and noises.

[0064] Normalization Layers

[0065] Batch Normalization is a commonly used normalization method. In a feature encoder, it can normalize the input of each layer, making the distribution of the input data more stable. For example, in a small batch of data, Batch Normalization calculates the mean and variance for each feature channel and then normalizes the data to a standard normal distribution with a mean of 0 and a variance of 1. This helps to accelerate the training process, improve the convergence speed and generalization ability of the model. During the process of fusing RGB and depth map features, since the distributions of the two types of data may be different, the normalization layer can enable them to work better together in subsequent processing.

[0066] Activation Functions

[0067] Activation functions are used to introduce non - linear factors into neurons. In a feature encoder, commonly used activation functions such as ReLU (Rectified Linear Unit). The ReLU function is defined as, which can make the neuron output retain only the positive part, increasing the non - linear expression ability of the network. When the features extracted by the convolutional layer pass through the activation function, they can better fit complex functional relationships. For example, when processing the fusion of texture features of RGB images and spatial features of depth maps, the activation function can help the network learn more complex associations between the two types of features.

[0068] Feature decoder structure

[0069] Deconvolution Layers or Transposed Convolution Layers

[0070] The deconvolution layer is an important component of the feature decoder for restoring the image size. It operates in the opposite way of the convolution layer and can upsample a low-resolution feature map into a high-resolution feature map. For example, if the size of the input feature map is H×W×C, through appropriate deconvolution kernels and parameter settings, the output feature map can become 2H×2W×C' (C' may be different from C, depending on the settings of the deconvolution layer). In the process of restoring the dense depth map and the ground plane spatial coefficients from the fused features, the deconvolution layer can gradually restore the features into a depth map and coefficient representation similar to the size of the original image. The transposed convolution layer is similar to the deconvolution layer in principle and is also an operation for upsampling.

[0071] Upsampling Layers

[0072] In addition to the deconvolution layer, there are some simple upsampling methods, such as bilinear interpolation upsampling. It calculates the pixel values of the high-resolution feature map by interpolating the pixel values in the low-resolution feature map. Suppose the coordinates of a point on the low-resolution feature map are, bilinear interpolation will calculate the weighted average of the pixel values of the four surrounding points of this point to obtain the corresponding pixel value on the high-resolution feature map. This method can play an auxiliary role in restoring the details of the depth map and, when combined with the deconvolution layer, etc., can better reconstruct a high-quality dense depth map.

[0073] Residual Connections

[0074] In the feature decoder, residual connections can effectively solve the problem of vanishing gradients and improve the performance of the network. Its basic idea is to directly add the input features to the output of some layers in the decoder. For example, suppose the input of a module in the decoder is, and after a series of complex operations, the output is obtained. Through the residual connection, the output can be changed to. In this way, during the backpropagation process, the gradients can be more easily passed through these connections, making the network easier to train. Especially in tasks such as depth map reconstruction that require precise restoration of details, residual connections help to better restore high-quality depth information and ground plane coefficients.

[0075] Step S3: Convert the dense depth map into three-dimensional point cloud data according to the internal parameters.

[0076] In this step, the internal parameters of the RGBD camera include parameter information such as focal length and principal point coordinates. These parameters describe the geometric relationship of camera imaging and determine the conversion rule from image plane coordinates to coordinates in the camera coordinate system. For each pixel point in the dense depth map, given its two-dimensional coordinates (u, v) in the image coordinate system and the corresponding depth value d (provided by the dense depth map), using the mathematical relationships defined by the camera internal parameter matrix (usually based on the projection formula of the pinhole camera model and its inverse transformation, etc.), the corresponding three-dimensional coordinates (X, Y, Z) of this pixel point in the camera coordinate system can be calculated. Summing up the three-dimensional coordinates of all pixel points forms the three-dimensional point cloud data that can represent the spatial structure of the scene. The point cloud data presents the actual spatial positions of objects in the scene in the form of numerous discrete three-dimensional points and is the basis for subsequent operations such as spatial analysis and plane fitting.

[0077] Step S4: Fit a ground plane according to the ground plane spatial coefficient and the three-dimensional point cloud data.

[0078] In this step, the ground plane can be represented by a plane equation under ideal circumstances (for example, the general form in a three-dimensional Cartesian coordinate system is Ax + By + Cz + D = 0). The ground plane spatial coefficient obtained previously can yield a plane. Using the points in the three-dimensional point cloud data that belong to the ground plane part (it can be roughly judged which points may be on the ground plane through some prior knowledge or based on spatial distribution characteristics, such as points in the area near the bottom and relatively flat areas, etc.), adopting a suitable fitting algorithm (such as the least squares method, etc.), substituting the coordinates of these points into the plane equation, and through the way of optimization and solution, determining the plane equation parameters that best conform to the distribution law of these points, thus obtaining another plane by fitting the point cloud. Fusing these two planes gives the final ground plane.

[0079] Step S5: Obtain the incorrect points below the ground plane in the three-dimensional point cloud according to the ground plane.

[0080] In this step, taking the fitted ground plane as the reference plane, due to possible errors and other situations in the actual depth measurement and reconstruction process, the positions represented by some point cloud data do not conform to the actual spatial logic. For example, the point cloud points corresponding to some objects that should be above the ground plane appear below the ground plane (possibly caused by reasons such as incorrect depth measurement and noise interference). By comparing the Z coordinate (height direction coordinate) of each point cloud point with the Z coordinate value of the ground plane determined by the ground plane equation, those points with Z coordinates less than the Z coordinate of the corresponding position of the ground plane can be screened out. These points are the incorrect points below the ground plane, and their existence will affect the subsequent accurate spatial understanding and application of the scene.

[0081] Step S6: Find the two-dimensional points (x, y) of the corresponding image coordinates according to the error points. Take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and find the depth z', so as to obtain the correct three-dimensional coordinates (X, Y, Z).

[0082] In this step, since the error points were originally converted from the image to the three-dimensional point cloud through a series of processes such as depth reconstruction, the two-dimensional coordinates (x, y) corresponding to these error points in the original image can be found by reverse lookup based on the previous conversion mapping relationship. According to the geometric projection principles such as the camera internal parameters and the pinhole camera model, for the known image coordinates (x, y), taking the depth value z as an unknown parameter temporarily, the expressions of the X and Y coordinates of this point in the camera coordinate system with respect to the depth z can be deduced (for example, deduced based on the principle of similar triangles, the projection relationship defined by the internal parameters, etc.). Then substitute these expressions of X and Y with respect to z into the previously fitted ground plane equation (the form of the known plane equation, such as Ax + By + Cz + D = 0). At this time, only the depth z is the unknown quantity in the equation, and the correct depth value z' can be obtained by solving the equation. Finally, substitute the obtained z' into the previously deduced expressions of X and Y, so as to obtain the correct three-dimensional coordinates (X, Y, Z) of this point. By performing such correction operations on all error points, the entire three-dimensional point cloud data can be further optimized, and the accuracy of the depth reconstruction result can be improved.

[0083] In some embodiments, according to the formula the depth z' is obtained, and then After finding the error points, find the two-dimensional points (x, y) of the corresponding image coordinates according to the error points. Take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and inversely find the correct z, and then deduce the correct three-dimensional coordinates (X, Y, Z) according to z. The formula process is as follows:

[0084]

[0085]

[0086] Z = z

[0087]

[0088]

[0089] Figure 2 is the step flowchart of another depth reconstruction method of the RGBD camera in the embodiments of the present invention. As Figure 2 shown, compared with the previous embodiments, another depth reconstruction method of the RGBD camera in the embodiments of the present invention further includes:

[0090] Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through the internal parameters to obtain the final depth prediction result.

[0091] In this step, for the corrected correct three-dimensional coordinates (X, Y, Z), use the projection matrix determined by the camera internal parameters (this matrix can be constructed according to the internal parameter values, which reflects the conversion relationship from three-dimensional coordinates to two-dimensional image coordinates), and perform operations according to the corresponding projection calculation formula. Specifically, generally through linear algebra operations such as matrix multiplication, multiply the three-dimensional coordinates (X, Y, Z) by the projection matrix to obtain the corresponding coordinate values on the two-dimensional plane. The two-dimensional plane coordinates obtained after projection, combined with the corresponding depth value z', jointly constitute the final depth prediction result. This result not only contains the position information of the object in the two-dimensional image but also carries accurate depth information (reflected by the depth value z'), which can more comprehensively and accurately describe the spatial position and depth distribution of the object relative to the camera in the scene, and can be used in many subsequent related application scenarios such as object detection, scene understanding, and 3D reconstruction visualization, providing a high-quality data basis for further analysis and processing.

[0092] Reproject the three-dimensional points back to the two-dimensional plane through the camera internal parameters to obtain the final depth map prediction result. The formula is as follows:

[0093] z = Z

[0094]

[0095]

[0096] Figure 3 is a flowchart of the steps for fitting a ground plane in an embodiment of the present invention. As Figure 3 shown, the steps for fitting a ground plane in an embodiment of the present invention include:

[0097] Step S41: Obtain a first ground plane according to the ground plane spatial coefficient.

[0098] In this step, based on the ground plane spatial coefficient, a theoretical ground plane can be directly determined, which is the first ground plane mentioned here. It has a clear position and attitude in three-dimensional space, representing the ground plane situation initially estimated according to the previous feature processing.

[0099] Step S42: Fit a second ground plane according to the three-dimensional point cloud data and calculate the average deviation value between the point cloud and the second ground plane.

[0100] In this step, using the acquired three-dimensional point cloud data, filter out those points that may belong to the ground plane. This can be judged based on some prior knowledge or spatial distribution characteristics. For example, points at the bottom of the scene and in relatively flat areas are likely to belong to the ground plane part. Then, adopt a suitable fitting algorithm, such as the least squares method commonly used. Substitute the coordinates of these filtered points into the plane equation, and determine the parameters of the plane equation through an optimization solution method, thereby fitting out the second ground plane. This process is to re-determine the position and attitude of the ground plane based on the actual point cloud data, making it more conform to the ground plane situation in the real scene.

[0101] After fitting out the second ground plane, it is necessary to measure the degree of conformity between the point cloud data and this second ground plane, which is reflected by calculating the average deviation value. For each point in the point cloud data, substitute its coordinates into the equation of the second ground plane to calculate the corresponding deviation value (which can be calculated according to the distance formula from a point to a plane, etc.). Then sum up the deviation values of all points and divide by the total number of points to obtain the average deviation value. The smaller the average deviation value, the more consistent the point cloud data is with the fitted second ground plane; on the contrary, it indicates a larger deviation.

[0102] Step S43: Perform weighted averaging on the first ground plane and the second ground plane to obtain a third ground plane; wherein, the weight of the second ground plane is inversely proportional to the average deviation value.

[0103] In this step, considering that the first ground plane is determined based on the coefficients obtained from the previous feature decoder and has a certain theoretical basis; while the second ground plane is fitted according to the actual point cloud data and is closer to the actual scene situation, but there may be certain fitting errors (reflected by the average deviation value). In order to integrate the advantages of both, a weighted averaging method is adopted to obtain a more accurate ground plane, that is, the third ground plane.

[0104] It is stipulated that the weight of the second ground plane is inversely proportional to the average deviation value, that is, the smaller the average deviation value, the more reliable the second ground plane and the greater its weight; on the contrary, the larger the average deviation value, the smaller its weight. In this embodiment, the equation of the third ground plane in three-dimensional space is determined, obtaining a final ground plane representation that is more in line with the actual situation and integrates various aspects of information, providing a more accurate basis for subsequent operations such as error point judgment based on the ground plane.

[0105] Figure 4 It is a network framework diagram combining time series in an embodiment of the present invention. As Figure 4 shown, this embodiment considers the association between the data of the previous frame and the data of the current frame. Figure 4 The model in [] is illustrated by taking PostNet as an example. t represents the current frame, and t - 1 represents the previous frame.

[0106] In step S4, the third ground plane is weighted and added to the ground plane coefficient of the previous frame as the ground plane coefficient of the current frame. The predicted ground plane spatial coefficients (a t , b t , c t , d t ) of the current frame are weighted and added to the ground plane coefficients (a m-1 , b m-1 , c m-1 , d m-1 ) stored in the memory to obtain the ground plane coefficients (a m , b m , c m , d m ) of the current frame, which are then stored in the memory for use in the next frame, as shown on the Figure 4 right side.

[0107] In practical application scenarios, such as when a camera continuously captures images and performs depth reconstruction and other operations in a dynamic environment, the ground planes between adjacent frames often have a certain continuity and correlation. By weighted adding the ground plane spatial coefficients calculated for the current frame to the ground plane coefficients of the previous frame to determine the ground plane coefficients of the current frame, it helps to utilize historical information to stabilize the estimation of the ground plane, reducing the inaccurate ground plane fitting caused by factors such as noise and local anomalies that may exist in single-frame data, making the determination of the ground plane smoother and more reliable in continuous video sequences or dynamic scenes.

[0108] Appropriate weight values are set according to actual application requirements, experience, etc., which are the weights for the coefficients of the current frame and the previous frame respectively, and usually satisfy certain conditions to ensure the rationality of the result after weighted addition. For example, the weights can be adjusted according to factors such as the expected speed of scene change and the stability of camera movement. If the scene is relatively stable and the camera movement is gentle, can be appropriately increased to make the coefficients of the previous frame play a greater role; conversely, if the scene changes rapidly, can be appropriately increased to give more emphasis to the coefficients calculated for the current frame. Calculate according to the rule of weighted addition to obtain the final ground plane coefficients of the current frame.

[0109] Through such a weighted addition method, the correlation of ground plane information between different frames can be fully utilized during the processing of consecutive image frames, improving the robustness and accuracy of the entire depth reconstruction and plane correction methods.

[0110] In some embodiments, the weight of the ground plane spatial coefficient and the ground plane coefficient of the previous frame is proportional to the corresponding number of point clouds. When fusing the ground plane coefficients using multiple frames of data, the number of point clouds is a crucial consideration factor. To a certain extent, the number of point clouds reflects the reliability of the current frame and the previous frame of data for ground plane estimation, or rather, the amount of information. Generally speaking, the more the number of point clouds, the richer the data for fitting the ground plane, and the more representative and accurate the calculated ground plane coefficient. Therefore, establishing a proportional relationship between the weight and the corresponding number of point clouds aims to make the coefficients obtained from a larger amount of point cloud data play a more important role in finally determining the current frame's ground plane coefficient when fusing the ground plane coefficients.

[0111] First, it is necessary to separately count the number of point clouds used to fit the ground plane in the current frame and the number of point clouds used to fit the ground plane in the previous frame. This counting process is usually completed by counting the point clouds that are screened out and may belong to the ground area before performing the ground plane fitting operation.

[0112] According to the principle that the weight is proportional to the number of point clouds, calculate their respective weights. When calculating the weights, first find the sum of the number of point clouds in the two frames, and then determine the weight of the ground plane spatial coefficient of the current frame and the weight of the ground plane coefficient of the previous frame in the following manner.

[0113] Through such a weight setting and weighted calculation process, the difference in data information content reflected by the number of point clouds in different frames is fully considered, making the fusion of the ground plane coefficients more reasonable and scientific, further improving the accuracy and stability of ground plane estimation during continuous multi-frame processing in a dynamic scene, and providing a more reliable basis for subsequent operations such as depth correction based on the ground plane.

[0114] Figure 5 This is a schematic structural diagram of a depth reconstruction system of an RGBD camera in an embodiment of the present invention. As Figure 5 shown, a depth reconstruction system of an RGBD camera in an embodiment of the present invention includes:

[0115] An acquisition module, configured to acquire an RGB image and a sparse depth map;

[0116] A super-resolution module, configured to fuse the RGB image and the sparse depth map through a feature encoder to obtain a fused feature, and decode the fused feature through a feature decoder to obtain a super-resolved dense depth map and a ground plane spatial coefficient; the depth value density of the dense depth map is higher than that of the sparse depth map;

[0117] A reconstruction module, configured to convert the dense depth map into three-dimensional point cloud data according to the internal parameters;

[0118] A fitting module, configured to fit a ground plane according to the ground plane spatial coefficient and the three-dimensional point cloud data;

[0119] A screening module, configured to obtain incorrect points below the ground plane from the three-dimensional point cloud according to the ground plane;

[0120] A correction module, configured to find the two-dimensional points (x, y) of the corresponding image coordinates according to the incorrect points, use the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, find the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).

[0121] This embodiment is based on an RGBD camera. According to the sparse depth map, a high-resolution dense depth map is obtained by using a deep learning network. Finally, the ground plane is corrected by smoothing the ground plane coefficients in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and achieving the binocular depth reconstruction task at a low computational cost.

[0122] In an embodiment of the present invention, there is also provided a depth reconstruction device for an RGBD camera, including a processor and a memory, in which executable instructions of the processor are stored. Wherein, the processor is configured to execute the steps of a depth reconstruction method for an RGBD camera by executing the executable instructions.

[0123] As above, this embodiment is based on an RGBD camera. According to the sparse depth map, a high-resolution dense depth map is obtained by using a deep learning network. Finally, the ground plane is corrected by smoothing the ground plane coefficients in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and achieving the binocular depth reconstruction task at a low computational cost.

[0124] Those skilled in the art to which the present invention pertains can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "platform" here.

[0125] Figure 6 It is a schematic structural diagram of a depth reconstruction device for an RGBD camera in an embodiment of the present invention. The following will refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present invention. Figure 6 The displayed electronic device 600 is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present invention.

[0126] Such as Figure 6As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0127] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the part of the depth reconstruction method of an RGBD camera above in this specification. For example, the processing unit 610 can execute as Figure 1 shown in.

[0128] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0129] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a grid environment.

[0130] The bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the multiple bus structures.

[0131] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 650. And, the electronic device 600 can also communicate with one or more grids (such as a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 6Not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0132] An embodiment of the present invention also provides a computer-readable storage medium for storing a program, and the steps of a depth reconstruction method of an RGBD camera are implemented when the program is executed. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned part of the depth reconstruction method of an RGBD camera in this specification.

[0133] As shown above, based on the RGBD camera, according to the sparse depth map, a high-resolution dense depth map is obtained by using a deep learning network. Finally, through the ground plane coefficient smoothing in the three-dimensional space, the ground plane is corrected, ensuring the flatness and stability of the ground during depth reconstruction, and achieving the binocular depth reconstruction task at a relatively low computational cost.

[0134] Figure 7 is a schematic structural diagram of the computer-readable storage medium in the embodiment of the present invention. Refer to Figure 7 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0135] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0136] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0137] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0138] This embodiment is based on an RGBD camera. According to the sparse depth map, a high-resolution dense depth map is obtained using a deep learning network. Finally, through the smoothing of the ground plane coefficients in three-dimensional space, the ground plane is corrected, ensuring the flatness and stability of the ground during depth reconstruction while achieving the binocular depth reconstruction task at a relatively low computational cost.

[0139] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0140] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A depth reconstruction method for an RGBD camera, characterized in that: include: Step S1: Obtain RGB image and sparse depth image; Step S2: fusing the RGB image and the sparse depth image through a feature encoder to obtain a fused feature, and decoding the fused feature through a feature decoder to obtain a dense depth image and a ground plane spatial coefficient; the depth value density of the dense depth image is higher than that of the sparse depth image; Step S3: converting the dense depth map into three-dimensional point cloud data according to the internal reference; Step S4: fitting a ground plane according to the ground plane spatial coefficient and the three-dimensional point cloud data; Step S5: obtaining, according to the ground plane, error points below the ground plane in the three-dimensional point cloud; Step S6: Find the corresponding two-point point (x, y) of the image coordinates according to the error point, take the depth z as the unknown parameter, calculate the expression of X and Y according to the internal parameter, substitute the expression into the plane formula, calculate the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).

2. The depth reconstruction method of an RGBD camera according to claim 1, characterized in that: Also includes: Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through internal references to obtain a final depth prediction result.

3. The depth reconstruction method of an RGBD camera according to claim 1, characterized in that: The feature encoder and the feature decoder in step S2 are both convolutional neural networks.

4. The depth reconstruction method of an RGBD camera according to claim 1, characterized in that: Step S4 includes: Step S41: obtaining a first ground plane according to the ground plane space coefficient; Step S42: fitting a second ground plane according to the three-dimensional point cloud data, and calculating an average deviation value between the point cloud and the second ground plane; Step S43: performing weighted averaging on the first ground plane and the second ground plane to obtain a third ground plane; wherein the weight of the second ground plane is inversely proportional to the average deviation value.

5. The depth reconstruction method of an RGBD camera according to claim 4, characterized in that: The third ground plane is weightedly added to the ground plane coefficient of the previous frame to obtain the ground plane coefficient of the current frame.

6. The depth reconstruction method of an RGBD camera according to claim 5, characterized in that: The weights of the third ground plane coefficients and the ground plane coefficients of the previous frame are proportional to the corresponding number of point clouds.

7. The depth reconstruction method of an RGBD camera according to claim 1, characterized in that: In step S6, according to the formula Obtain the depth z', and then obtain 8. A depth reconstruction system for an RGBD camera, used to implement the depth reconstruction method for an RGBD camera according to any one of claims 1 to 7, characterized in that: include: Acquisition module, used to obtain RGB image and sparse depth image; A super-resolution module is used to fuse the RGB image and the sparse depth image through a feature encoder to obtain a fusion feature, and decode the fusion feature through a feature decoder to obtain a super-resolved dense depth map and a ground plane spatial coefficient; the depth value density of the dense depth map is higher than that of the sparse depth map; A reconstruction module, used for converting the dense depth map into three-dimensional point cloud data according to an internal parameter; A fitting module, used for fitting a ground plane according to the ground plane spatial coefficient and the three-dimensional point cloud data; A screening module, used for obtaining, in the three-dimensional point cloud, erroneous points below the ground plane according to the ground plane; The correction module is used to find the corresponding two-point point (x, y) of the image coordinates according to the error point, take the depth z as an unknown parameter, calculate the expression of X and Y according to the internal parameter, substitute the expression into the plane formula, calculate the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).

9. A depth reconstruction device for an RGBD camera, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to perform the steps of the depth reconstruction method of the RGBD camera according to any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the depth reconstruction method of the RGBD camera described in any one of claims 1 to 7 are implemented.