Depth reconstruction, super-resolution and plane correction method, system and device of binocular camera and storage medium
Through the depth reconstruction method based on binocular camera, the deep learning network is used to supersegment the high-resolution depth map and perform plane correction, which solves the problems of poor ground scene prediction effect and high computing power requirements in the prior art, and realizes the smoothness and stability of the ground and the effect of real-time depth reconstruction.
Patent Information
- Application Number
- CN202411961955.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The existing bi-purpose depth prediction model has poor prediction effects in common ground scenarios, especially in highlights, mirrors and other scenarios, which will lead to ground bending. At the same time, it requires high computing power, making it difficult to achieve real-time depth reconstruction.
Using a depth reconstruction method based on a binocular camera, a low-resolution depth map is obtained by obtaining the parallax of the left and right images, and a deep learning network is used to super-segment the high-resolution depth map and perform plane correction to ensure the flatness and stability of the ground.
It realizes the leveling and stability of the ground during deep reconstruction at a low computing cost, while improving the real-time deep reconstruction capability in a variety of scenarios.
Smart Images

Figure CN120047516A_ABST
Abstract
Description
Background Art
[0002] In recent years, the demand for depth reconstruction based on images has been increasing. It is a great technical challenge to achieve depth reconstruction within a target scene, ensure the flatness of the ground, and maintain its stability over time.
[0003] Currently, for typical binocular depth prediction models, although the accuracy has been improved, the prediction effect for some common ground scenes is very poor. For example, for CRE, although the effect is better in the details of depth prediction, serious ground bending will occur during the reconstruction of scenes such as highlights and mirrors on the ground.
[0004] In addition, to achieve real-time depth reconstruction in diverse scenes, there are currently many binocular three-dimensional depth reconstruction algorithms, but they all require high computing power.
[0005] The disclosure of the above background art content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0006] Therefore, based on a binocular camera, the present invention uses the left and right images to obtain a low-resolution depth map, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the ground plane coefficient smoothing in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction while achieving the binocular depth reconstruction task at a low computational cost.
[0007] In a first aspect, the present invention provides a method for depth reconstruction, super-resolution, and plane correction of a binocular camera, which is characterized by including:
[0008] Step S1: Obtain the left image and the right image, and calculate the disparity between the left image and the right image to obtain a first depth map;
[0009] Step S2: Fuse the left image and the first depth map, infer a fusion feature, and decode the fusion feature through a feature decoder to obtain a super-resolved second depth map and a ground plane spatial coefficient; the resolution of the second depth map is higher than that of the first depth map;
[0010] Step S3: Convert the second depth map into three-dimensional point cloud data according to the internal parameters;
[0011] Step S4: Fit out the ground plane;
[0012] Step S5: Obtain error points below the ground plane according to the ground plane;
[0013] Step S6: Find the two-dimensional points (x, y) of the corresponding image coordinates according to the error point. Take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and find the depth z' to obtain the correct three-dimensional coordinates (X, Y, Z).
[0014] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, it further includes:
[0015] Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through the internal parameters to obtain the final depth prediction result.
[0016] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, the SGM algorithm is used in Step S1 to obtain the first depth map.
[0017] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, in Step S4, the ground plane spatial coefficient is weighted and added to the ground plane coefficient of the previous frame as the ground plane coefficient of the current frame.
[0018] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, the weights of the ground plane spatial coefficient and the ground plane coefficient of the previous frame are proportional to the corresponding point cloud numbers.
[0019] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, in Step S5, the three-dimensional point cloud data is substituted into the plane formula aX + bY + cZ + d = 0 for calculation, and the points below the ground plane are filtered out as the error points.
[0020] Optionally, in the method for depth reconstruction, super-resolution, and plane correction of a binocular camera, in Step S6, according to the formula obtain the depth z', and further obtain
[0021] In a second aspect, the present invention provides a system for depth reconstruction, super-resolution, and plane correction of a binocular camera, which is used to implement the method for depth reconstruction, super-resolution, and plane correction of a binocular camera described in any one of the foregoing items. It is characterized by including:
[0022] An acquisition module, configured to acquire a left image and a right image, and calculate the disparity between the left image and the right image to obtain a first depth map;
[0023] The super-resolution module is used to fuse the left image and the first depth map, infer the fused features, decode the fused features through a feature decoder to obtain the super-resolved second depth map and the ground plane spatial coefficient; the resolution of the second depth map is higher than that of the first depth map;
[0024] The reconstruction module is used to convert the second depth map into three-dimensional point cloud data according to the internal parameters;
[0025] The fitting module is used to fit out the ground plane;
[0026] The screening module is used to obtain the error points below the ground plane according to the ground plane;
[0027] The correction module is used to find the two-dimensional points (x, y) of the corresponding image coordinates according to the error points, take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and find the depth z' to obtain the correct three-dimensional coordinates (X, Y, Z).
[0028] In a third aspect, the present invention provides a depth reconstruction, super-resolution and plane correction device for a binocular camera, which is characterized by including:
[0029] A processor;
[0030] A memory, in which executable instructions of the processor are stored;
[0031] Wherein, the processor is configured to execute the steps of the depth reconstruction, super-resolution and plane correction method of the binocular camera described in any one of the foregoing through executing the executable instructions.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, which is characterized in that when the program is executed, the steps of the depth reconstruction, super-resolution and plane correction method of the binocular camera described in any one of the foregoing are realized.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] Based on the binocular camera, the present invention uses the left and right images to obtain a low-resolution depth map, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the smoothing of the ground plane coefficient in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and realizing the binocular depth reconstruction task with a relatively low computational cost. Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings. By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more obvious:
[0036] Figure 1 It is a flowchart of the steps of a method for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention;
[0037] Figure 2 It is a flowchart of the steps of another method for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention;
[0038] Figure 3 It is a network framework diagram combining time series in an embodiment of the present invention;
[0039] Figure 4 It is a schematic structural diagram of a system for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention;
[0040] Figure 5 It is a schematic structural diagram of a device for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention; and
[0041] Figure 6 It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. Specific Embodiments
[0042] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0043] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0044] A method for depth reconstruction, super-resolution and plane correction of a binocular camera provided by an embodiment of the present invention aims to solve the problems existing in the prior art.
[0045] The following uses specific embodiments to elaborate in detail on the technical solutions of the present invention and how the technical solutions of the present application solve the above-mentioned technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present invention with reference to the drawings.
[0046] Based on a binocular camera, the present invention uses the left and right images to obtain a low-resolution depth map, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the ground plane coefficient smoothing in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction while achieving the binocular depth reconstruction task at a relatively low computational cost.
[0047] Figure 1 It is a flowchart of the steps of a method for depth reconstruction, super-resolution and plane correction of a binocular camera in an embodiment of the present invention. As Figure 1 shown, the steps of a method for depth reconstruction, super-resolution and plane correction of a binocular camera in an embodiment of the present invention include:
[0048] Step S1: Obtain the left image and the right image, and calculate the disparity between the left image and the right image to obtain a first depth map.
[0049] In this step, it is first necessary to synchronously collect scene information through the two lenses of the binocular camera to obtain the left image and the right image respectively. These two images are taken of the same object or scene from different perspectives and have a certain parallax relationship. They are the basic data for subsequent depth calculation. Parallax refers to the horizontal position offset of the same object in the two images. According to the principle of triangulation, given the baseline length of the binocular camera (the distance between the optical centers of the two lenses) and internal parameter information such as the focal length, the depth value corresponding to each pixel can be calculated using the calculated parallax, thereby constructing the first depth map. This depth map reflects the distance relationship of each object in the scene relative to the camera. However, there may be certain limitations in its resolution, etc., and subsequent steps will further optimize and process it.
[0050] In some embodiments, the SGM algorithm is used to obtain the first depth map. The SGM algorithm is an efficient method widely used in stereo matching to calculate parallax and then obtain the depth map. It is based on the idea of global energy optimization. While ensuring a certain matching accuracy, it can reduce the computational complexity through effective strategies and is applicable to binocular image parallax calculation in various scenarios. The processing includes:
[0051] Calculate the matching cost: For each pixel block in the image, calculate its matching cost at different parallaxes. In this process, the absolute difference (AD) can be used as the matching cost, that is, calculate the pixel value difference between the corresponding pixel blocks in the left image and the right image.
[0052] Cost aggregation: Associate the matching cost of each pixel block with the surrounding pixel blocks and aggregate to obtain a global cost function. The SGM algorithm realizes path aggregation through cumulative costs in multiple directions, which usually involves 8 directions of paths, including horizontal, vertical, and diagonal directions.
[0053] Parallax calculation: Calculate the parallax of each pixel block by minimizing the global cost function. The SGM algorithm uses the method of dynamic programming to optimize the parallax map and obtain the final matching result.
[0054] Generate the depth map: Generate the first depth map according to the calculated parallax values. The parallax values represent the horizontal displacement of pixels in the two perspectives and are proportional to the depth information in the scene. Each pixel value of the depth map represents the distance from the corresponding physical point to the camera.
[0055] The advantages of the SGM algorithm lie in its high accuracy and stability. It considers the global information in the image in a semi-global way, thereby improving the accuracy of depth estimation while maintaining the computational efficiency. Through the above steps, the SGM algorithm can effectively extract depth information from binocular images and provide basic data for subsequent depth map super-resolution and plane correction.
[0056] Step S2: Fuse the left image and the first depth map, infer the fused features, and decode the fused features through a feature decoder to obtain the super-resolved second depth map and the ground plane spatial coefficients; the resolution of the second depth map is higher than that of the first depth map.
[0057] In this step, the left image (containing rich appearance information such as texture and color) and the first depth map (containing depth information) are fused. The purpose is to integrate features from multiple aspects so that more comprehensive and valuable features can be mined for subsequent optimization of the depth map and other related calculations. The fusion method can be simple weighted fusion. For example, the pixel values of the left image and the depth values of the corresponding first depth map are combined according to certain weights; or more complex fusion methods based on convolutional neural networks can be used to let the network automatically learn the best fused feature representation. With the help of a deep learning model (such as a convolutional neural network), the fused features are input into it for inference. The network gradually extracts, abstracts, and transforms the features through multiple convolutional layers, pooling layers, etc., and mines deep-level fused features, which contain important information such as the correlation between the image appearance and depth. Use a feature decoder (usually also part of the network structure) to decode the inferred fused features. On the one hand, the decoder gradually restores a depth map with a higher resolution, that is, the second depth map, through operations such as upsampling. Compared with the first depth map, it has a more detailed depth information representation and can better reflect the depth of scene details; on the other hand, the ground plane spatial coefficients (a t , b t , c t , d t ) can also be obtained during the decoding process. These coefficients can be used for subsequent ground plane-related fitting, calibration, etc. operations, such as describing the attitude and position of the ground plane in three-dimensional space.
[0058] Step S3: Convert the second depth map into three-dimensional point cloud data according to the internal parameters.
[0059] In this step, the camera internal parameters include parameters such as focal length and principal point coordinates, which describe the imaging characteristics of the camera itself and the relationship between the image coordinates and the camera coordinate system. These internal parameter information has been determined during the camera calibration stage. For each pixel point in the second depth map, its pixel coordinates are known (for example, represented by (x, y)), and at the same time, there is a corresponding depth value (obtained from the depth map). According to the pinhole imaging model and the camera internal parameters, the pixel coordinates and depth values can be converted into three-dimensional coordinates (X, Y, Z) in the camera coordinate system through the corresponding geometric relationship calculation formula. Z = z, After such a conversion is performed on all the pixel points in the second depth map, three-dimensional point cloud data representing the three-dimensional structure of the scene is obtained. Each point in the point cloud data corresponds to an actual position in the scene, and can intuitively display the spatial distribution of the objects in the scene.
[0060] Step S4: Fit the ground plane.
[0061] In this step, based on the converted three-dimensional point cloud data, the least squares method or other optimization algorithms can be used to fit an optimal plane equation as the ground plane model. Commonly used methods for fitting the ground plane include the least squares method, the Random Sample Consensus algorithm (RANSAC), etc. For example, when using the least squares method, assuming the general form of the ground plane equation is aX + bY + cZ + d = 0 (in three-dimensional space), substitute the point coordinates in the three-dimensional point cloud data into this equation, and solve the coefficients a, b, c, d of the plane equation by constructing an error function and minimizing it, so as to determine the specific equation of the ground plane, that is, fit the ground plane. The RANSAC algorithm, on the other hand, fits the ground plane equation robustly through operations such as multiple random samplings, calculating the model, and evaluating inliers in the presence of noise and outliers (points that do not conform to the ground plane model), and it can better handle the interference of abnormal points that may exist in the point cloud data. No matter which method is used, the ultimate goal is to find the plane equation that best represents the ground plane based on the points that belong to or are close to the ground plane in the point cloud data, determine the position and attitude of the ground plane in three-dimensional space, and provide a basis for subsequent further calibration and other operations. In comparison, the least squares method has a smaller computational amount, is conducive to rapid calculation, and has a relatively low requirement for computing power.
[0062] Step S5: Obtain the incorrect points below the ground plane according to the ground plane.
[0063] In this step, by comparing the difference between the actually observed ground height and the theoretically predicted value, the points that are below the true ground are determined, that is, the depth estimates considered to be incorrect. Using the fitted ground plane equation, for each point in the 3D point cloud data, substitute its coordinates into the ground plane equation to calculate its distance to the ground plane (according to the distance formula from a point to a plane). If this distance is less than 0, it means that this point is below the plane (judging the up and down relationship according to the established conventions such as the direction of the plane normal vector), and these points can be determined to be incorrect points below the ground plane. Generally, in the actual scenario, there should theoretically be no points corresponding to valid scene objects below the ground plane (unless there are special cases such as underground passages that are considered separately during modeling), so these points below the ground plane are often abnormal points caused by reasons such as noise and depth calculation errors. Traverse the entire 3D point cloud data, filter out all the points below the ground plane, and form a set of incorrect points. Subsequently, correction processing will be carried out for these incorrect points.
[0064] Step S6: Find the two-dimensional coordinates (x, y) of the corresponding image coordinates according to the incorrect points, take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, find the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).
[0065] In this step, for the incorrect points marked as below the ground plane, recalculate their correct depth values based on the known internal and external camera parameters. Specifically, it is to use the previously mentioned plane equation to correct the Z-axis coordinates of these points so that they conform to the true position of the ground in physical reality.
[0066] When determining the image coordinates, given the coordinates of the incorrect points in the 3D point cloud data, through the inverse transformation of camera imaging and the camera internal parameters, the image coordinates (x, y) corresponding to these incorrect points in the original image (such as the left image, because the left image was involved in operations such as fusion before) can be calculated, establishing the association between the 3D space and the 2D image.
[0067] According to the camera imaging model and the geometric relationships related to the internal parameters, for the given image coordinates (x, y), an expression relating them to the depth value z can be derived. This expression reflects the constraint relationship between them during the camera imaging process, for example, derived through geometric principles such as similar triangles.
[0068] Substitute the derived expressions for x and y into the previously fitted ground plane equation. At this time, there is only one unknown parameter, the depth z, in the equation. By solving the equation, the correct depth value z' can be obtained, thus correcting the original incorrect depth.
[0069] Finally, by combining the image coordinates (x, y) and the corrected depth z’, the correct three-dimensional coordinates (X, Y, Z) are obtained, completing the three-dimensional coordinate correction of these incorrect points, making the entire three-dimensional point cloud data and the corresponding depth information more accurate and reliable, and more in line with the actual situation of the real scene.
[0070] In some embodiments, according to the formula the depth z’ is obtained, and then After finding the incorrect points, the two-bit coordinates (x, y) of the corresponding image coordinates are found according to the incorrect points. Taking the depth z as an unknown parameter, the expressions of X and Y are calculated according to the internal parameters. Substitute the expressions into the plane formula to inversely calculate the correct z, and then deduce the correct three-dimensional coordinates (X, Y, Z) according to z. The formula process is as follows:
[0071]
[0072] Z = z
[0073]
[0074]
[0075] Figure 2 This is the flowchart of the steps of another method for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention. As Figure 2 shown, compared with the foregoing embodiments, another method for depth reconstruction, super-resolution, and plane correction of a binocular camera in an embodiment of the present invention further includes:
[0076] Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through the internal parameters to obtain the final depth prediction result.
[0077] In this step, for the already corrected correct three-dimensional coordinates (X, Y, Z), use the projection matrix determined by the camera internal parameters (this matrix can be constructed according to the internal parameter values and reflects the conversion relationship from three-dimensional coordinates to two-dimensional image coordinates), and perform operations according to the corresponding projection calculation formula. Specifically, generally through linear algebraic operations such as matrix multiplication, multiply the three-dimensional coordinates (X, Y, Z) by the projection matrix to obtain the corresponding coordinate values on the two-dimensional plane. The two-dimensional plane coordinates obtained after projection, combined with the corresponding depth value z’, together constitute the final depth prediction result. This result not only contains the position information of the object on the two-dimensional image but also carries accurate depth information (reflected by the depth value z’), and can more comprehensively and accurately describe the spatial position and depth distribution of the object in the scene relative to the camera, and can be used in many subsequent related application scenarios such as target detection, scene understanding, and three-dimensional reconstruction visualization, providing a high-quality data basis for further analysis and processing.
[0078] The three points are reprojected back to the two-dimensional plane through the camera internal parameters to obtain the final depth map prediction result. The formula is as follows:
[0079] z = Z
[0080]
[0081] Figure 3 This is a network framework diagram combining time series in an embodiment of the present invention. As Figure 3 shown, this embodiment considers the association between the data of the previous frame and the current frame data. Figure 3 The model in
[0082] is illustrated by taking PostNet as an example. t represents the current frame, and t - 1 represents the previous frame. t , b t , c t , d t ) and the ground plane coefficients (a m-1 , b m-1 , c m-1 , d m-1 ) stored in the memory are weighted and added to be used as the ground plane coefficients (a m , b m , c m , d m ) of the current frame, and are stored in the memory for the next frame to use, as shown on the right side of Figure 3 .
[0083] In an actual application scenario, such as when the camera continuously captures images and performs depth reconstruction and other operations in a dynamic environment, the ground plane between adjacent frames often has a certain continuity and correlation. By weighted adding the ground plane spatial coefficients calculated from the current frame and the ground plane coefficients of the previous frame to determine the ground plane coefficients of the current frame, it helps to use historical information to stabilize the estimation of the ground plane, reduce the inaccurate ground plane fitting caused by factors such as noise and local anomalies that may exist in single-frame data, and make the determination of the ground plane smoother and more reliable in a continuous video sequence or dynamic scene.
[0084] Set appropriate weight values according to actual application requirements, experience, etc., which are the weights of the current frame coefficients and the previous frame coefficients respectively, and usually satisfy [conditions not specified in the original] to ensure the rationality of the result after weighted summation. For example, the weights can be adjusted according to factors such as the expected speed of scene change and the stability of camera movement. If the scene is relatively stable and the camera movement is gentle, [weight] can be appropriately increased to make the coefficients of the previous frame play a greater role; on the contrary, if the scene changes rapidly, [weight] can be appropriately increased to focus more on the coefficients calculated from the current frame. Calculate according to the rule of weighted summation to obtain the final ground plane coefficient of the current frame.
[0085] Through such a weighted summation method, during the continuous image frame processing, the correlation of ground plane information between different frames can be fully utilized to improve the robustness and accuracy of the entire depth reconstruction and plane correction methods.
[0086] In some embodiments, the weight of the ground plane spatial coefficient and the ground plane coefficient of the previous frame is proportional to the corresponding number of point clouds. When fusing the ground plane coefficients using multiple frames of data, the number of point clouds is a very crucial consideration factor. The number of point clouds to a certain extent reflects the reliability or the amount of information of the current frame and the previous frame data for ground plane estimation. Generally speaking, the more the number of point clouds, the richer the data for fitting the ground plane, and the more representative and accurate the calculated ground plane coefficients. Therefore, establishing a proportional relationship between the weight and the corresponding number of point clouds is to make the coefficients obtained based on a larger amount of point cloud data play a more important role when finally determining the ground plane coefficient of the current frame during the fusion of ground plane coefficients.
[0087] First, it is necessary to separately count the number of point clouds used to fit the ground plane in the current frame and the number of point clouds used to fit the ground plane in the previous frame. This counting process is usually completed by counting the point clouds that are screened out and may belong to the ground area before performing the ground plane fitting operation.
[0088] According to the principle that the weight is proportional to the number of point clouds, calculate their respective weights. When calculating the weights, first find the sum of the number of point clouds in the two frames, and then determine the weight of the ground plane spatial coefficient of the current frame and the weight of the ground plane coefficient of the previous frame in the following way.
[0089] Through such a weight setting and weighted calculation process, the difference in data information content reflected by the number of point clouds in different frames is fully considered, making the fusion of ground plane coefficients more reasonable and scientific, further improving the accuracy and stability of ground plane estimation during continuous multi-frame processing in a dynamic scene, and providing a more reliable basis for subsequent operations such as depth correction based on the ground plane.
[0090] Figure 4 This is a schematic structural diagram of a depth reconstruction, super-resolution, and plane correction system for a binocular camera in an embodiment of the present invention. AsFigure 4 As shown in Figure 4 , in an embodiment of the present invention, a depth reconstruction, super-resolution, and plane correction system for a binocular camera includes:
[0091] An acquisition module, configured to acquire a left image and a right image, and calculate the disparity between the left image and the right image to obtain a first depth map;
[0092] A super-resolution module, configured to fuse the left image and the first depth map, infer a fused feature, decode the fused feature through a feature decoder to obtain a super-resolved second depth map and a ground plane spatial coefficient; the resolution of the second depth map is higher than that of the first depth map;
[0093] A reconstruction module, configured to convert the second depth map into three-dimensional point cloud data according to the internal parameters;
[0094] A fitting module, configured to fit out a ground plane;
[0095] A screening module, configured to obtain error points below the ground plane according to the ground plane;
[0096] A correction module, configured to find the two-dimensional points (x, y) of the corresponding image coordinates according to the error points, take the depth z as an unknown parameter, calculate the expressions of X and Y according to the internal parameters, substitute the expressions into the plane formula, and obtain the depth z' to obtain the correct three-dimensional coordinates (X, Y, Z).
[0097] Based on the binocular camera, this embodiment uses the left and right images to obtain a low-resolution depth map, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the smoothing of the ground plane coefficient in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and achieving the binocular depth reconstruction task at a relatively low computational cost.
[0098] In an embodiment of the present invention, there is also provided a depth reconstruction, super-resolution, and plane correction device for a binocular camera, including a processor and a memory, in which executable instructions of the processor are stored. Among them, the processor is configured to execute the steps of a depth reconstruction, super-resolution, and plane correction method for a binocular camera by executing the executable instructions.
[0099] As above, based on the binocular camera, this embodiment uses the left and right images to obtain a low-resolution depth map, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the smoothing of the ground plane coefficient in the three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and achieving the binocular depth reconstruction task at a relatively low computational cost.
[0100] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.
[0101] Figure 5 It is a schematic structural diagram of a depth reconstruction, super-resolution, and plane correction device for a binocular camera in an embodiment of the present invention. The following will be described with reference to Figure 5 to describe the electronic device 600 according to this embodiment of the present invention. Figure 5 The displayed electronic device 600 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0102] As Figure 5 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0103] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the part of the method for depth reconstruction, super-resolution, and plane correction of a binocular camera in the above description of this specification. For example, the processing unit 610 can execute the steps as Figure 1 shown in.
[0104] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0105] The storage unit 620 may further include a program / utility 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a grid environment.
[0106] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0107] The electronic device 600 can also communicate with one or more external devices 700 (such as keyboards, pointing devices, Bluetooth devices, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as routers, modems, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as local area network (LAN), wide area network (WAN), and / or public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 5 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0108] In an embodiment of the present invention, a computer-readable storage medium is also provided for storing a program, and when the program is executed, it implements the steps of a depth reconstruction, super-resolution, and plane correction method for a binocular camera. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned part of the depth reconstruction, super-resolution, and plane correction method for a binocular camera in this specification.
[0109] As shown above, this embodiment is based on a binocular camera, obtains a low-resolution depth map using the left and right images, then uses a deep learning network to obtain a high-resolution depth map, and finally corrects the ground plane through the ground plane coefficient smoothing in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction, and at the same time achieving the binocular depth reconstruction task with a relatively low computational cost.
[0110] Figure 6 is a schematic structural diagram of the computer-readable storage medium in an embodiment of the present invention. Refer to Figure 6 shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described. It can use a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0111] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0113] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0114] This embodiment is based on a binocular camera. A low-resolution depth map is obtained using the left and right images, and then a high-resolution depth map is obtained using a deep learning network. Finally, the ground plane is corrected by smoothing the ground plane coefficients in three-dimensional space, ensuring the flatness and stability of the ground during depth reconstruction while achieving the binocular depth reconstruction task at a relatively low computational cost.
[0115] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0116] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention.
Claims
1. A method for depth reconstruction, super-resolution and plane correction of a binocular camera, characterized in that: include: Step S1: obtaining a left image and a right image, and calculating the disparity between the left image and the right image to obtain a first depth map; Step S2: Fusing the left image and the first depth map, inferring to obtain fused features, decoding the fused features through a feature decoder, and obtaining a second depth map and ground plane spatial coefficients after super-resolution; the resolution of the second depth map is higher than that of the first depth map; Step S3: converting the second depth map into three-dimensional point cloud data according to the internal reference; Step S4: fitting the ground plane; Step S5: obtaining error points below the ground plane according to the ground plane; Step S6: Find the corresponding two-point point (x, y) of the image coordinates according to the error point, take the depth z as the unknown parameter, calculate the expression of X and Y according to the internal parameter, substitute the expression into the plane formula, calculate the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).
2. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 1, characterized in that: Also includes: Step S7: Project the correct three-dimensional coordinates (X, Y, Z) onto a two-dimensional plane through internal references to obtain a final depth prediction result.
3. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 1, characterized in that: In step S1, the SGM algorithm is used to obtain a first depth map.
4. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 1, characterized in that: In step S4, the ground plane spatial coefficient is weightedly added to the ground plane coefficient of the previous frame to obtain the ground plane coefficient of the current frame.
5. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 4, characterized in that: The weight of the ground plane spatial coefficient and the ground plane coefficient of the previous frame is proportional to the corresponding number of point clouds.
6. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 1, characterized in that: In step S5, the three-dimensional point cloud data is substituted into the plane formula aX+bY+cZ+d=0 for calculation, and points below the ground plane are filtered out, namely, error points.
7. The method for depth reconstruction, super-resolution and plane correction of a binocular camera according to claim 1, characterized in that: In step S6, according to the formula Obtain the depth z', and then obtain 8. A binocular camera depth reconstruction, super-resolution and plane correction system, used to implement the binocular camera depth reconstruction, super-resolution and plane correction method according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used to acquire a left image and a right image, and calculate the disparity between the left image and the right image to obtain a first depth map; A super-resolution module is used to fuse the left image and the first depth map, infer fusion features, decode the fusion features through a feature decoder, and obtain a super-resolved second depth map and ground plane spatial coefficients; the resolution of the second depth map is higher than that of the first depth map; A reconstruction module, used for converting the second depth map into three-dimensional point cloud data according to an internal parameter; A fitting module, used to fit the ground plane; A screening module, used for obtaining error points below the ground plane according to the ground plane; The correction module is used to find the corresponding two-point point (x, y) of the image coordinates according to the error point, take the depth z as an unknown parameter, calculate the expression of X and Y according to the internal parameter, substitute the expression into the plane formula, calculate the depth z', and obtain the correct three-dimensional coordinates (X, Y, Z).
9. A binocular camera depth reconstruction, super-resolution and plane correction device, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to execute the steps of the depth reconstruction, super-resolution and plane correction method of the binocular camera of any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the depth reconstruction, super-resolution and plane correction method of the binocular camera described in any one of claims 1 to 7 are implemented.