Multi-scale information fusion depth reconstruction method, system and device based on binocular camera and storage medium

Images of different resolutions are acquired through binocular cameras, depth maps are generated and synthesized, and high-precision depth maps are generated using deep learning networks, which solves the problem of poor depth reconstruction effect in various scenarios in the existing technology, and realizes real-time and accurate depth reconstruction.

CN120147126APending Publication Date: 2025-06-13SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213458.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time depth reconstruction in a variety of scenarios, while taking into account close-range depth detection and detail accuracy, especially in ultra-close (within 12cm) scenarios, the depth reconstruction effect is not good.

Method used

The left and right images of different resolutions are obtained by a binocular camera, depth images of different resolutions are generated, and the initial depth images are synthesized through an effective depth synthesis scheme from near to far, and a deep learning network is used to generate high-precision and high-resolution dense depth images.

Benefits of technology

Real-time deep reconstruction in a variety of scenarios is achieved, breaking through the lower limit of detection distance of traditional binocular solutions, taking into account close-range depth detection and detail accuracy, and significantly improving the accuracy and fidelity of reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147126A_ABST
    Figure CN120147126A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale information fusion depth reconstruction method, system and device based on a binocular camera, and a storage medium. The method comprises the following steps: S1, obtaining a left image and a right image through the binocular camera; performing down-sampling on the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image and a second-resolution right image; s2, obtaining a first resolution depth image according to the first resolution left image and the first resolution right image; obtaining a second resolution depth image according to the second resolution left image and the second resolution right image; s3, synthesizing the first resolution depth map and the second resolution depth map to obtain an initial depth map which is high in precision and takes different depth distances into consideration; s4, fusing the first resolution left image and the initial depth image by using a feature encoder to obtain fused features; and S5, decoding the fused features through a feature decoder to obtain a super-divided high-resolution depth map.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In recent years, the demand for image-based depth reconstruction has been increasing. It is a great technical challenge to achieve the depth reconstruction of various objects in the target scene, ensure the effective detection of ultra-close obstacles, and at the same time ensure its accuracy. With a fixed maximum disparity, different resolutions of image input will result in different depth details and different effective detection distances: the larger the resolution, the better the depth details, but the lower limit of depth detection will become larger; the smaller the resolution, the closer the effective distance that can be detected, but there will be a degradation of depth details. Currently, there is no dedicated binocular three-dimensional depth reconstruction algorithm that can balance these two aspects. Therefore, it is of great significance to effectively fuse multi-scale binocular information for depth reconstruction to achieve real-time depth reconstruction in diverse scenarios while taking into account close-range depth detection and detail accuracy.

[0003] Currently, typical binocular-based depth prediction models can only perform detection under the same resolution input and cannot well balance the depth reconstruction effects at various distances, especially for ultra-close distances (within 12 cm). For some common close-range scene depth reconstructions, they are basically ineffective. For example, CRE, although it has good results in the details of depth prediction, has no ability to effectively detect obstacles with a distance less than 12 cm, and serious depth prediction error problems will occur.

[0004] The disclosure of the above background art content is only used to assist in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0005] To this end, based on a binocular camera, the present invention obtains depth maps with different resolutions from the left and right images obtained by downsampling. The depth maps with different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, the depths of different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained using a deep learning network, which not only realizes the binocular depth reconstruction task at a relatively low computational cost but also ensures the detection accuracy of depths at different distances during depth reconstruction, breaking through the lower limit of the detection distance of the traditional binocular scheme.

[0006] In a first aspect, the present invention provides a multi-scale information fusion depth reconstruction method based on a binocular camera, which is characterized by including:

[0007] Step S1: Obtain a left image and a right image through a binocular camera; downsample the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image;

[0008] Step S2: Obtain a first-resolution depth map based on the left image of the first resolution and the right image of the first resolution; obtain a second-resolution depth map based on the left image of the second resolution and the right image of the second resolution.

[0009] Step S3: Start synthesizing from the first-resolution depth map with the second-resolution depth map to obtain an initial depth map with high accuracy and taking into account different depth distances.

[0010] Step S4: Use a feature encoder to fuse the left image of the first resolution and the initial depth map to obtain fused features.

[0011] Step S5: Decode the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.

[0012] Optionally, in the method for multi-scale information fusion depth reconstruction based on a binocular camera, in step S1, the downsampling is implemented by the Gaussian pyramid method, where the resolution of the left image of the first resolution and the right image of the first resolution is 1 / 2 of the left image of the second resolution and the right image of the second resolution.

[0013] Optionally, in the method for multi-scale information fusion depth reconstruction based on a binocular camera, in step S2, the first-resolution depth map and the second-resolution depth map are calculated through a stereo matching algorithm, and the stereo matching algorithm includes but is not limited to semi-global matching (SGM), block matching, or adaptive window matching.

[0014] Optionally, in the method for multi-scale information fusion depth reconstruction based on a binocular camera, in step S3, the iterative synthesis includes the following steps:

[0015] Step S31: Obtain a first detection distance lower limit of the first-resolution depth map and a second detection distance lower limit of the second-resolution depth map.

[0016] Step S32: For pixel points with depth values greater than the first detection distance lower limit, perform weighted fusion on the corresponding pixel points of the first-resolution depth map and the second-resolution depth map; for pixel points with depth values less than the first detection distance lower limit and greater than the second detection distance lower limit, use the depth value of the second-resolution depth value to obtain an initial depth map.

[0017] Optionally, in the method for multi-scale information fusion depth reconstruction based on a binocular camera, in step S4, the feature encoder is a convolutional neural network, and the number of layers and the number of convolutional kernels in each layer of the convolutional neural network are optimized and configured according to the size and feature complexity of the input image.

[0018] Optionally, in the multi-scale information fusion depth reconstruction method based on a binocular camera, the convolutional neural network structure includes at least one of ResNet, VGG, or U-Net architectures.

[0019] Optionally, in the multi-scale information fusion depth reconstruction method based on a binocular camera, in step S5, the feature decoder includes a deconvolution layer and an upsampling layer for gradually restoring the fused features to a high-resolution depth map.

[0020] In a second aspect, the present invention provides a multi-scale information fusion depth reconstruction system based on a binocular camera for implementing the multi-scale information fusion depth reconstruction method based on a binocular camera described in any one of the foregoing items, which is characterized by including:

[0021] An acquisition module for obtaining a left image and a right image through a binocular camera; downsampling the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image;

[0022] A depth module for obtaining a first-resolution depth map based on the first-resolution left image and the first-resolution right image; obtaining a second-resolution depth map based on the second-resolution left image and the second-resolution right image;

[0023] A synthesis module for synthesizing the first-resolution depth map with the second-resolution depth map to obtain an initial depth map with high accuracy and taking into account different depth distances;

[0024] A fusion module for fusing the first-resolution left image and the initial depth map using a feature encoder to obtain fused features;

[0025] A decoding module for decoding the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.

[0026] In a third aspect, the present invention provides a multi-scale information fusion depth reconstruction device based on a binocular camera, which is characterized by including:

[0027] A processor;

[0028] A memory storing executable instructions of the processor;

[0029] Wherein, the processor is configured to execute the steps of the multi-scale information fusion depth reconstruction method based on a binocular camera described in any one of the foregoing items by executing the executable instructions.

[0030] Fourthly, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of the multi-scale information fusion depth reconstruction method based on a binocular camera described in any one of the foregoing are implemented.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] By generating depth maps of different resolutions and synthesizing them, the present invention can effectively balance the accuracy of different depth distances. The low-resolution depth map can obtain depth information at closer distances, while the high-resolution depth map can capture more accurate depth values. The combination of the two makes the reconstruction result more accurate and comprehensive.

[0033] The present invention uses a feature encoder and a decoder to fuse and super-resolution reconstruct the depth map, and can generate a high-resolution depth map. Compared with traditional methods, the present invention can significantly improve the resolution of the depth map while maintaining the computational efficiency. By using the feature encoder to extract and fuse the features of the image and the depth map, the present invention can make full use of the semantic information and spatial information of the image, and further improve the detail performance of the depth map.

[0034] Through multi-scale information fusion and feature fusion, the present invention can better retain the high-frequency details in the depth map. This detail retention is particularly important for the reconstruction of complex scenes and can significantly improve the fidelity of the reconstruction result. Through multi-scale information fusion, the present invention can effectively process the depth information of distant and close objects and is applicable to various scenarios such as indoor and outdoor. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more obvious:

[0036] Figure 1 It is a flowchart of the steps of a multi-scale information fusion depth reconstruction method based on a binocular camera in an embodiment of the present invention;

[0037] Figure 2 It is a flowchart of the steps of an iterative synthesis in an embodiment of the present invention;

[0038] Figure 3 It is a schematic diagram of a binocular multi-scale depth prediction network framework in an embodiment of the present invention;

[0039] Figure 4 Schematic diagram of the structure of a multi-scale information fusion depth reconstruction system based on a binocular camera in an embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the structure of a multi-scale information fusion depth reconstruction device based on a binocular camera in an embodiment of the present invention; and

[0041] Figure 6 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Specific embodiments

[0042] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0043] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] A multi-scale information fusion depth reconstruction method based on a binocular camera provided by an embodiment of the present invention aims to solve the problems existing in the prior art.

[0045] The technical solutions of the present invention and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.

[0046] Due to the binocular depth calculation algorithm, the binocular camera has a detection distance lower limit related to its resolution. For target objects with a depth less than the detection distance lower limit, their depth values cannot be calculated, even if the target objects appear in the binocular images. The applicant found that the detection distance lower limit is related to the resolution of the binocular images. When the resolution of the binocular images decreases, the detection distance lower limit also decreases, so that depth values at closer distances can be obtained.

[0047] Based on the binocular camera, the present invention uses the left and right images with different resolutions obtained by downsampling to obtain depth maps with different resolutions. The depth maps with different resolutions have different detection distance lower limits. Through an effective depth synthesis scheme from near to far, the depths at different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained by using a deep learning network. While achieving the binocular depth reconstruction task at a relatively low computational cost, the detection accuracy of depths at different distances during depth reconstruction is ensured, and the detection distance lower limit of the traditional binocular scheme is broken through.

[0048] Figure 1 It is a flowchart of the steps of a multi-scale information fusion depth reconstruction method based on a binocular camera in an embodiment of the present invention. As Figure 1 shown, the steps of a multi-scale information fusion depth reconstruction method based on a binocular camera in an embodiment of the present invention include:

[0049] Step S1: Obtain a left image and a right image through the binocular camera; perform downsampling on the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image.

[0050] In this step, the binocular camera simulates the human eye vision principle, and takes pictures of the same object from different angles through two cameras, so as to obtain an image pair with parallax information. Downsampling is an operation to reduce the image resolution. By reducing the number of pixels in the image, the resolution is reduced. The original image data for depth reconstruction is obtained, and images with different resolutions are generated through downsampling, providing a basis for subsequent multi-scale information fusion.

[0051] The binocular camera is used to simultaneously capture the scene to obtain a left image and a right image. The Gaussian pyramid downsampling method is adopted, and by adjusting the size and step of the Gaussian kernel, the left image and the right image are respectively downsampled to obtain a first-resolution left image and a first-resolution right image with a lower resolution, and a second-resolution left image and a second-resolution right image with a relatively higher resolution.

[0052] Step S2: Obtain a first-resolution depth map according to the first-resolution left image and the first-resolution right image; obtain a second-resolution depth map according to the second-resolution left image and the second-resolution right image; wherein, the first resolution is greater than the second resolution.

[0053] In this step, based on the parallax principle, the depth information is calculated using the position differences of corresponding pixel points in the left and right images. The semi-global matching algorithm is a commonly used stereo matching algorithm. By performing energy optimization within a global range, the best matching points in the left and right images are found, thereby calculating the depth map.

[0054] The left image and the right image at the first resolution are input into the calculation module based on the semi-global matching algorithm. After steps such as matching point search and parallax calculation, the depth map at the first resolution is obtained. Similarly, the left image and the right image at the second resolution are processed in the same way to obtain the depth map at the second resolution.

[0055] Step S3: Starting from the depth map at the first resolution, it is synthesized with the depth map at the second resolution to obtain an initial depth map with high accuracy and taking into account different depth distances.

[0056] In this step, the low-resolution depth map has a smaller lower limit of detection distance but relatively lower accuracy. The high-resolution depth map is richer in details but has a larger lower limit of detection distance. By means of iterative synthesis, the advantages of both are combined to gradually improve the accuracy of the depth map and its adaptability to different depth distances. This step generates an initial depth map with high accuracy and capable of taking into account different depth distances, providing better-quality data for subsequent feature fusion and super-resolution processing.

[0057] Starting from the depth map at the first resolution, it is synthesized with the depth map at the second resolution. In each iteration, according to the error situation of the current synthesized depth map, the fusion weights of the depth map at the first resolution and the depth map at the second resolution are dynamically adjusted. For example, within the lower limit of the detection distance of the high-resolution depth map, there are no valid depth values in the high-resolution depth map, and the weight of the depth value of the high-resolution depth map is set to 0, while the weight of the low-resolution depth map is set to 1; outside the lower limit of the detection distance of the high-resolution depth map, the high-resolution depth map has depth values with high accuracy and high confidence, and the weight of the depth value of the high-resolution depth map is set to 1, while the weight of the low-resolution depth map is set to 0.

[0058] Step S4: Use the feature encoder to fuse the left image at the first resolution and the initial depth map to obtain fused features.

[0059] In this step, the feature encoder usually adopts a convolutional neural network (CNN). Through structures such as convolutional layers and pooling layers, the CNN can automatically extract the features of the image. The left image at the first resolution and the initial depth map are input into the feature encoder, and using the powerful feature extraction ability of the CNN, the feature information of both is fused to obtain fused features containing image and depth information.

[0060] Construct a convolutional neural network as a feature encoder, and optimize the configuration of the number of layers and the number of convolutional kernels in each layer according to the size and feature complexity of the input image. Input the left image at the first resolution and the initial depth map into the feature encoder at the same time. After a series of operations such as convolution and pooling, fused features are output.

[0061] Step S5: Decode the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.

[0062] In this step, the feature decoder adopts a combination of transposed convolution operations and skip connections. The transposed convolution operation is the inverse process of the convolution operation, which can restore a low-resolution feature map to a high-resolution image. The skip connection is used to fuse feature information at different levels to avoid losing important information during the transposed convolution process, thereby realizing the reconstruction of the super-resolved high-resolution depth map.

[0063] Input the fused features into the feature decoder. The feature decoder first gradually enlarges the size of the feature map through transposed convolution operations, and at the same time introduces the feature information at different levels before by using skip connections. After multiple transposed convolutions and feature fusions, a super-resolved high-resolution depth map is finally output.

[0064] This embodiment obtains the initial depths at different limit distances based on multi-scale binocular images, and synthesizes them according to the lower limit of the effective detection distance of depth maps at different resolutions, breaking through the lower limit of the detection distance of the original resolution to obtain the initial depth at a super-close distance, while maintaining the detail accuracy of other depth distances. Finally, the low-resolution depth map is repaired and super-resolved through a deep learning network in combination with the RGB image to obtain a depth reconstruction effect that can take into account both super-close distances and maintain the detail accuracy at a distance.

[0065] In some embodiments, in step S1, the downsampling is implemented by the Gaussian pyramid method, where the resolutions of the left image and the right image at the first resolution are 1 / 2 of the resolutions of the left image and the right image at the second resolution. The Gaussian pyramid method is a common image downsampling technique. By performing Gaussian filtering and downsampling operations on the image, it can reduce the image resolution while trying to retain the main features of the image. The resolutions of the left image and the right image at the first resolution are one-half of the resolutions of the left image and the right image at the second resolution. This specific resolution ratio relationship provides a clear scale basis for subsequent operations such as depth map calculation and fusion based on images at different resolutions, and can make the change in resolution slow, so that the depth values change successively, and relatively accurate depth values can be obtained.

[0066] In some embodiments, in step S2, the first-resolution depth map and the second-resolution depth map are calculated through a stereo matching algorithm, which includes but is not limited to semi-global matching (SGM), block matching, or adaptive window matching. The stereo matching algorithm is the core method for calculating the depth map. Its principle is based on the parallax principle, using the position difference of corresponding pixel points in the left and right images to calculate the depth information. The stereo matching algorithms mentioned here include various ones, such as semi-global matching (SGM), block matching, or adaptive window matching, etc. Semi-global matching (SGM) is a commonly used and effective stereo matching algorithm. It calculates the depth map by performing energy optimization within a global range to find the best matching points. The block matching algorithm divides the image into multiple small blocks and determines the parallax by searching for the best match of the corresponding small blocks in the left and right images. The adaptive window matching algorithm adaptively adjusts the size of the matching window according to the local features of the image to improve the matching accuracy. These different algorithms have their own characteristics and applicable scenarios. When calculating the first-resolution depth map and the second-resolution depth map, an appropriate stereo matching algorithm can be selected according to specific requirements and image characteristics, providing an accurate data basis for the subsequent fusion of the depth maps and the entire depth reconstruction process.

[0067] Figure 2 This is a flowchart of the steps of an iterative synthesis in an embodiment of the present invention. As Figure 2 shown, the steps of an iterative synthesis in an embodiment of the present invention include the following steps:

[0068] Step S31: Obtain the lower limit of the first detection distance of the first-resolution depth map and obtain the lower limit of the second detection distance of the second-resolution depth map.

[0069] In this step, the lower limit of the first detection distance is greater than the lower limit of the second detection distance. For a binocular camera, when its internal parameters and external parameters are fixed, the lower limit of the detection distance corresponding to the depth map of the corresponding resolution is fixed. That is, the lower limit of the first detection distance of the first-resolution depth map is known, and the lower limit of the second detection distance of the second-resolution depth map is also known.

[0070] Step S32: For the pixel points with depth values greater than the lower limit of the first detection distance, perform weighted fusion on the corresponding pixel points of the first-resolution depth map and the second-resolution depth map; for the pixel points with depth values less than the lower limit of the first detection distance and greater than the lower limit of the second detection distance, use the depth value of the second-resolution depth value to obtain an initial depth map.

[0071] In this step, for a depth camera, the depth range can be divided into three intervals according to the lower limit of the first detection distance and the lower limit of the second detection distance: greater than the lower limit of the first detection distance, less than the lower limit of the first detection distance and greater than the lower limit of the second detection distance, and less than the lower limit of the second detection distance.

[0072] For pixel points greater than the lower limit of the first detection distance, weighted fusion is performed based on the depth values corresponding to the first-resolution depth map and the second-resolution depth map. The fusion method can be freely selected. For example, different weights are assigned according to the change of the depth value. The greater the depth value, the greater the weight of the first-resolution depth map; the smaller the depth value, the smaller the weight of the first-resolution depth map. For another example, the weight of the first-resolution depth map is set to 1, and the weight of the second-resolution depth map is set to 0, that is, all the depth data of the first-resolution depth map is adopted.

[0073] For pixel points less than the lower limit of the first detection distance and greater than the lower limit of the second detection distance, the depth value corresponding to the second-resolution depth map is adopted. At this time, the depth value in the first-resolution depth map is empty or an incorrect depth value. Directly adopting the depth value corresponding to the second-resolution depth map can obtain an accurate depth value.

[0074] For pixel points less than the lower limit of the second detection distance, there is no reliable depth value.

[0075] According to the description of this embodiment, those skilled in the art can easily think of adopting more levels of downsampling to obtain more different resolution depth maps, so as to obtain the lower limit of the third detection distance, the lower limit of the fourth detection distance, etc., so as to continuously approach the detection limit of the depth camera and obtain depth values at closer distances. These are all obtained under the inspiration of the present invention and do not deviate from the core inventive points of the present invention, and also belong to the protection scope of the present invention.

[0076] In some embodiments, in step S4, the feature encoder is a convolutional neural network, and the number of layers and the number of convolutional kernels in each layer of the convolutional neural network are optimized and configured according to the size and feature complexity of the input image. This embodiment takes into account the differences in the size and feature complexity of different input images. For example, when the input image is large in size and contains rich details and complex textures, in order to be able to fully capture these feature information, it may be necessary to increase the number of layers of the convolutional neural network so that the network can extract features from the image more deeply; at the same time, the number of convolutional kernels in each layer may also increase accordingly to dig out the features of the image from different angles and scales. On the contrary, if the input image is small in size and the features are relatively simple, too many layers and convolutional kernels may lead to problems such as waste of computing resources and overfitting. At this time, the number of layers and convolutional kernels can be appropriately reduced.

[0077] This method of optimizing and configuring according to the characteristics of the input image can enable the convolutional neural network to better adapt to different input data, improve the efficiency and quality of feature extraction, and thus more effectively fuse the feature information of the first-resolution left image and the initial depth map, laying a solid foundation for generating a high-quality super-resolved high-resolution depth map in the subsequent steps.

[0078] In some embodiments, the convolutional neural network structure includes at least one of ResNet, VGG, or U-Net architectures. ResNet solves the problems of vanishing gradients and network degradation in deep neural networks by introducing residual connections, enabling the network to be constructed deeper to extract more advanced features; the VGG network has a simple and unified convolutional layer structure, increasing the depth and non-linear expression ability of the network by stacking multiple small-sized convolutional kernels; the U-Net architecture is a symmetric encoder-decoder structure that directly transfers the feature information of the encoder part to the decoder part through skip connections, helping to restore the detailed information of the image. These different architectures all have their unique advantages and are suitable for different feature extraction and fusion tasks.

[0079] In some embodiments, in step S5, the feature decoder includes a transposed convolutional layer and an upsampling layer for gradually restoring the fused features to a high-resolution depth map. The transposed convolutional layer, also known as the deconvolutional layer, is not truly the inverse operation of the convolutional operation in essence, but its function is to convert a low-resolution feature map into a high-resolution feature map. In depth map reconstruction, the transposed convolutional layer performs specific operation on the input fused features, inserts new pixel values between the pixels of the feature map, thereby increasing the size of the feature map and gradually restoring the resolution of the image. For example, after feature extraction and fusion by the convolutional neural network, the resulting fused feature map has a low resolution, and the transposed convolutional layer can process it to make the size of the feature map gradually approach the size of the original high-resolution image.

[0080] The upsampling layer also plays an important role in enhancing the resolution of the feature map. The upsampling layer inserts new pixel values between the pixels of the low-resolution feature map through methods such as nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation to increase the resolution. Different from the transposed convolutional layer, the upsampling layer mainly focuses on increasing the number of pixels through interpolation algorithms, while the transposed convolutional layer not only increases the size but also has a certain feature transformation ability.

[0081] The transposed convolutional layer and the upsampling layer in the feature decoder cooperate with each other to gradually restore the fused features to a high-resolution depth map. They start from the low-resolution fused features, through multiple transposed convolutional and upsampling operations, continuously increase the resolution of the feature map, while retaining and restoring the detailed information of the image, and finally generate a super-resolved high-resolution depth map to meet the requirements of depth reconstruction for high-precision depth information, providing accurate depth data support for subsequent applications such as obstacle detection in autonomous driving and scene construction in virtual reality.

[0082] Figure 3 Schematic diagram of a binocular multi-scale depth prediction network framework in an embodiment of the present invention.Figure 3 Three - level downsampling is adopted to obtain the left and right images at scales of 1 / 2, 1 / 4, and 1 / 8 respectively. The following is combined with Figure 3 to elaborate in detail.

[0083] The left and right images obtained by the binocular camera are downsampled to obtain the left and right images at scales of 1 / 2, 1 / 4, and 1 / 8;

[0084] The SGM algorithm with the same 128 - disparity is used to process the left and right images at resolutions of 1 / 2, 1 / 4, and 1 / 8 respectively, and depth maps at each resolution are obtained quickly;

[0085] The lower limit of the depth limit at each resolution is calculated according to the minimum detection depth formula. The minimum detection depth calculation formula is as follows:

[0086] (focallength / scale*baseline) / Maxdisparity

[0087] Where focallength is the camera focal length, scale is the downsampling scale, baseline is the baseline length of the binocular camera, and Maxdisparity is the maximum disparity of binocular matching. Example: When the original image resolution is 1280*800, the focal length is 611, the maximum disparity is 128, and the baseline distance is 100: The lower limit of the limit depth at 1 / 2 scale is about 24 cm, the lower limit of the limit depth at 1 / 4 scale is about 12 cm, and the lower limit of the limit depth at 1 / 8 scale is about 6 cm.

[0088] The 1 / 8 - resolution depth map is upsampled by 2 times, the calculated lower limit of the limit depth is taken, and pixels with depths in the range of [6 cm, 12 cm] in the 1 / 8 - resolution depth map are used to generate a mask range. According to this mask, the depths in the 1 / 4 - resolution depth map are replaced with the depths in the 1 / 8 - resolution depth map to obtain a synthesized 1 / 4 - resolution depth map.

[0089] The synthesized 1 / 4 - resolution depth map is further upsampled by 2 times. Pixels with depths in the range of [6 cm, 24 cm] in the synthesized 1 / 4 - resolution depth map are used to generate a mask range. According to this mask, the depths in the 1 / 2 - resolution depth map are replaced with the depths in the synthesized 1 / 4 - resolution depth map to obtain a synthesized 1 / 2 - resolution depth map. The lower limit of the effective limit detection distance of this synthesized 1 / 2 - resolution depth map reaches 6 cm, and at the same time, the accuracy of other long - distance depths is retained.

[0090] Using Figure 3 the feature encoder shown in the figure to fuse the left image obtained by the binocular camera and the synthesized depth map, and infer the fused features of the left image and the initial depth;

[0091] The fused features are decoded by the feature decoder to obtain a corrected and super-resolved high-precision and high-resolution depth map.

[0092] Figure 4 This is a schematic structural diagram of a multi-scale information fusion depth reconstruction system based on a binocular camera in an embodiment of the present invention. As Figure 4 shown, a multi-scale information fusion depth reconstruction system based on a binocular camera in an embodiment of the present invention includes:

[0093] An acquisition module, configured to obtain a left image and a right image through a binocular camera; perform downsampling on the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image;

[0094] A depth module, configured to obtain a first-resolution depth map according to the first-resolution left image and the first-resolution right image; obtain a second-resolution depth map according to the second-resolution left image and the second-resolution right image;

[0095] A synthesis module, configured to start synthesizing from the first-resolution depth map and the second-resolution depth map to obtain an initial depth map with high accuracy and taking into account different depth distances;

[0096] A fusion module, configured to fuse the first-resolution left image and the initial depth map by using a feature encoder to obtain fused features;

[0097] A decoding module, configured to decode the fused features by a feature decoder to obtain a super-resolved high-resolution depth map.

[0098] Specifically, the acquisition module obtains a pair of image data, namely a left image and a right image, through a binocular camera. Then, downsampling operations are respectively performed on the obtained left image and right image to obtain images with different resolutions, namely a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image. The purpose of the downsampling operation is to provide image information at different scales for subsequent processing, so as to perform depth calculation and information fusion at different resolutions. It is the starting module of the entire system and provides input data, namely left and right images with different resolutions, for the depth module.

[0099] The depth module receives the image pairs with different resolutions output by the acquisition module, and obtains a first-resolution depth map according to the first-resolution left image and the first-resolution right image through corresponding algorithms or calculation methods; similarly, a second-resolution depth map is obtained according to the second-resolution left image and the second-resolution right image. The main function of this module is to calculate the corresponding depth information at different resolutions, providing a basis for subsequent depth map synthesis. It inputs different-resolution images from the acquisition module and outputs depth maps with different resolutions to the synthesis module.

[0100] The synthesis module receives the first-resolution depth map and the second-resolution depth map output by the depth module, and starts the synthesis operation with the second-resolution depth map from the first-resolution depth map. Through synthesis, an initial depth map with high accuracy and taking into account different depth distances is obtained. This synthesis operation can integrate the advantages of depth maps with different resolutions and improve the quality and accuracy of the depth map. It inputs depth maps with different resolutions from the depth module and outputs the synthesized initial depth map to the fusion module.

[0101] The fusion module uses the feature encoder to perform fusion processing on the first-resolution left image output by the acquisition module and the initial depth map output by the synthesis module to obtain fusion features. The fusion process combines the visual features and depth information of the image so that these comprehensive information can be better utilized in subsequent processing to generate a high-resolution depth map. It inputs the first-resolution left image from the acquisition module and the initial depth map of the synthesis module and outputs the fusion features to the decoding module.

[0102] The decoding module receives the fusion features output by the fusion module and performs a decoding operation on the fusion features through the feature decoder, and finally obtains the super-resolved high-resolution depth map. The decoding process restores the fusion features to the form of a high-resolution depth map, completing the ultimate goal of the entire depth reconstruction system. It inputs the fusion features from the fusion module and outputs the super-resolved high-resolution depth map, which is the output module of the entire system.

[0103] This embodiment is based on a binocular camera. Different-resolution depth maps are obtained by using the left and right images with different resolutions obtained by downsampling. The depth maps with different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, the depths with different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained by using a deep learning network. While achieving the binocular depth reconstruction task at a lower computational cost, it ensures the detection accuracy of depths at different distances during depth reconstruction and breaks through the lower limit of the detection distance of the traditional binocular scheme.

[0104] In the embodiment of the present invention, a multi-scale information fusion depth reconstruction device based on a binocular camera is further provided, including a processor and a memory in which executable instructions of the processor are stored. Among them, the processor is configured to execute the steps of a multi-scale information fusion depth reconstruction method based on a binocular camera by executing the executable instructions.

[0105] As described above, in this embodiment, based on a binocular camera, left and right images with different resolutions obtained by downsampling are used to obtain depth maps with different resolutions. The depth maps with different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, depths with different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained using a deep learning network. While achieving the binocular depth reconstruction task at a relatively low computational cost, the detection accuracy of depths at different distances during depth reconstruction is ensured, breaking through the lower limit of the detection distance of the traditional binocular scheme.

[0106] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "platform" here.

[0107] Figure 5 It is a schematic structural diagram of a multi-scale information fusion depth reconstruction device based on a binocular camera in an embodiment of the present invention. The following refers to Figure 5 to describe the electronic device 600 according to this embodiment of the present invention. Figure 5 The electronic device 600 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0108] As Figure 5 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, and the like.

[0109] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the part of the method for multi-scale information fusion depth reconstruction based on a binocular camera in the above description of this specification. For example, the processing unit 610 can execute the steps as Figure 1 shown.

[0110] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0111] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a grid environment.

[0112] The bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0113] The electronic device 600 may also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 650. Also, the electronic device 600 may communicate with one or more grids (such as a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through the grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 5 not shown, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0114] An embodiment of the present invention also provides a computer-readable storage medium for storing a program, and the steps of a multi-scale information fusion depth reconstruction method based on a binocular camera are implemented when the program is executed. In some possible implementation manners, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above part of the multi-scale information fusion depth reconstruction method based on a binocular camera of this specification.

[0115] As shown above, this embodiment is based on a binocular camera. Different-resolution left and right images obtained by downsampling are used to obtain depth maps with different resolutions. The depth maps with different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, depths with different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained using a deep learning network. While achieving the binocular depth reconstruction task at a relatively low computational cost, the detection accuracy of depths at different distances during depth reconstruction is ensured, breaking through the lower limit of the detection distance of the traditional binocular scheme.

[0116] Figure 6 It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. Refer to Figure 6 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0117] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0118] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0119] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0120] This embodiment is based on a binocular camera. Different-resolution left and right images obtained by downsampling are used to obtain depth maps of different resolutions. The depth maps of different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, the depths of different limit distances are synthesized. Then, a high-precision and high-resolution dense depth map is obtained by using a deep learning network. While achieving the binocular depth reconstruction task at a relatively low computational cost, the detection accuracy of depths at different distances during depth reconstruction is ensured, breaking through the lower limit of the detection distance of the traditional binocular scheme.

[0121] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0122] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments. Those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A multi-scale information fusion depth reconstruction method based on a binocular camera, characterized in that: include: Step S1: Obtain a left image and a right image through a binocular camera; Down-sampling the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image; Step S2: obtaining a first-resolution depth map according to the first-resolution left image and the first-resolution right image; Obtain a second-resolution depth map according to the second-resolution left image and the second-resolution right image; Step S3: synthesizing the first resolution depth map with the second resolution depth map to obtain an initial depth map with high accuracy and taking into account different depth distances; Step S4: using a feature encoder to fuse the first resolution left image and the initial depth image to obtain a fused feature; Step S5: Decode the fused features through a feature decoder to obtain a high-resolution depth map after super-resolution.

2. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 1 is characterized in that: In step S1, the downsampling is implemented by a Gaussian pyramid method, wherein the resolutions of the first-resolution left image and the first-resolution right image are 1 / 2 of the second-resolution left image and the second-resolution right image.

3. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 1 is characterized in that: In step S2, the first resolution depth map and the second resolution depth map are calculated by a stereo matching algorithm, and the stereo matching algorithm includes but is not limited to semi-global matching (SGM), block matching or adaptive window matching.

4. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 1 is characterized in that: In step S3, the iterative synthesis includes the following steps: Step S31: obtaining a first detection distance lower limit of a first resolution depth map, and obtaining a second detection distance lower limit of a second resolution depth map; Step S32: For pixel points whose depth values ​​are greater than the first detection distance lower limit, weighted fusion is performed on the pixel points corresponding to the first resolution depth map and the second resolution depth map; for pixel points whose depth values ​​are less than the first detection distance lower limit and greater than the second detection distance lower limit, the depth value of the second resolution depth value is used to obtain an initial depth map.

5. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 1, characterized in that: In step S4, the feature encoder is a convolutional neural network, and the number of layers of the convolutional neural network and the number of convolution kernels in each layer are optimized according to the size and feature complexity of the input image.

6. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 5, characterized in that: The convolutional neural network structure includes at least one of ResNet, VGG or U-Net architectures.

7. The multi-scale information fusion depth reconstruction method based on binocular camera according to claim 1, characterized in that: In step S5, the feature decoder includes a deconvolution layer and an upsampling layer, which is used to gradually restore the fused features to a high-resolution depth map.

8. A multi-scale information fusion depth reconstruction system based on a binocular camera, used to implement the multi-scale information fusion depth reconstruction method based on a binocular camera according to any one of claims 1 to 7, characterized in that: include: An acquisition module is used to obtain a left image and a right image through a binocular camera; Down-sampling the left image and the right image to obtain a first-resolution left image, a first-resolution right image, a second-resolution left image, and a second-resolution right image; A depth module, configured to obtain a first-resolution depth map according to the first-resolution left image and the first-resolution right image; Obtain a second-resolution depth map according to the second-resolution left image and the second-resolution right image; A synthesis module, configured to synthesize the first resolution depth map with the second resolution depth map to obtain an initial depth map with high accuracy and taking into account different depth distances; A fusion module, used for fusing the first resolution left image and the initial depth image using a feature encoder to obtain a fusion feature; The decoding module is used to decode the fused features through a feature decoder to obtain a high-resolution depth map after super-resolution.

9. A multi-scale information fusion depth reconstruction device based on a binocular camera, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to execute the steps of the multi-scale information fusion depth reconstruction method based on a binocular camera as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the multi-scale information fusion depth reconstruction method based on a binocular camera described in any one of claims 1 to 7 are implemented.