Multi-size fused binocular camera and electronic equipment
By using the depth maps of different resolutions obtained by downsampling, and using an effective depth synthesis solution, the problem of difficulty in real-time deep reconstruction in the existing technology in a variety of scenarios is solved, breaking through the lower limit of detection distance of traditional binocular solutions, and achieving high-precision depth reconstruction effect.
Patent Information
- Application Number
- CN202510213536.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to achieve real-time depth reconstruction in a variety of scenarios, while taking into account close-range depth detection and detail accuracy, especially in ultra-close (within 12cm) scenarios, the depth reconstruction effect is not good.
By using the depth maps of different resolutions obtained by downsampling, an effective depth synthesis scheme from near to far is synthesized to produce a correction depth map, breaking through the lower limit of detection distance of the traditional binocular scheme.
Real-time deep reconstruction in a variety of scenarios is achieved, taking into account the accuracy of different depth distances, breaking through the lower limit of detection distances of traditional binocular solutions, and significantly improving the accuracy and fidelity of reconstruction results.
Smart Images

Figure CN120147127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of depth cameras, and specifically, to a binocular camera and an electronic device with multi-size fusion. Background Art
[0002] In recent years, the demand for image-based depth reconstruction has been increasing. It is a great technical challenge to achieve depth reconstruction of various objects in a target scene, ensure effective detection of ultra-close obstacles, and guarantee its accuracy. At a fixed maximum disparity, different resolutions of image input will result in different depth details and different effective detection distances: the larger the resolution, the better the depth details, but the lower limit of depth detection will increase; the smaller the resolution, the closer the effective distance that can be detected, but there will be a degradation of depth details. Currently, there is no dedicated binocular three-dimensional depth reconstruction algorithm to balance these two aspects. Therefore, to achieve real-time depth reconstruction in diverse scenarios, while taking into account close-range depth detection and detail accuracy, it is of great significance to effectively fuse binocular information at multiple scales for depth reconstruction.
[0003] Currently, typical binocular-based depth prediction models can only perform detection with the same resolution input and cannot well balance the depth reconstruction effects at various distances, especially for ultra-close ranges (within 12 cm). For some common close-range scene depth reconstructions, they are basically ineffective. For example, CRE, although it has good effects on the details of depth prediction, has no ability to effectively detect obstacles with a distance less than 12 cm, and serious depth prediction error problems will occur.
[0004] The disclosure of the above background art content is only used to assist in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0005] To this end, the present invention uses depth maps with different resolutions obtained by downsampling. The depth maps with different resolutions have different lower limits of detection distances. Through an effective depth synthesis scheme from near to far, the depths at different limit distances are synthesized, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the lower limit of the detection distance of the traditional binocular scheme.
[0006] In a first aspect, the present invention provides a binocular camera with multi-size fusion, characterized by comprising:
[0007] A first receiver for obtaining a first image;
[0008] A second receiver for obtaining a second image;
[0009] A processor for respectively downsampling the first image and the second image to obtain a third image and a fourth image, splicing the first image and the third image to obtain a fifth image, and splicing the second image and the fourth image to obtain a sixth image; calculating according to the fifth image and the sixth image to obtain a spliced depth map, and then fusing and cropping the spliced depth map according to a detection distance lower limit to obtain a corrected depth map.
[0010] Optionally, in the multi-size fusion binocular camera, the processor includes, when processing:
[0011] Step S1: Downsample the first image and the second image to obtain a third image and a fourth image; wherein, the first image and the second image have a first resolution, the third image and the fourth image have a second resolution, and the second resolution is less than the first resolution;
[0012] Step S2: Splice the first image and the third image to obtain a fifth image, and splice the second image and the fourth image to obtain a sixth image;
[0013] Step S3: Calculate according to the fifth image and the sixth image to obtain a spliced depth map, and then fuse and crop the spliced depth map according to a detection distance lower limit to obtain a corrected depth map;
[0014] Step S4: Use a feature encoder to fuse the first image and the corrected depth map to obtain a fused feature;
[0015] Step S5: Decode the fused feature through a feature decoder to obtain a super-resolved high-resolution depth map.
[0016] Optionally, in the multi-size fusion binocular camera, in step S1, the method for downsampling the first image and the second image includes but is not limited to nearest neighbor interpolation, bilinear interpolation or bicubic interpolation.
[0017] Optionally, in the multi-size fusion binocular camera, in step S2, the splicing is to splice the first image and the third image vertically to obtain a fifth image, and splice the second image and the fourth image vertically to obtain a sixth image.
[0018] Optionally, in the multi-size fusion binocular camera, step S3 includes:
[0019] Step S31: Calculate according to the fifth image and the sixth image to obtain a spliced depth map;
[0020] Step S32: Obtain the lower limit of the first detection distance at the first resolution and obtain the lower limit of the second detection distance at the second resolution;
[0021] Step S33: For the part with the first resolution on the stitched depth map, if the depth value of a pixel point is less than the lower limit of the first detection distance, set its depth value to the depth value of the corresponding pixel point in the part with the second resolution;
[0022] Step S34: Crop the part with the first resolution from the stitched depth map to obtain a corrected depth map.
[0023] Optionally, in the multi-size fusion binocular camera, it is characterized in that between Step S33 and Step S34, it further includes:
[0024] Step S35: For the part with the first resolution on the stitched depth map, if the depth value of a pixel point is greater than the lower limit of the first detection distance, perform weighted fusion on the depth value of the part with the first resolution and the corresponding pixel points with the second resolution.
[0025] Optionally, in the multi-size fusion binocular camera, it is characterized in that in Step S4, the feature encoder adopts a convolutional neural network structure, which is used to extract the feature information in the first image and the corrected depth map and fuse the extracted feature information.
[0026] Optionally, in the multi-size fusion binocular camera, it is characterized in that in Step S5, the feature decoder is a transposed convolutional neural network decoder, and this transposed convolutional neural network decoder includes multiple transposed convolutional layers and upsampling layers, which are used to decode and upsample the fused features to obtain a super-resolved high-resolution depth map.
[0027] Optionally, in the multi-size fusion binocular camera, it is characterized in that when the processor is processing, it further includes:
[0028] Step S6: Perform post-processing on the super-resolved high-resolution depth map, and the post-processing includes filtering processing. Smooth the super-resolved high-resolution depth map through the Gaussian filtering algorithm to remove noise and artifacts.
[0029] In a second aspect, the present invention provides an electronic device, which is characterized in that it includes any one of the foregoing multi-size fusion binocular cameras.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] By generating depth maps with different resolutions and synthesizing them, the present invention can effectively balance the accuracy at different depth distances. The low-resolution depth map can obtain depth information at closer distances, while the high-resolution depth map can capture depth values with higher precision. The combination of the two makes the reconstruction result more accurate and comprehensive, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the lower limit of the detection distance of the traditional binocular solution.
[0032] Through multi-scale information fusion and feature fusion, the present invention can better retain the high-frequency details in the depth map. This detail retention is particularly important for the reconstruction of complex scenes and can significantly improve the fidelity of the reconstruction result. Through multi-scale information fusion, the present invention can effectively process the depth information of objects at long and short distances and is applicable to various scenarios such as indoor and outdoor. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:
[0034] Figure 1 It is a schematic structural diagram of a binocular camera with multi-size fusion in an embodiment of the present invention;
[0035] Figure 2 It is a step flow chart of a processor during processing in an embodiment of the present invention;
[0036] Figure 3 It is a step flow chart of obtaining a calibrated depth map in an embodiment of the present invention;
[0037] Figure 4 It is another step flow chart of obtaining a calibrated depth map in an embodiment of the present invention;
[0038] Figure 5 It is another step flow chart of a processor during processing in an embodiment of the present invention;
[0039] Figure 6 It is a schematic diagram of a binocular multi-scale depth prediction network framework in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0041] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0042] A binocular camera with multi-size fusion provided by an embodiment of the present invention aims to solve the problems existing in the prior art.
[0043] The technical solution of the present invention and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the drawings.
[0044] Due to the reason of the binocular depth calculation algorithm, the binocular camera has a detection distance lower limit related to its resolution. For a target object with a depth less than the detection distance lower limit, its depth value cannot be calculated, even if the target object appears in the binocular image. The applicant found that the detection distance lower limit is related to the resolution of the binocular image. When the resolution of the binocular image decreases, the detection distance lower limit also decreases, so that the depth value at a relatively close distance can be obtained.
[0045] The present invention uses depth maps with different resolutions obtained by downsampling. The depth maps with different resolutions have different detection distance lower limits. Through an effective depth synthesis scheme from near to far, the depths at different limit distances are synthesized, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the detection distance lower limit of the traditional binocular scheme.
[0046] Figure 1 It is a schematic structural diagram of a binocular camera with multi-size fusion in an embodiment of the present invention. As Figure 1As shown in the figure, a binocular camera with multi - size fusion in an embodiment of the present invention includes:
[0047] A first receiver for obtaining a first image.
[0048] Specifically, the first receiver is an image sensor of a camera. It captures light and converts it into an electrical signal, and then converts it into digital image data to obtain the first image. This process is based on the photoelectric effect. The pixel points on the sensor respond to light of different intensities and colors, and after processing such as analog - to - digital conversion, the first image is output. For example, in a common CMOS image sensor, each pixel unit in the pixel array can sense the intensity of the incident light, and then convert it into a digital signal through a circuit, finally forming the first image. The first receiver can be an RGB receiver, an infrared receiver, or other types of receivers.
[0049] A second receiver for obtaining a second image.
[0050] Specifically, the second receiver is of the same type as the first receiver. Preferably, the first receiver and the second receiver have exactly the same model. The image parameters generated by the first receiver and the second receiver are exactly the same, and the only difference lies in the shooting angle. There is a certain parallax between the first receiver and the second receiver, simulating the human eyes, so a depth map can be calculated through binocular technology.
[0051] A processor for respectively down - sampling the first image and the second image to obtain a third image and a fourth image, splicing the first image and the third image to obtain a fifth image, splicing the second image and the fourth image to obtain a sixth image; calculating according to the fifth image and the sixth image to obtain a spliced depth map, and then fusing and cropping the spliced depth map according to the lower limit of the detection distance to obtain a corrected depth map.
[0052] Specifically, the processor is the core processing component of the binocular camera, undertaking multiple important image - processing tasks. Specifically, it includes down - sampling the first image and the second image to obtain a third image and a fourth image; splicing the first image and the third image to obtain a fifth image, splicing the second image and the fourth image to obtain a sixth image; calculating according to the fifth image and the sixth image to obtain a spliced depth map; and then fusing and cropping the spliced depth map according to the lower limit of the detection distance to obtain a corrected depth map.
[0053] The processor downsamples the first image and the second image by executing a downsampling algorithm. The downsampling algorithm can be simple average sampling, maximum sampling, or more complex filtering sampling, etc. For example, average sampling will average the pixel values within a certain area of the image to obtain the pixel values after downsampling, thereby reducing the resolution of the image and generating the third image and the fourth image.
[0054] After obtaining the third image and the fourth image, the processor stitches the first image and the third image, and stitches the second image and the fourth image. The stitching process requires precisely aligning the coordinates and pixels of the images to ensure that there are no misalignments or overlaps in the stitched image. An image registration algorithm can be used to achieve image alignment, and then the corresponding pixels are merged to generate the fifth image and the sixth image.
[0055] The processor performs depth calculation based on the fifth image and the sixth image, which is based on the parallax principle of binocular vision. By finding matching feature points in the two images, calculating the parallax between these feature points, and combining known camera parameters (such as focal length, baseline distance, etc.), the depth value corresponding to each feature point is calculated using mathematical formulas, thereby generating a stitched depth map.
[0056] Finally, the processor performs fusion and cropping operations on the stitched depth map according to the set lower limit of the detection distance. The fusion operation can include smoothing the depth values, weighted averaging, etc. to improve the accuracy and reliability of the depth map. The cropping operation is to crop the stitched image so that its size is the same as that of the first image, obtaining the final corrected depth map, providing more accurate depth information for subsequent applications.
[0057] The processor plays a central control and data processing role in the binocular camera. It performs a series of complex processes on the original image data acquired by the first receiver and the second receiver, and finally obtains accurate depth information, providing important support for the application of the camera in fields such as 3D reconstruction, object recognition, and robot navigation.
[0058] Figure 2 It is a flowchart of the steps when a processor processes in an embodiment of the present invention. As Figure 2 shown, the steps when a processor processes in an embodiment of the present invention include:
[0059] Step S1: Downsample the first image and the second image to obtain the third image and the fourth image; wherein, the first image and the second image have a first resolution, the third image and the fourth image have a second resolution, and the second resolution is less than the first resolution.
[0060] In this step, the processor performs downsampling operations on the first image from the first receiver and the second image from the second receiver respectively. Downsampling is a technical means to reduce the image resolution. There are various common downsampling methods. For example, average downsampling will average all the pixel values within a small area of the image (such as an n×n pixel block) and use this average value as a pixel value at the corresponding position after downsampling. There is also max downsampling, which selects the pixel with the largest pixel value within the small area as the result after downsampling. In these ways, the first image and the second image with the first resolution are converted into the third image and the fourth image with the second resolution, and the second resolution is less than the first resolution.
[0061] Step S2: Stitch the first image and the third image to obtain a fifth image, and stitch the second image and the fourth image to obtain a sixth image.
[0062] In this step, after obtaining the first image, the third image, the second image, and the fourth image, the processor performs stitching operations on the first image and the third image, and on the second image and the fourth image respectively. During the stitching process, it is necessary to accurately determine the positional relationship and alignment between the images. Usually, according to the coordinate system of the images, the third image is placed at a suitable position of the first image (such as the center position or according to a certain proportional relationship), and then the pixel data of the two are merged. Similarly, the fourth image is placed at a suitable position of the second image to complete the merging of pixel data, thus obtaining the fifth image and the sixth image. The stitched fifth image and sixth image contain both the original high-resolution information (the first image and the second image parts) and the downsampled low-resolution information (the third image and the fourth image parts). This multi-size fused image enables obtaining depth maps of different sizes through a single depth calculation.
[0063] Step S3: Calculate based on the fifth image and the sixth image to obtain a stitched depth map, and then fuse and crop the stitched depth map according to the lower limit of the detection distance to obtain a corrected depth map.
[0064] In this step, based on the binocular vision principle, the processor analyzes and calculates the fifth image and the sixth image. By finding matching feature points in these two images (for example, using feature extraction algorithms such as SIFT, SURF, ORB, etc., first extract features such as corners and edges in the images, and then perform feature matching), the disparity of these feature points in the two images is calculated. According to the inverse relationship between the disparity and the object depth, and the known camera parameters (such as focal length, baseline distance, etc.), the depth value corresponding to each feature point is calculated through corresponding mathematical formulas, and then a stitched depth map is generated.
[0065] Process the stitched depth map according to the preset lower limit of the detection distance. The fusion operation may include performing smoothing filtering (such as Gaussian filtering) on the depth values to remove noise points and abnormal depth values; or performing weighted averaging on the depth values obtained from images with different resolutions, etc., to make the depth map smoother and more accurate. The cropping operation is to crop the stitched image so that its size is the same as that of the first image, obtaining the final corrected depth map, which provides more accurate depth information for subsequent applications.
[0066] Step S4: Use a feature encoder to fuse the first image and the corrected depth map to obtain fused features.
[0067] In this step, the feature encoder is a module that can extract image features and perform fusion. It receives the first image and the corrected depth map as inputs and extracts features from them respectively. For the first image, its visual features (such as color, texture, shape, etc. features) are extracted; for the corrected depth map, its depth-related features (such as the distribution of depth, the changing trend, etc. features) are extracted. Then, the two types of extracted features are fused, for example, through vector concatenation, weighted summation, etc., to obtain fused features. Based on deep learning or traditional feature extraction and fusion methods, using the structure and algorithm of the feature encoder, the visual information and depth information of the image are organically combined, so that the fused features contain more comprehensive scene information.
[0068] Step S5: Decode the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.
[0069] In this step, the feature decoder receives the fused features obtained in step S4 as input and performs decoding operations on them. Through a series of operations and transformations (which may include convolution, deconvolution, upsampling, etc. operations in deep learning), the fused features are converted into a high-resolution depth map. The feature decoder maps the information in the fused features back to the image space according to its designed structure and algorithm, restoring the high-resolution depth information and realizing the super-resolution processing of the depth map. Obtaining a depth map with a higher resolution can represent the depth information of objects in the scene more finely, meeting some application scenarios with higher requirements for the accuracy of the depth map, such as high-precision 3D reconstruction, fine operations of robots, etc.
[0070] In some embodiments, the methods for downsampling the first image and the second image include, but are not limited to, nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation. During the process of downsampling the first image and the second image, a variety of different methods are adopted, including but not limited to nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation. Nearest neighbor interpolation is a relatively simple downsampling method. When downsampling an image, for each pixel point in the output image, the pixel point closest to it in the original image (i.e., the first image or the second image) is found, and then the pixel value of this closest pixel point is assigned to the corresponding pixel point in the output image. For example, if the size of the image is to be reduced to half of the original, at a certain position in the new image, by calculating its corresponding position in the original image (since the size is reduced, one pixel in the new image corresponds to a region in the original image), and then selecting the color value of the original pixel point closest to the position of the new image pixel point as the value of the new pixel point.
[0071] Bilinear interpolation calculates the pixel value of a pixel point in the output image using the pixel values of the surrounding four pixel points. For each pixel point in the output image, first find its corresponding position in the original image, which generally does not exactly fall on a pixel point of the original image. Then find the four nearest pixel points around this position, and calculate the value of the target pixel point through linear weighting according to the relative distances of these four pixel points from the target position. For example, for a two-dimensional image, linear interpolation calculations are performed separately in the horizontal and vertical directions to finally obtain the color value of the target pixel point.
[0072] Bicubic interpolation is a more complex but better-performing downsampling method. It not only considers the pixel values of the surrounding four pixel points but also takes into account the influence of some pixel points around these four pixel points. The pixel value of a pixel point in the output image is calculated through a bicubic function, which performs a more refined weighting calculation based on the relative positions and distances of the target pixel point from multiple surrounding pixel points. In practical applications, bicubic interpolation can better preserve the details and high-frequency information of the image.
[0073] In some embodiments, in step S2, the splicing is to splice the first image and the third image vertically to obtain a fifth image, and splice the second image and the fourth image vertically to obtain a sixth image. For the operation of splicing the first image and the third image to obtain the fifth image and the operation of splicing the second image and the fourth image to obtain the sixth image, the operation of splicing the first image and the third image to obtain the fifth image is taken as an example for illustration. First, we have obtained the first image with the first resolution and the third image with the second resolution (and the second resolution is less than the first resolution) after downsampling through the previous steps. These two images are the basic data for splicing. When splicing vertically, it is necessary to ensure that the sizes (widths) of the two images in the horizontal direction are matched, or they can be made to match through appropriate adjustment. Then, place the third image above or below the first image (the specific placement position is determined according to a preset rule). For example, if the third image is placed above the first image, then from the perspective of the image coordinate system, the bottom edge of the third image is aligned with the top edge of the first image vertically. Then, in the order of pixels, combine the pixels of the third image row by row with the pixels of the first image. That is to say, first arrange the first row of pixels of the third image in front, and then immediately arrange the first row of pixels of the first image, and so on, until all the row pixels of the third image and the first image are arranged, thus forming a new image, that is, the fifth image. The size of this new image in the vertical direction is the sum of the vertical sizes of the first image and the third image (the size in the horizontal direction is the same as the larger of the two widths or the adjusted width), and it contains the information of the original high-resolution first image and the downsampled third image information.
[0074] Figure 3 The flowchart shows the steps for obtaining a corrected depth map in an embodiment of the present invention. As Figure 3 shown, the steps for obtaining a corrected depth map in an embodiment of the present invention include:
[0075] Step S31: Calculate a spliced depth map based on the fifth image and the sixth image.
[0076] In this step, this step is based on the parallax principle of binocular vision. The two perspectives of the binocular camera (corresponding to the fifth image and the sixth image) result in a difference in the positions of the same object in the two images, and this difference is called parallax. Parallax is inversely proportional to the distance from the object to the camera. By calculating the parallax and combining the internal parameters of the camera (such as focal length, baseline distance, etc.), the depth information of the object can be calculated.
[0077] Step S32: Obtain a first detection distance lower limit with the first resolution and a second detection distance lower limit with the second resolution.
[0078] In this step, the lower detection distance limit refers to the minimum distance at which the camera can accurately detect the depth of an object. Due to the characteristics of the binocular algorithm, images with different resolutions have different lower detection distance limits. For a fixed binocular camera, when its resolution is fixed, its lower detection distance limit can also be determined.
[0079] Step S33: For the part of the stitched depth map with the first resolution, if the depth value of a pixel point is less than the first lower detection distance limit, set its depth value to the depth value of the corresponding pixel point in the part with the second resolution.
[0080] In this step, for a binocular camera, the depth range can be divided into three intervals according to the first lower detection distance limit and the second lower detection distance limit: greater than the first lower detection distance limit, less than the first lower detection distance limit and greater than the second lower detection distance limit, and less than the second lower detection distance limit.
[0081] For pixel points less than the first lower detection distance limit and greater than the second lower detection distance limit, use the depth value corresponding to the second-resolution depth map. At this time, the depth value in the first-resolution depth map is empty or an incorrect depth value. Directly using the depth value corresponding to the second-resolution depth map can obtain an accurate depth value.
[0082] For pixel points less than the second lower detection distance limit, there is no reliable depth value.
[0083] Step S34: Crop the part of the stitched depth map with the first resolution to obtain a corrected depth map.
[0084] In this step, after the fusion and correction of the depth values are completed, only the part of the stitched depth map with the first resolution is retained to obtain the final corrected depth map. First, determine the area to be cropped according to the position and size information of the part with the first resolution in the stitched depth map. Use an image cropping algorithm to crop the part of the stitched depth map with the first resolution from the entire image to obtain a corrected depth map. The corrected depth map only contains the depth information of the first resolution and has undergone the fusion and correction in the previous steps, and its depth value is more accurate and reliable.
[0085] According to the description of this embodiment, those skilled in the art can easily think of using more levels of downsampling to obtain more depth maps with different resolutions, so as to obtain the third lower detection distance limit, the fourth lower detection distance limit, etc., so as to continuously approach the detection limit of the depth camera and obtain depth values at closer distances. These are all obtained under the inspiration of the present invention and do not depart from the core inventive points of the present invention, and also fall within the protection scope of the present invention.
[0086] Figure 4This is a flowchart of another step for obtaining a corrected depth map in an embodiment of the present invention. As Figure 4 shown, compared with the foregoing embodiment, another step for obtaining a corrected depth map in an embodiment of the present invention further includes, between step S33 and step S34:
[0087] Step S35: For the part with the first resolution on the stitched depth map, if the depth value of a pixel point is greater than the lower limit of the first detection distance, then perform weighted fusion on the depth value of the part with the first resolution and the pixel points corresponding to the second resolution.
[0088] In this step, for pixel points greater than the lower limit of the first detection distance, weighted fusion is performed according to the depth values corresponding to the first-resolution depth map and the second-resolution depth map. The fusion method can be freely selected. For example, different weights are assigned according to the change of the depth value. The greater the depth value, the greater the weight of the first-resolution depth map; the smaller the depth value, the smaller the weight of the first-resolution depth map. For another example, the weight of the first-resolution depth map is set to 1, and the weight of the second-resolution depth map is set to 0, that is, all depth data of the first-resolution depth map is adopted.
[0089] In some embodiments, the feature encoder in step S4 adopts a convolutional neural network structure, which is used to extract the feature information in the first image and the corrected depth map, and fuse the extracted feature information. In step S4, the feature encoder plays a crucial role. Its main task is to extract features from the first image and the corrected depth map, and fuse this feature information. The first image contains rich visual information of the scene, such as color, texture, and shape, etc.; while the corrected depth map provides the depth information of the objects in the scene. By fusing these two different types of information, more comprehensive and representative features can be obtained, laying a foundation for generating a high-quality super-resolution depth map in subsequent steps.
[0090] The convolutional neural network structure is selected as the feature encoder because CNN has significant advantages in image feature extraction:
[0091] Local perception: The convolutional layer of CNN slides a small convolutional kernel on the image for convolution operation, which can capture local features in the image. For the first image and the corrected depth map, this local perception ability can effectively extract local feature information such as edges and corners.
[0092] Parameter sharing: In the convolutional layer, the same convolutional kernel shares parameters throughout the image, greatly reducing the number of model parameters, reducing the computational complexity, and at the same time helping to improve the generalization ability of the model.
[0093] Hierarchical Feature Extraction: A CNN is typically composed of multiple convolutional layers and pooling layers stacked together, enabling hierarchical feature extraction. From simple features at the bottom layer (such as edges and textures) to abstract features at the high layer (such as object categories and semantic information), higher-level feature representations are gradually extracted.
[0094] The feature extraction of the first image mainly includes the following aspects:
[0095] Input Layer: The first image is passed as input to the convolutional neural network.
[0096] Convolutional Layer: The convolutional layer contains multiple convolutional kernels, and each convolutional kernel extracts different features from the first image through convolution operations. For example, a 3x3 convolutional kernel can extract edge features in the image. The convolution operation generates multiple feature maps, and each feature map corresponds to a type of feature.
[0097] Activation Function: After the convolutional layer, an activation function (such as ReLU) is usually applied to introduce non-linearity and enhance the model's expressive power. The ReLU function sets values less than 0 to 0 and keeps values greater than 0 unchanged, effectively alleviating the vanishing gradient problem.
[0098] Pooling Layer: The pooling layer is used to downsample the feature maps, reducing the size of the feature maps while retaining important feature information. Common pooling operations include max pooling and average pooling. Max pooling selects the maximum value in each pooling window as the output, highlighting the saliency of features.
[0099] Feature Extraction of the Rectified Depth Map
[0100] The feature extraction process of the rectified depth map is similar to that of the first image. Similarly, through the combination of convolutional layers, activation functions, and pooling layers, depth-related feature information is extracted from the rectified depth map. Due to the different data characteristics of the depth map from the first image, the convolutional kernels will learn features related to depth changes and the distance relationship of objects.
[0101] Feature Fusion Process
[0102] After extracting the features of the first image and the rectified depth map respectively, these features need to be fused. The common feature fusion methods are as follows:
[0103] Concatenation Fusion: The feature maps extracted from the first image and the rectified depth map are concatenated in the channel dimension. For example, if the feature map of the first image has 64 channels and the feature map of the rectified depth map has 32 channels, the concatenated feature map will have 96 channels. This method is simple and direct, and can retain all the feature information of the two inputs.
[0104] Weighted fusion: Different weights are assigned to the feature maps of the first image and the corrected depth map, and then they are weighted and summed. The weights can be learned through training to adaptively adjust the importance of the two input features. For example, in some scenarios, the features of the first image may be more important, and a larger weight can be assigned to it.
[0105] Output result
[0106] After feature extraction and fusion, the feature encoder outputs fused features. These fused features contain the visual information of the first image and the depth information of the corrected depth map, and have richer semantics and stronger expressive power. The fused features will be passed as input to the subsequent feature decoder for generating the super-resolved high-resolution depth map.
[0107] In some embodiments, in step S5, the feature decoder is a deconvolution neural network decoder, which includes multiple deconvolution layers and upsampling layers for decoding and upsampling the fused features to obtain the super-resolved high-resolution depth map. In step S5, the feature decoder undertakes a key task, and its core objective is to decode and upsample the fused features output by the feature encoder in step S4 to obtain the super-resolved high-resolution depth map. Although the fused features contain rich information, they are an abstract representation processed by the feature encoder and need to be restored and upsampled to a higher resolution by the feature decoder to meet the high-precision requirements for depth maps in practical applications.
[0108] The deconvolution neural network decoder is mainly composed of multiple deconvolution layers and upsampling layers. The functions and working principles of these two types of layers are introduced below.
[0109] The deconvolution layer is also called the transposed convolution layer, and its function is the opposite of that of the convolution layer. The convolution layer reduces the size of the input feature map through convolution operations, while the deconvolution layer enlarges the size of the input feature map through specific operations. From a mathematical principle perspective, the deconvolution operation is an inverse operation of the convolution operation, but it is not strictly reversible because some information may be lost during the convolution process.
[0110] The deconvolution layer enlarges the size of the feature map by padding zeros around each pixel of the input feature map and then performing convolution operations using a learnable convolution kernel. For example, a feature map with an input of H×W×C (height H, width W, number of channels C) may increase in height and width after passing through the deconvolution layer to obtain a larger-sized feature map.
[0111] Function: The transposed convolution layer can learn how to recover more detailed information from low-resolution fused features. While expanding the size of the feature map, it can adjust the parameters of the convolution kernel according to the training data to meet different requirements for depth information recovery.
[0112] The main purpose of the upsampling layer is to directly increase the spatial size of the feature map. Common upsampling methods include nearest-neighbor interpolation, bilinear interpolation, etc. Nearest-neighbor interpolation directly copies the value of each pixel in the input feature map to multiple adjacent pixel positions in the output feature map; bilinear interpolation calculates the value of the pixel in the output feature map by weighted averaging the values of adjacent pixels in the input feature map.
[0113] Taking nearest-neighbor interpolation as an example, assume that a 2×2 feature map needs to be upsampled to a 4×4 feature map. Then each pixel in the input feature map will be copied to the corresponding 2×2 area in the output feature map. Bilinear interpolation will calculate the weighted coefficients according to the position and distance of the input pixels, and then perform weighted summation to obtain the value of the output pixel.
[0114] The upsampling layer can quickly increase the size of the feature map, providing more space for the subsequent transposed convolution layer to further refine the features. Without introducing too many parameters, it can initially improve the resolution of the feature map.
[0115] When the transposed convolution neural network decoder processes the fused features, usually the transposed convolution layer and the upsampling layer are alternated. The specific process is as follows:
[0116] Initial stage: The fused features obtained in step S4 are passed as input to the decoder.
[0117] Alternating processing: First, the upsampling layer may be used to initially expand the size of the fused features and quickly improve the resolution of the feature map. Then, the upsampled feature map is input into the transposed convolution layer, and the transposed convolution layer further refines and adjusts the features through the learned convolution kernel parameters to recover more depth details. Such a process is repeated multiple times. Each alternation makes the resolution of the feature map continuously increase, and at the same time, the depth information becomes more abundant and accurate.
[0118] Output result: After being processed by multiple transposed convolution layers and upsampling layers, the finally output feature map is the super-resolved high-resolution depth map. The resolution of this depth map has been significantly improved, and it can represent the depth information of objects in the scene more clearly, providing more accurate data support for subsequent applications (such as 3D reconstruction, object detection, etc.).
[0119] The performance of the deconvolution neural network decoder largely depends on the optimization of its parameters. During the training process, a loss function is usually used to measure the difference between the super-resolved high-resolution depth map and the true high-resolution depth map. Common loss functions include mean squared error (MSE) loss, structural similarity index (SSIM) loss, etc. Through the backpropagation algorithm, the convolution kernel parameters of the deconvolution layer are continuously adjusted to minimize the value of the loss function, thereby improving the performance of the decoder and enabling it to generate high-quality super-resolved depth maps more accurately.
[0120] Figure 5 This is the flowchart of the steps when another processor in the embodiment of the present invention is processing. As Figure 5 shown, compared with the foregoing embodiment, the steps when another processor in the embodiment of the present invention is processing further include:
[0121] Step S6: Perform post-processing on the super-resolved high-resolution depth map. The post-processing includes filtering processing, and the super-resolved high-resolution depth map is smoothed through the Gaussian filtering algorithm to remove noise and artifacts.
[0122] In this step, noise and artifacts are processed. Noise may be generated due to factors such as the imperfection of the image acquisition device and errors in the feature extraction and fusion process; artifacts may be inappropriate results that occur when the deconvolution neural network decoder learns and reconstructs depth information. These noise and artifacts will affect the quality and accuracy of the depth map, and thus have an adverse impact on subsequent applications based on the depth map (such as 3D reconstruction, object recognition, etc.). Therefore, it is necessary to perform post-processing on the super-resolved high-resolution depth map to improve its quality.
[0123] Filtering processing is a commonly used method in image post-processing for removing noise and artifacts in images. In this step, the Gaussian filtering algorithm is selected because it has the following advantages:
[0124] Good smoothing effect: Gaussian filtering is a linear smoothing filter that convolves the image based on the Gaussian function. The Gaussian function has the shape of a bell curve, and its characteristic is that the points closer to the center have larger weights, and the points farther from the center have smaller weights. Through this weighted average method, the image can be effectively smoothed, the influence of noise can be reduced, and the edges and details of the image can be retained as much as possible.
[0125] Strong adjustability: The smoothing degree of Gaussian filtering can be controlled by adjusting the standard deviation of the Gaussian function (also known as the size of the Gaussian kernel). A larger standard deviation will make the filtering effect more obvious, and it can remove noise in a larger range, but it may cause the image edges to be blurred; a smaller standard deviation can better retain the details of the image while removing noise.
[0126] The core of the Gaussian filtering algorithm is to perform a convolution operation on the super-resolved high-resolution depth map using a two-dimensional Gaussian kernel. The specific steps are as follows:
[0127] Define the Gaussian kernel: According to the required degree of smoothness, determine the size of the Gaussian kernel (usually a square matrix of odd size, such as 3x3, 5x5, etc.) and the standard deviation. The value of each element in the Gaussian kernel is calculated by the Gaussian function, and the formula is:
[0128]
[0129] where (x, y) are the coordinates of the element in the Gaussian kernel, and σ is the standard deviation.
[0130] Normalization processing: In order to ensure that the brightness range of the filtered image remains unchanged, it is necessary to perform normalization processing on the Gaussian kernel, that is, make the sum of all elements in the Gaussian kernel equal to 1.
[0131] Convolution operation: Slide the defined and normalized Gaussian kernel point by point on the super-resolved high-resolution depth map, and perform weighted summation on each pixel and its neighborhood. Specifically, for each pixel in the depth map, multiply the pixel values in its neighborhood by the corresponding elements in the Gaussian kernel, and then sum all the products. The resulting value is used as the new value of the pixel after filtering. For example, for a 3x3 Gaussian kernel, when processing a certain pixel in the depth map, the values of the pixel and its surrounding 8 neighboring pixels will be considered, and weighted summation will be performed according to the weights of the Gaussian kernel.
[0132] By performing smoothing processing on the super-resolved high-resolution depth map through the Gaussian filtering algorithm, noise and artifacts in the image can be significantly removed. The pixel values of noise points will be weighted and averaged by the surrounding pixel values, making them closer to the normal pixel values of the surrounding area; the artifact part will also be improved due to the smoothing operation, making the overall visual effect of the depth map more natural and clear. At the same time, since Gaussian filtering has certain advantages in preserving edges and details, the processed depth map can still retain the important structural information of the objects in the scene, providing more reliable data support for subsequent applications.
[0133] After post-processing, the quality of the super-resolved high-resolution depth map is further improved, making it more in line with the requirements of practical applications. Whether it is to construct a more accurate 3D model in 3D reconstruction or to more accurately judge the position and shape of an object in object recognition, a high-quality depth map plays a crucial role. Therefore, the post-processing in step S6 is an indispensable link in the entire multi-scale fusion binocular camera processing flow.
[0134] Figure 6 This is a schematic diagram of a binocular multi-scale depth prediction network framework in an embodiment of the present invention. Figure 6Three - level downsampling is adopted to obtain the left second images at scales of 1 / 2, 1 / 4, and 1 / 8 respectively. The following combines with Figure 6 for a detailed description.
[0135] Downsample the left second image obtained by the binocular camera to get the left second images at scales of 1 / 2, 1 / 4, and 1 / 8;
[0136] Use the SGM algorithm with the same 128 - disparity to process the left second images at resolutions of 1 / 2, 1 / 4, and 1 / 8 respectively, and quickly obtain the depth maps of each resolution;
[0137] Calculate the lower limit of the depth limit for each resolution according to the minimum detection depth formula. The minimum detection depth calculation formula is as follows:
[0138] (focallength / scale*baseline) / Maxdisparity
[0139] Where focallength is the camera focal length, scale is the downsampling scale, baseline is the baseline length of the binocular camera, and Maxdisparity is the maximum disparity of binocular matching. Example: When the original image resolution is 1280*800, the focal length is 611, the maximum disparity is 128, and the baseline distance is 100: The lower limit of the limit depth at the 1 / 2 scale is about 24 cm, the lower limit of the limit depth at the 1 / 4 scale is about 12 cm, and the lower limit of the limit depth at the 1 / 8 scale is about 6 cm.
[0140] Upsample the 1 / 8 - resolution depth map by 2 times, calculate the obtained lower limit of the limit depth, take the pixels in the 1 / 8 - resolution depth map with depths between [6 cm, 12 cm] to generate a mask range, and according to this mask, replace the depths in the 1 / 4 - resolution depth map with the depths in the 1 / 8 - resolution depth map to obtain a synthesized 1 / 4 - resolution depth map.
[0141] Further upsample the synthesized 1 / 4 - resolution depth map by 2 times, take the pixels in the synthesized 1 / 4 - resolution depth map with depths between [6 cm, 24 cm] to generate a mask range, and according to this mask, replace the depths in the 1 / 2 - resolution depth map with the depths in the synthesized 1 / 4 - resolution depth map to obtain a synthesized 1 / 2 - resolution depth map. The lower limit of the effective limit detection distance of this synthesized 1 / 2 - resolution depth map reaches 6 cm, while retaining the accuracy of other long - distance depths.
[0142] Use Figure 6 the feature encoder shown in to fuse and process the first image obtained by the binocular camera and the synthesized depth map, and infer the fused features of the first image and the initial depth;
[0143] The fused features are decoded by the feature decoder to obtain a corrected and super-resolved depth map with high precision and high resolution.
[0144] This specification also provides an electronic device, including the multi-size fused binocular camera in any of the foregoing embodiments. It should be noted that the electronic device in this embodiment is only exemplary and is set for the convenience of those skilled in the art to understand, and should not constitute any limitation to the protection scope of the present invention.
[0145] This kind of electronic device combines the mobility of a mobile device with the depth perception ability of a multi-size fused binocular camera. The mobile device provides the mobility of the device, enabling it to collect and process data at different spatial positions; while the multi-size fused binocular camera is responsible for acquiring the image and depth information of the scene, providing data support for various applications of the electronic device.
[0146] The electronic device includes a mobile device and a multi-size fused binocular camera.
[0147] Description of the mobile device:
[0148] Functional overview: The mobile device is the basic support part of the electronic device. It provides the mobile ability for the entire device, enabling the device to move freely in different environments and expanding the scope of data collection.
[0149] Possible types:
[0150] Wheel-type mobile platform: For example, a robotic vehicle, which is usually equipped with multiple wheels and can move flexibly on a flat ground. The driving system of the wheels can be motor-driven, and by precisely controlling the rotation speed and steering of the motor, actions such as the device moving forward, backward, and turning can be achieved.
[0151] Flying mobile platform: Such as a drone, which uses propellers to generate lift and thrust and can fly freely in the air. Drones have high mobility and can collect data at different heights and angles, and are suitable for large-scale scene monitoring and mapping.
[0152] Internal composition:
[0153] Power system: Provides power for the mobile device to enable it to move. For a wheel-type mobile platform, the power system usually includes a motor, a battery, and a transmission device; for a flying mobile platform, the power system includes a motor, propellers, and a battery.
[0154] Control system: Responsible for controlling the movement direction, speed, and attitude of the mobile device. It usually consists of a controller, sensors (such as gyroscopes, accelerometers, etc.), and actuators. The controller calculates appropriate control instructions based on the information fed back by the sensors and sends them to the actuators to achieve precise motion control.
[0155] Communication system: Used for data transmission and communication with other devices. It can achieve data interaction with a multi-size fusion binocular camera, transmit the data collected by the camera to the processing unit of the mobile device for processing; at the same time, it can also communicate with external devices (such as remote control terminals) to achieve remote control and data sharing.
[0156] Description of multi-size fusion binocular camera:
[0157] Function overview: The multi-size fusion binocular camera is the core perception component of the electronic device. It obtains images from different perspectives through two receivers, and uses a processor to process the images. Finally, a super-resolved high-resolution depth map is obtained, providing three-dimensional information of the scene for the electronic device.
[0158] Specific composition and working principle:
[0159] The first receiver and the second receiver: Used to obtain the first image and the second image respectively. They are usually cameras based on image sensor technology (such as CCD or CMOS). By receiving light and converting it into an electrical signal, a digital image is generated.
[0160] Processor: Performs a series of processes on the first image and the second image, including downsampling, stitching, depth calculation, feature fusion, decoding, and post-processing. Finally, a super-resolved high-resolution depth map is obtained. The specific processing process is as follows:
[0161] Downsample the first image and the second image to obtain the third image and the fourth image to reduce the data volume and obtain information at different scales.
[0162] Stitch the first image and the third image to obtain the fifth image, and stitch the second image and the fourth image to obtain the sixth image to achieve the fusion of multi-size images.
[0163] Calculate based on the fifth image and the sixth image to obtain a stitched depth map, and perform fusion and cropping on it according to the lower limit of the detection distance to obtain a corrected depth map.
[0164] Use a feature encoder to fuse the first image and the corrected depth map to obtain fused features.
[0165] Decode the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.
[0166] Perform post-processing on the super-resolved high-resolution depth map, such as Gaussian filtering, to remove noise and artifacts.
[0167] Workflow
[0168] Data acquisition: The mobile device moves according to a preset path or instruction, while the first receiver and the second receiver of the multi-size fusion binocular camera continuously acquire image data of the scene.
[0169] Data processing: The processor of the camera processes the acquired image data and generates a super-resolved high-resolution depth map according to the above steps.
[0170] Application execution: The electronic device executes corresponding application tasks according to the depth map information obtained after processing. For example, in robot navigation, the position and distance of obstacles ahead are judged based on the depth map, and an appropriate moving path is planned; in 3D reconstruction, a 3D model of the scene is constructed using the depth map.
[0171] Feedback control: During the application execution process, the electronic device will perform feedback control on the movement of the mobile device according to the actual situation. For example, if an obstacle is detected ahead, the mobile device will automatically adjust its movement direction to avoid the obstacle.
[0172] The main application scenarios are as follows:
[0173] Robot navigation: The electronic device can act as an autonomous mobile robot, utilize the depth map information obtained by the multi-size fusion binocular camera, sense the surrounding environment in real time, avoid obstacles, plan the optimal moving path, and achieve autonomous navigation.
[0174] 3D reconstruction: In the fields of building surveying and mapping, cultural relic protection, etc., the electronic device can collect image and depth information of the scene by moving, and then use this information to construct a high-precision 3D model to provide basic data for subsequent analysis and research.
[0175] Intelligent security: The electronic device can be installed in the monitoring area, utilize the depth perception ability of the binocular camera, accurately detect and identify the position, posture and movement trajectory of personnel and objects, and achieve intelligent security monitoring.
[0176] Augmented reality (AR) and virtual reality (VR): In AR and VR applications, the electronic device can provide users with more real scene depth information, enhancing the user's immersion and interaction experience. For example, in an AR game, the fusion and interaction of virtual objects and the real scene are realized according to the depth map information.
[0177] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
[0178] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention.
Claims
1. A multi-size fusion binocular camera, characterized in that: include: A first receiver, configured to obtain a first image; a second receiver, configured to obtain a second image; The processor is used to downsample the first image and the second image respectively to obtain a third image and a fourth image, stitch the first image and the third image to obtain a fifth image, and stitch the second image and the fourth image to obtain a sixth image; perform calculations based on the fifth image and the sixth image to obtain a stitched depth map, and then fuse and crop the stitched depth map based on a lower limit of the detection distance to obtain a corrected depth map.
2. The multi-size fusion binocular camera according to claim 1, characterized in that: The processor includes: Step S1: downsampling the first image and the second image to obtain a third image and a fourth image; wherein the first image and the second image have a first resolution, the third image and the fourth image have a second resolution, and the second resolution is smaller than the first resolution; Step S2: stitching the first image and the third image to obtain a fifth image, and stitching the second image and the fourth image to obtain a sixth image; Step S3: performing calculations based on the fifth image and the sixth image to obtain a stitched depth map, and then fusing and cropping the stitched depth map according to a lower limit of the detection distance to obtain a corrected depth map; Step S4: using a feature encoder to fuse the first image and the corrected depth map to obtain a fused feature; Step S5: Decode the fused features through a feature decoder to obtain a high-resolution depth map after super-resolution.
3. The multi-scale fusion binocular camera according to claim 2, characterized in that: In step S1, the manner of downsampling the first image and the second image includes but is not limited to nearest neighbor interpolation, bilinear interpolation or bicubic interpolation.
4. The multi-scale fusion binocular camera according to claim 2, characterized in that: In step S2, the stitching is to stitch the first image and the third image in a vertical direction to obtain a fifth image, and to stitch the second image and the fourth image in a vertical direction to obtain a sixth image.
5. The multi-scale fusion binocular camera according to claim 2, characterized in that: Step S3 includes: Step S31: performing calculation according to the fifth image and the sixth image to obtain a spliced depth map; Step S32: obtaining a first detection distance lower limit of a first resolution, and obtaining a second detection distance lower limit of a second resolution; Step S33: for a portion of the stitched depth map having the first resolution, if the depth value of a pixel point is less than the first detection distance lower limit, setting the depth value of the pixel point corresponding to the portion having the second resolution; Step S34: cropping a portion having a first resolution from the stitched depth map to obtain a corrected depth map.
6. The multi-scale fusion binocular camera according to claim 5, characterized in that: The steps between step S33 and step S34 also include: Step S35: For a portion of the stitched depth map having a first resolution, if the depth value of a pixel is greater than the first detection distance lower limit, weighted fusion is performed on the depth value of the first resolution portion and the pixel of the second resolution portion.
7. The multi-scale fusion binocular camera according to claim 2, characterized in that: The feature encoder in step S4 adopts a convolutional neural network structure to extract feature information from the first image and the corrected depth map, and fuse the extracted feature information.
8. The multi-scale fusion binocular camera according to claim 2, characterized in that: In step S5, the feature decoder is a deconvolutional neural network decoder, which includes multiple deconvolution layers and upsampling layers, and is used to decode and upsample the fused features to obtain a high-resolution depth map after super-resolution.
9. The multi-scale fusion binocular camera according to claim 2, characterized in that: The processor also includes: Step S6: performing post-processing on the super-resolution high-resolution depth map, wherein the post-processing includes filtering processing, and performing smoothing processing on the super-resolution high-resolution depth map by using a Gaussian filtering algorithm to remove noise and artifacts.
10. An electronic device, characterized in that: A binocular camera with multiple sizes fusion comprising any one of claims 1-9.
Citation Information
Cited By
End-to-end infrared super-resolution reconstruction method
CN121258794A