Multi-scale information fusion structured light camera and electronic equipment
By downsampling and parallax calculation of structured light images, and combining with the fusion method of detecting the lower limit of distance, a correction depth map is generated, which solves the problem that deep reconstruction algorithms in the prior art are difficult to take into account resolution and close-range detection, and achieves high-precision multi-scale depth reconstruction.
Patent Information
- Application Number
- CN202510213581.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The existing structured light three-dimensional depth reconstruction algorithm is difficult to take into account the needs of depth details and close-range depth detection of high-resolution images, and there is a lack of effective multi-scale information fusion method.
The second structured light image is obtained by downsampling the first structured light image, and the first depth map and the second depth map are obtained using parallax calculation, and fuse according to the lower limit of the detection distance to generate a correction depth map.
Real-time deep reconstruction in a variety of scenarios is achieved, taking into account close-range depth detection and detail accuracy, and breaking through the lower limit of detection distance of traditional structured light solutions.
Smart Images

Figure CN120151497A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of structured light cameras, and in particular, to a multi-scale information fusion structured light camera and an electronic device. Background Art
[0002] In the current field of three-dimensional vision measurement, structured light cameras are widely used in many fields such as industrial inspection, robot navigation, virtual reality, etc. due to their advantages of high precision and non-contact. The structured light camera projects structured light onto the surface of the target object, and then the receiver receives the reflected signal, thereby obtaining the three-dimensional information of the target object.
[0003] In recent years, the demand for depth reconstruction based on images has been increasing. It is a great technical challenge to achieve depth reconstruction of various targets in the target scene, ensure effective detection of ultra-close obstacles, and at the same time ensure its accuracy. With a fixed maximum disparity, different resolutions of image input will result in different depth details and different effective detection distances: the larger the resolution, the better the depth details, but the lower limit of depth detection will become larger; the smaller the resolution, the closer the effective distance that can be detected, but there will be a degradation of depth details. Currently, there is no dedicated structured light three-dimensional depth reconstruction algorithm to balance these two aspects. Therefore, to achieve real-time depth reconstruction in diverse scenarios, while taking into account close-range depth detection and detail accuracy, it is of great significance to effectively fuse multi-scale structured light information for depth reconstruction.
[0004] The disclosure of the above background art content is only used to assist in understanding the inventive concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0005] To this end, based on the structured light camera, the present invention downsamples the first structured light image to obtain a second structured light image, respectively calculates the first depth map and the second depth map using disparity, and then performs fusion according to the lower limit of the detection distance, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the lower limit of the detection distance of the traditional structured light scheme.
[0006] In a first aspect, the present invention provides a multi-scale information fusion structured light camera, which is characterized by comprising:
[0007] A structured light projector for projecting structured light;
[0008] A structured light receiver for receiving the reflected signal of the structured light;
[0009] A processor, configured to generate a first structured light image based on the reflected signal, downsample the first structured light image to obtain a second structured light image; calculate a disparity between the first structured light image and a first reference image to obtain a first depth map, and calculate a disparity between the second structured light image and a second reference image to obtain a second depth map; fuse the first depth map and the second depth map according to a lower detection distance limit to obtain a corrected depth map.
[0010] Optionally, in the multi-scale information fusion structured light camera, the processor includes, when processing:
[0011] Step S1: Generate a first structured light image based on the reflected signal; downsample the structured light image to obtain a second structured light image;
[0012] Step S2: Calculate a disparity between the first structured light image and a first reference image to obtain a first depth map, and calculate a disparity between the second structured light image and a second reference image to obtain a second depth map;
[0013] Step S3: Fuse the first depth map and the second depth map according to a lower detection distance limit to obtain a corrected depth map with high accuracy and taking into account different depth distances;
[0014] Step S4: Use a feature encoder to fuse the first structured light image and the corrected depth map to obtain fused features;
[0015] Step S5: Decode the fused features through a feature decoder to obtain a super-resolved high-resolution depth map.
[0016] Optionally, in the multi-scale information fusion structured light camera, when downsampling the structured light image in Step S1 to obtain a second structured light image, the downsampling method used is one of mean downsampling, max downsampling, or bilinear interpolation downsampling.
[0017] Optionally, in the multi-scale information fusion structured light camera, the first structured light image has the same resolution as the first reference image; the second structured light image has the same resolution as the second reference image.
[0018] Optionally, in the multi-scale information fusion structured light camera, Step S3 includes the following steps:
[0019] Step S41: Obtain a first lower detection distance limit of the first depth map and a second lower detection distance limit of the second depth map;
[0020] Step S42: For the pixel points whose depth values are less than the lower limit of the first detection distance and greater than the lower limit of the second detection distance, use the depth values of the depth values of the second resolution to obtain a corrected depth map.
[0021] Optionally, in the multi-scale information fusion structured light camera, between step S41 and step S42, it further includes:
[0022] Step S43: For the pixel points whose depth values are greater than the lower limit of the first detection distance, perform weighted fusion on the pixel points corresponding to the first depth map and the second depth map.
[0023] Optionally, in the multi-scale information fusion structured light camera, in step S4, the feature encoder includes a plurality of convolutional layers and pooling layers. The convolutional layers are used to extract the features of the first structured light image and the corrected depth map, and the pooling layers are used to perform downsampling on the extracted features to generate the fusion features.
[0024] Optionally, in the multi-scale information fusion structured light camera, in step S5, the feature decoder includes a plurality of upsampling layers and convolutional layers. The upsampling layers are used to perform upsampling on the fusion features, and the convolutional layers are used to refine the upsampled features to finally obtain the super-resolved high-resolution depth map.
[0025] In a second aspect, the present invention provides a multi-scale information fusion structured light camera, which includes:
[0026] A structured light projector for projecting structured light;
[0027] A structured light receiver for receiving the reflected signal of the structured light;
[0028] A processor for generating a first structured light image according to the reflected signal, downsampling the first structured light image to obtain a second structured light image, and downsampling the second structured light image to obtain a third structured light image; calculating the disparity between the first structured light image and a first reference image to obtain a first depth map, calculating the disparity between the second structured light image and a second reference image to obtain a second depth map, and calculating the disparity between the third structured light image and a third reference image to obtain a third depth map; fusing the first depth map, the second depth map, and the third depth map according to the lower limit of the detection distance to obtain a corrected depth map.
[0029] In a third aspect, the present invention provides an electronic device, which includes the multi-scale information fusion structured light camera described in any one of the foregoing.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] Based on a structured light camera, the present invention downsamples the first structured light image to obtain a second structured light image, calculates the first depth map and the second depth map respectively using disparity, and then performs fusion according to the lower limit of the detection distance, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the lower limit of the detection distance of the traditional structured light scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:
[0033] Figure 1 It is a schematic structural diagram of a multi-scale information fusion structured light camera in an embodiment of the present invention;
[0034] Figure 2 It is a step flow chart of a processor during processing in an embodiment of the present invention;
[0035] Figure 3 It is a step flow chart of obtaining a calibrated depth map in an embodiment of the present invention;
[0036] Figure 4 It is another step flow chart of obtaining a calibrated depth map in an embodiment of the present invention;
[0037] Figure 5 It is a schematic diagram of a structured light multi-scale depth prediction network framework in an embodiment of the present invention; Figure 6 It is a schematic structural diagram of a multi-scale information fusion system in an embodiment of the present invention; Figure 7 It is a schematic structural diagram of a multi-scale information fusion device in an embodiment of the present invention; and Figure 8 It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.
[0039] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0040] A multi-scale information fusion structured light camera provided by an embodiment of the present invention aims to solve the problems existing in the prior art.
[0041] The technical solution of the present invention and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.
[0042] Due to the reason of the structured light depth calculation algorithm, the structured light camera has a lower detection distance limit related to its resolution. For a target object with a depth less than the lower detection distance limit, its depth value cannot be calculated, even if the target object appears in the structured light image. The applicant found that the lower detection distance limit is related to the resolution of the structured light image. When the resolution of the structured light image decreases, the lower detection distance limit also decreases, so that the depth value at a closer distance can be obtained.
[0043] Based on the structured light camera, the present invention downsamples the first structured light image to obtain the second structured light image, respectively calculates the first depth map and the second depth map using parallax, and then performs fusion according to the lower detection distance limit, ensuring the detection accuracy of depths at different distances during depth reconstruction and breaking through the lower detection distance limit of the traditional structured light scheme.
[0044] Figure 1 It is a schematic structural diagram of a multi-scale information fusion structured light camera in an embodiment of the present invention. As Figure 1 shown, a multi-scale information fusion structured light camera in an embodiment of the present invention includes:
[0045] A structured light projector for projecting structured light.
[0046] Specifically, the structured light projector is an important component of the entire camera system. Its main function is to project structured light of a specific pattern onto the target object. This structured light can be stripe-shaped, grid-shaped, or other light patterns with specific coding. By projecting the structured light, it provides the basic visual features for subsequent depth information acquisition.
[0047] Inside the structured light projector, there are usually a light source (such as a laser diode, LED, etc.) and optical elements (such as a diffraction grating, projector lens, etc.). The light emitted by the light source is modulated and shaped by the optical elements to form a light pattern with a specific structure and distribution, and then projected onto the surface of the target object. When the structured light irradiates the object surface, it will be deformed due to factors such as the shape and distance of the object surface. Subsequently, by analyzing the deformed structured light, the depth information of the object can be obtained.
[0048] The structured light projector determines the pattern and quality of the projected light, directly affecting the accuracy and reliability of subsequent depth calculations. Different structured light patterns are suitable for different application scenarios. For example, for high-precision three-dimensional measurements, a more refined stripe structured light may be required; while for some scenarios with high real-time requirements, a simple grid-shaped structured light may be more appropriate.
[0049] The structured light receiver is used to receive the reflected signal of the structured light.
[0050] Specifically, the main task of the structured light receiver is to receive the structured light signal reflected after being projected onto the target object by the structured light projector. It converts the reflected light signal into an electrical signal or a digital signal for subsequent processing and analysis.
[0051] The structured light receiver usually consists of an image sensor (such as a CCD or CMOS sensor) and related optical lenses. The optical lens focuses the reflected structured light onto the image sensor. The image sensor converts the light signal into an electrical signal, and after processing such as analog-to-digital conversion, it outputs a digital image signal. These digital image signals contain the deformation information of the structured light after reflection on the object surface and are an important basis for calculating the depth of the object.
[0052] The performance of the structured light receiver (such as resolution, sensitivity, etc.) has a great impact on the acquisition of depth information. A high-resolution image sensor can capture more subtle deformations of the structured light, thereby improving the accuracy of depth calculation; while a high-sensitivity sensor can work properly in a darker environment, expanding the applicable range of the camera.
[0053] A processor is configured to generate a first structured light image based on the reflected signal, downsample the first structured light image to obtain a second structured light image; calculate the disparity between the first structured light image and a first reference image to obtain a first depth map, and calculate the disparity between the second structured light image and a second reference image to obtain a second depth map; fuse the first depth map and the second depth map according to a detection distance lower limit to obtain a corrected depth map.
[0054] Specifically, the processor is the core control and data processing unit of the entire structured light camera. It is responsible for processing the signals output by the structured light receiver, generating a structured light image, and further calculating depth information to ultimately achieve the three-dimensional reconstruction of the target object.
[0055] The processor receives the digital signal output by the structured light receiver, and after image preprocessing (such as denoising, enhancement, etc.), generates a first structured light image. This image contains the complete information of the structured light reflected on the object surface. A downsampling operation is performed on the first structured light image to reduce the image resolution and obtain a second structured light image. The purpose of downsampling is to reduce the data volume, improve the speed of subsequent processing, and to a certain extent, highlight the low-frequency information of the image, which is suitable for processing distant objects.
[0056] The first structured light image is compared with the first reference image, and the first depth map is obtained by calculating the disparity (i.e., the position difference of the same object point in different images). Similarly, the second depth map is obtained by calculating the disparity between the second structured light image and the second reference image. Disparity calculation is based on the principle of triangulation. By using known camera parameters (such as focal length, baseline distance, etc.) and disparity information, the depth of each point on the object surface can be calculated.
[0057] The first depth map and the second depth map are fused according to the detection distance lower limit to obtain a corrected depth map. The corrected depth map has a wider range of depth values and more accurate depth values.
[0058] The performance of the processor (such as computing speed, processing ability, etc.) directly affects the real-time performance and accuracy of the entire camera system. Efficient algorithms and powerful computing capabilities can quickly and accurately process a large amount of image data to achieve real-time three-dimensional reconstruction of the target object.
[0059] Figure 2 This is a flowchart of the steps of a processor in the embodiment of the present invention during processing. As Figure 2 shown, the steps of a processor in the embodiment of the present invention during processing include:
[0060] Step S1: Generate a first structured light image based on the reflected signal; downsample the structured light image to obtain a second structured light image.
[0061] In this step, the main purpose is to convert the reflected signal received by the structured light receiver into an image for subsequent processing, and obtain structured light images with different resolutions through downsampling operations.
[0062] After the processor receives the reflected signals from the structured light receiver, it performs a series of operations on these signals, such as signal amplification, filtering, denoising, etc., to remove interference and noise in the signals. Then, the processed signals are converted into digital image data to form the first structured light image. This image is an original structured light image with a high resolution and contains rich detailed information.
[0063] To obtain information at different scales, the processor performs downsampling operations on the first structured light image. Downsampling refers to reducing the resolution of an image by decreasing the number of pixels in the image. Common downsampling methods include mean downsampling, max downsampling, or bilinear interpolation, etc.
[0064] Step S2: Calculate the disparity between the first structured light image and the first reference image to obtain the first depth map, and calculate the disparity between the second structured light image and the second reference image to obtain the second depth map.
[0065] In this step, based on the first structured light image and the second structured light image obtained in step S1, the depth information of the target object is obtained by calculating the disparity with the corresponding reference images, and the first depth map and the second depth map are generated.
[0066] Match and compare the first structured light image with the first reference image, and calculate the disparity between them. Disparity refers to the position difference of the same object point in different images, and it is inversely proportional to the distance from the object to the camera. By using known camera parameters (such as focal length, baseline distance, etc.) and the calculated disparity information, the depth values of each point on the object surface can be calculated using the principle of triangulation, and then the first depth map is generated. Since the first structured light image has a high resolution, the first depth map can usually provide more detailed depth information of nearby objects. The first structured light image and the first reference image have the same resolution.
[0067] Similarly, calculate the disparity between the second structured light image and the second reference image to obtain the second depth map. Since the second structured light image has been downsampled, its resolution is low. The second structured light image and the second reference image have the same resolution.
[0068] Step S3: Fuse the first depth map and the second depth map according to the lower limit of the detection distance to obtain a corrected depth map with high accuracy and taking into account different depth distances.
[0069] In this step, there is a lower limit for the detection distance, that is, the closest distance that can be accurately detected. For the part of the depth map with a depth value less than the lower limit of the detection distance, the data may be inaccurate due to measurement errors or other reasons and needs to be processed. Depth maps with different resolutions have different lower limits for the detection distance. Therefore, in this step, by combining images with different resolutions and using different lower limits for the detection distance, the two depth maps are fused to give full play to the advantages of both, and a calibrated depth map with high accuracy and covering different depth distances is obtained.
[0070] Step S4: Use a feature encoder to fuse the first structured light image and the calibrated depth map to obtain a fused feature.
[0071] In this step, a feature encoder is used to extract and fuse features from the first structured light image and the calibrated depth map, and integrate the information of both into a more representative fused feature. The feature encoder is usually a deep neural network model that can automatically learn the feature information in the image and the depth map. First, the first structured light image and the calibrated depth map are input into the feature encoder, and the encoder will perform operations such as convolution and pooling on them to extract their respective feature representations. Then, these two feature representations are fused through a specific fusion method (such as concatenation, weighted summation, etc.) to obtain a fused feature. The fused feature contains information such as the texture and color of the structured light image and the depth information of the calibrated depth map, providing a richer information basis for the subsequent generation of high-resolution depth maps.
[0072] Step S5: Decode the fused feature through a feature decoder to obtain a super-resolved high-resolution depth map.
[0073] In this step, a decoding operation is performed on the fused feature obtained in step S4 through a feature decoder to restore and increase the resolution of the depth map, and finally a super-resolved high-resolution depth map is obtained. The feature decoder is also a deep neural network model, which corresponds to the feature encoder. The decoder receives the fused feature as input and gradually restores and enlarges the size of the feature map through operations such as deconvolution and upsampling, converting the low-resolution fused feature into a high-resolution depth map. During the decoding process, the decoder will learn how to reconstruct accurate depth information from the fused feature, thereby generating a depth map with higher resolution and richer details to meet the requirements of high-precision three-dimensional reconstruction of the target object.
[0074] Figure 3 This is a flowchart of the steps for obtaining a calibrated depth map in an embodiment of the present invention. As Figure 3 shown, the steps for obtaining a calibrated depth map in an embodiment of the present invention include:
[0075] Step S41: Obtain the first lower detection distance limit of the first depth map and the second lower detection distance limit of the second depth map.
[0076] In this step, the first lower detection distance limit is greater than the second lower detection distance limit. For a structured light camera, when its internal parameters and external parameters are fixed, the lower detection distance limit corresponding to a depth map with a corresponding resolution is fixed. That is, the first lower detection distance limit of the first resolution depth map is known, and the second lower detection distance limit of the second resolution depth map is also known.
[0077] Step S42: For the pixel points whose depth values are less than the first lower detection distance limit and greater than the second lower detection distance limit, use the depth values of the second resolution depth values to obtain a corrected depth map.
[0078] In this step, for a structured light camera, the depth range can be divided into three intervals according to the first lower detection distance limit and the second lower detection distance limit: greater than the first lower detection distance limit, less than the first lower detection distance limit and greater than the second lower detection distance limit, and less than the second lower detection distance limit.
[0079] For the pixel points that are less than the first lower detection distance limit and greater than the second lower detection distance limit, use the depth values corresponding to the second resolution depth map. At this time, the depth values in the first resolution depth map are empty or incorrect depth values. Directly using the depth values corresponding to the second resolution depth map can obtain accurate depth values. Since the first resolution is higher than the second resolution, for one pixel point of the second resolution, it corresponds to two or more points of the first resolution. When modifying the depth values, the depth values of at least two pixel points of the first resolution will be the same.
[0080] For the pixel points that are less than the second lower detection distance limit, there are no reliable depth values.
[0081] Figure 4 This is the flowchart of another step for obtaining a corrected depth map in the embodiment of the present invention. As Figure 4 shown, another step for obtaining a corrected depth map in the embodiment of the present invention further includes between step S41 and step S42:
[0082] Step S43: For the pixel points whose depth values are greater than the first lower detection distance limit, perform weighted fusion on the pixel points corresponding to the first depth map and the second depth map.
[0083] In this step, for pixel points greater than the lower limit of the first detection distance, weighted fusion is performed based on the depth values corresponding to the first-resolution depth map and the second-resolution depth map. The fusion method can be freely selected. For example, different weights can be assigned according to the change of the depth value. The greater the depth value, the greater the weight of the first-resolution depth map; the smaller the depth value, the smaller the weight of the first-resolution depth map. Another example is to set the weight of the first-resolution depth map to 1 and the weight of the second-resolution depth map to 0, that is, all the depth data of the first-resolution depth map is adopted.
[0084] In some embodiments, in step S4, the feature encoder includes a plurality of convolutional layers and pooling layers. The convolutional layers are used to extract the features of the first structured light image and the corrected depth map, and the pooling layers are used to downsample the extracted features to generate the fused features. The feature encoder plays a key role in the data processing flow of the entire multi-scale information fusion structured light camera. Its main task is to extract valuable features from the first structured light image and the corrected depth map, and integrate these features to generate fused features. The convolutional layers and pooling layers are the basic components that make up the feature encoder, and they work together to complete the tasks of feature extraction and downsampling.
[0085] The convolutional layer is the core part of the feature encoder for extracting the features of images and depth maps. It can automatically learn the local feature patterns in the input data, such as edge, texture, shape and other information. For the first structured light image, the convolutional layer can extract the texture and color features of the object surface; for the corrected depth map, the convolutional layer can capture the geometric shape and depth change features of the object.
[0086] The convolutional layer performs a sliding convolution operation on the input data through a set of learnable convolutional kernels (also called filters). Each convolutional kernel has specific weight parameters. During the convolution process, the convolutional kernel multiplies element by element with the local area of the input data and sums them up to obtain a convolutional output value. By sliding the convolutional kernel over the entire input data, a feature map can be obtained. Different convolutional kernels can extract different types of features. Therefore, multiple convolutional kernels are usually used in the convolutional layer to generate multiple feature maps. These feature maps contain the feature information of the input data at different scales and directions.
[0087] The pooling layer is mainly used to downsample the features extracted by the convolutional layer. The purpose of downsampling is to reduce the size of the feature map, reduce the data volume, and at the same time retain the main feature information in the feature map. By downsampling, the complexity of subsequent calculations can be reduced, the computational efficiency of the model can be improved, and to a certain extent, the robustness of the model can be enhanced, reducing the risk of overfitting.
[0088] Common pooling operations include max pooling and average pooling. Taking max pooling as an example, the pooling layer divides the input feature map into several non-overlapping regions (usually rectangular regions), and then selects the maximum value in each region as the output value of that region. In this way, the pooling layer can retain the most significant feature information in the feature map while reducing the size of the feature map. The pooling operation is usually performed after the convolutional layer to downsample the feature map output by the convolutional layer.
[0089] The process of generating the fused feature is as follows:
[0090] First, the first structured light image and the calibrated depth map are respectively input into the convolutional layer of the feature encoder. The convolutional layer extracts features from these two input data and generates corresponding feature maps respectively. These feature maps contain the local feature information of the first structured light image and the calibrated depth map.
[0091] The feature maps output by the convolutional layer are input into the pooling layer for downsampling. The pooling layer performs a dimensionality reduction operation on the feature maps, reducing the size and data volume of the feature maps. After being processed by the pooling layer, the obtained feature maps not only retain the main feature information but also have a lower dimension, which is convenient for subsequent processing and fusion.
[0092] Finally, the feature maps of the first structured light image and the calibrated depth map after being processed by the pooling layer are fused. Common fusion methods include concatenation and weighted summation, etc. For example, in concatenation fusion, the two feature maps are concatenated in the channel dimension to form a new feature map, and this new feature map is the fused feature. The fused feature contains the comprehensive feature information of the first structured light image and the calibrated depth map, providing a richer and more representative input for subsequent feature decoding and high-resolution depth map generation.
[0093] In summary, through the collaborative work of the convolutional layer and the pooling layer, the feature encoder effectively extracts the features of the first structured light image and the calibrated depth map and generates the fused feature, providing an important intermediate result for the depth map super-resolution processing of the entire multi-scale information fusion structured light camera.
[0094] In some embodiments, in step S5, the feature decoder includes multiple upsampling layers and convolutional layers. The upsampling layer is used to upsample the fused feature, and the convolutional layer is used to refine the upsampled feature, and finally obtain the super-resolved high-resolution depth map. In the data processing flow of the entire multi-scale information fusion structured light camera, the feature decoder undertakes the important task of decoding the fused feature generated by the feature encoder, thereby restoring and improving the resolution of the depth map. The upsampling layer and the convolutional layer are the core components of the feature decoder, and they cooperate with each other to gradually achieve the amplification and refinement of the features, and finally generate the high-resolution depth map.
[0095] The main function of the upsampling layer is to perform an amplification operation on the fused features, increase the size of the feature map, and gradually make it approach the size of the target high-resolution depth map. Since the pooling layer in the feature encoder downsamples the features, resulting in a reduction in the size of the feature map and a certain degree of compression of information. Therefore, the role of the upsampling layer is to restore and expand this compressed feature information, laying the foundation for the subsequent generation of a high-resolution depth map.
[0096] Common upsampling methods include bilinear interpolation, nearest neighbor interpolation, and transposed convolution (also known as deconvolution), etc.
[0097] Bilinear interpolation: It is an interpolation method based on weighted averaging of the surrounding four pixel values. For the new pixel points that need to be upsampled, according to their relative positions in the original feature map, calculate the weights of the surrounding four pixels, and then perform weighted summation to obtain the value of the new pixel point. This method can generate relatively smooth upsampling results.
[0098] Nearest neighbor interpolation: This method directly selects the pixel value in the original feature map that is closest to the new pixel point as the value of the new pixel point. This method is computationally simple, but may cause the upsampled image to have jagged edges.
[0099] Transposed convolution: Transposed convolution is a learnable upsampling method. It performs a convolution operation on the input feature map through a set of trainable convolution kernels, thereby achieving an increase in the size of the feature map. Transposed convolution can automatically learn the optimal upsampling method according to the data, so it is widely used in many deep learning models.
[0100] The convolution layer refines the features after the upsampling layer. Although the upsampling operation increases the size of the feature map, it may introduce some blurred or inaccurate information. The convolution layer can further extract and enhance the feature information by performing a convolution operation on these upsampled features, remove the noise and artifacts generated during the upsampling process, make the features clearer and more accurate, thereby improving the quality of the finally generated high-resolution depth map.
[0101] The working principle of the convolution layer is similar to that of the convolution layer in the feature encoder. It performs a sliding convolution operation on the input upsampled feature map through a set of trainable convolution kernels. Each convolution kernel multiplies element by element with the local area of the input feature map and sums them to obtain a convolution output value. By sliding the convolution kernel over the entire input feature map, a new feature map can be obtained. Different convolution kernels can extract different types of features. Through the combination of multiple convolution kernels, the features can be refined and enhanced in multiple dimensions.
[0102] The process of generating the super-resolution high-resolution depth map is as follows:
[0103] The fused features generated by the feature encoder are input into the upsampling layer of the feature decoder. The upsampling layer performs an amplification operation on the fused features according to the selected upsampling method (such as bilinear interpolation, transposed convolution, etc.), gradually increasing the size of the feature map. Usually, the feature decoder will contain multiple upsampling layers. Through multiple upsamplings, the size of the feature map gradually approaches the size of the target high-resolution depth map.
[0104] The feature map processed by the upsampling layer is input into the convolutional layer for refinement. The convolutional layer performs a convolution operation on these upsampled features, extracts and enhances the feature information, and removes possible noise and artifacts. In the feature decoder, there are usually multiple convolutional layers, which can continuously refine and optimize the features, making the features more accurate and clear.
[0105] After multiple upsampling and convolution operations, the feature decoder finally outputs a high-resolution depth map. This depth map has a higher resolution and richer detail information, and can more accurately reflect the three-dimensional structure of the target object. Compared with the corrected depth map, the super-resolved high-resolution depth map has significant improvements in accuracy and clarity, meeting the requirements of multi-scale information fusion structured light cameras in application scenarios such as high-precision three-dimensional measurement and object recognition.
[0106] In summary, through the collaborative work of the upsampling layer and the convolutional layer, the feature decoder effectively converts the fused features into a super-resolved high-resolution depth map, providing high-quality output results for the depth information processing of multi-scale information fusion structured light cameras.
[0107] This specification also provides a multi-scale information fusion structured light camera, which is characterized by including:
[0108] A structured light projector for projecting structured light;
[0109] A structured light receiver for receiving the reflected signal of the structured light;
[0110] A processor for generating a first structured light image according to the reflected signal, downsampling the first structured light image to obtain a second structured light image, and downsampling the second structured light image to obtain a third structured light image; calculating the disparity between the first structured light image and a first reference image to obtain a first depth map, calculating the disparity between the second structured light image and a second reference image to obtain a second depth map, and calculating the disparity between the third structured light image and a third reference image to obtain a third depth map; fusing the first depth map, the second depth map, and the third depth map according to the detection distance lower limit to obtain a corrected depth map.
[0111] Compared with the foregoing embodiments, in this embodiment, on the basis of downsampling the first structured light image, the second structured light image is also downsampled, thereby generating three structured light images with different resolutions. Therefore, three depth maps with different resolutions can be finally generated, and depth maps with a closer depth range can be obtained.
[0112] The first depth map corresponds to the first detection distance lower limit, the second depth map corresponds to the second detection distance lower limit, and the third depth map corresponds to the third detection distance lower limit. Since the resolution of the first depth map is greater than that of the second depth map, and the resolution of the second depth map is greater than that of the third depth map, the first detection distance lower limit is greater than the second detection distance lower limit, and the second detection distance lower limit is greater than the third detection distance lower limit. According to the first detection distance lower limit, the second detection distance lower limit, and the third detection distance lower limit, the depth range can be divided into four parts:
[0113] (1) The depth range greater than the first detection distance lower limit.
[0114] (2) The depth range less than the first detection distance lower limit and greater than the second detection distance lower limit.
[0115] (3) The depth range less than the second detection distance lower limit and greater than the third detection distance lower limit.
[0116] (4) The depth range less than the third detection distance lower limit.
[0117] When fusing the depth maps, various strategies can be adopted. For example:
[0118] For (1) the depth range greater than the first detection distance lower limit, directly use the value of the first depth map.
[0119] For (2) the depth range less than the first detection distance lower limit and greater than the second detection distance lower limit, directly use the value of the second depth map.
[0120] For (3) the depth range less than the second detection distance lower limit and greater than the third detection distance lower limit, directly use the value of the third depth map.
[0121] For (4) the depth range less than the third detection distance lower limit, set the depth value to 0.
[0122] For another example:
[0123] For (1) the depth range greater than the first detection distance lower limit, directly use the value of the first depth map.
[0124] For (2) the depth range less than the first detection distance lower limit and greater than the second detection distance lower limit, weight-average the depth values of the second depth map and the third depth map.
[0125] For the depth range where (3) is less than the lower limit of the second detection distance and greater than the lower limit of the third detection distance, directly use the value of the third depth map.
[0126] For the depth range where (4) is less than the lower limit of the third detection distance, set the depth value to 0.
[0127] Figure 5 It is a schematic diagram of a structured light multi-scale depth prediction network framework in an embodiment of the present invention. Figure 5 Three-level downsampling is adopted to obtain the first structured light images at scales of 1 / 2, 1 / 4, and 1 / 8 respectively. The following will be described in detail in conjunction with Figure 5 for elaboration.
[0128] Downsample the first structured light image obtained by the structured light camera to obtain the first structured light images at scales of 1 / 2, 1 / 4, and 1 / 8;
[0129] Use the SGM algorithm with the same 128 disparities to process the first structured light images at resolutions of 1 / 2, 1 / 4, and 1 / 8 respectively to quickly obtain the depth maps of each resolution;
[0130] Calculate the lower limit of the depth limit of each resolution according to the minimum detection depth formula. The minimum detection depth calculation formula is as follows:
[0131] (focallength / scale*baseline) / Maxdisparity
[0132] Among them, focallength is the camera focal length, scale is the downsampling scale, baseline is the baseline length of the structured light camera, and Maxdisparity is the maximum disparity of structured light matching. Example: When the original image resolution is 1280*800, the focal length is 611, the maximum disparity is 128, and the baseline distance is 100: the lower limit of the limit depth at the 1 / 2 scale is about 24 cm, the lower limit of the limit depth at the 1 / 4 scale is about 12 cm, and the lower limit of the limit depth at the 1 / 8 scale is about 6 cm.
[0133] Upsample the 1 / 8 resolution depth map by 2 times, calculate the obtained lower limit of the limit depth, take the pixels in the 1 / 8 resolution depth map where the depth is between [6 cm, 12 cm] to generate a mask range, and replace the depth in the 1 / 4 resolution depth map with the depth in the 1 / 8 resolution depth map according to this mask to obtain a synthesized 1 / 4 resolution depth map.
[0134] The synthesized quarter-resolution depth map is further upsampled by 2 times. Pixels with depths in the range of [6 cm, 24 cm] in the quarter-resolution synthesized depth map are used to generate a mask range. According to this mask, the depths in the half-resolution depth map are replaced with the depths in the synthesized quarter-resolution depth map, obtaining a synthesized half-resolution depth map. The lower limit of the effective limit detection distance of this synthesized half-resolution depth map reaches 6 cm, while retaining the accuracy of other long-distance depths.
[0135] Using Figure 5 the feature encoder shown in the figure to perform fusion processing on the first structured light image and the synthesized depth map obtained by the structured light camera, and inferring the fusion features of the first structured light image and the initial depth;
[0136] The fusion features are decoded by a feature decoder to obtain a corrected and super-resolved high-precision and high-resolution depth map.
[0137] This specification also provides an electronic device, including the multi-scale information fusion structured light camera in any of the foregoing embodiments. It should be noted that this embodiment is only an exemplary illustration for those skilled in the art to better understand the role of the multi-scale information fusion structured light camera in the electronic device, and should not constitute any limitation to the protection scope of the present invention.
[0138] In this embodiment, the electronic device mainly consists of two major parts: a mobile device and a multi-scale information fusion structured light camera:
[0139] Mobile device: It is the carrier of the electronic device and provides the mobile ability for the multi-scale information fusion structured light camera. The mobile device can be in various forms, such as a robot chassis, a drone platform, or a mobile part of a wearable device, etc. Taking the robot chassis as an example, it usually includes a power system (such as motors, batteries, etc.), a transmission system (such as gears, chains, etc.), and a control system (such as a microcontroller, sensors, etc.), and can move autonomously according to a preset path or based on environmental feedback.
[0140] Multi-scale information fusion structured light camera: It consists of a structured light projector, a structured light receiver, and a processor. The structured light projector is used to project a specific pattern of structured light onto the target object; the structured light receiver receives the reflected structured light signal; the processor processes the reflected signal to generate structured light images and depth maps with different resolutions, and performs fusion and super-resolution processing to finally obtain a high-precision and high-resolution depth map.
[0141] Functions of the electronic device
[0142] The functions of the electronic device depend on its application scenarios. The following are several common functions:
[0143] 3D Modeling: During movement, an electronic device can use a multi-scale information fusion structured light camera to perform an all-round scan of the surrounding environment or target object. By acquiring depth maps and structured light images from different angles and through processing and stitching, a 3D model of the target object or environment can be constructed. This has important application value in fields such as industrial design, cultural relic protection, and architectural surveying and mapping.
[0144] Object Recognition and Detection: With the high-precision depth information and rich texture features provided by the multi-scale information fusion structured light camera, an electronic device can identify and detect target objects. For example, in a logistics warehouse, the electronic device can identify the types, locations, and sizes of goods to achieve automated goods sorting and management.
[0145] Navigation and Obstacle Avoidance: During the movement of an electronic device, the multi-scale information fusion structured light camera can continuously sense the depth information of the surrounding environment. By analyzing this information, the electronic device can identify the positions and distances of obstacles and plan a safe movement path to achieve autonomous navigation and obstacle avoidance functions. This has broad application prospects in fields such as intelligent robots and autonomous driving vehicles.
[0146] The functions of the multi-scale information fusion structured light camera are as follows:
[0147] Providing High-Precision Depth Information: By projecting structured light and analyzing the reflected signals, the multi-scale information fusion structured light camera can accurately calculate the depth information of each point on the surface of the target object. Using the method of multi-scale information fusion and combining depth maps of different resolutions improves the accuracy and reliability of the depth information, especially having obvious advantages when dealing with complex scenes and objects at different distances.
[0148] Enhancing Environmental Perception Ability: The camera can not only acquire depth information but also provide feature information such as the texture and color of the target object through the structured light image. These rich information helps the electronic device to more comprehensively sense the surrounding environment, improve the ability to identify and understand objects, and thus better complete various tasks.
[0149] Supporting Real-Time Processing: The processor of the camera has efficient data processing capabilities and can complete operations such as the generation of structured light images, the calculation, fusion, and super-resolution of depth maps in a short time. This enables the electronic device to obtain and process environmental information in real time, meeting the requirements of application scenarios such as real-time navigation and real-time monitoring.
[0150] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0151] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention.
Claims
1. A multi-scale information fusion structured light camera, characterized in that: include: A structured light projector, used for projecting structured light; A structured light receiver, used for receiving a reflection signal of the structured light; A processor, configured to generate a first structured light image according to the reflection signal, and downsample the first structured light image to obtain a second structured light image; The first structured light image and the first reference image are used to calculate the parallax to obtain a first depth map, and the second structured light image and the second reference image are used to calculate the parallax to obtain a second depth map; the first depth map and the second depth map are fused according to the lower limit of the detection distance to obtain a corrected depth map.
2. The multi-scale information fusion structured light camera according to claim 1, characterized in that: The processor includes: Step S1: generating a first structured light image according to the reflection signal; downsampling the structured light image to obtain a second structured light image; Step S2: calculating the parallax between the first structured light image and the first reference image to obtain a first depth map, and calculating the parallax between the second structured light image and the second reference image to obtain a second depth map; Step S3: fusing the first depth map and the second depth map according to the lower limit of the detection distance to obtain a corrected depth map with high accuracy and taking into account different depth distances; Step S4: using a feature encoder to fuse the first structured light image and the corrected depth map to obtain a fusion feature; Step S5: Decode the fused features through a feature decoder to obtain a high-resolution depth map after super-resolution.
3. The multi-scale information fusion structured light camera according to claim 2, characterized in that: In the step S1, the structured light image is downsampled to obtain a second structured light image, and the downsampling method adopted is one of mean downsampling, maximum downsampling or bilinear interpolation downsampling.
4. The multi-scale information fusion structured light camera according to claim 2, characterized in that: The first structured light image has the same resolution as the first reference image; the second structured light image has the same resolution as the second reference image.
5. The multi-scale information fusion structured light camera according to claim 2, characterized in that: Step S3 includes the following steps: Step S41: obtaining a first detection distance lower limit of the first depth map, and obtaining a second detection distance lower limit of the second depth map; Step S42: for pixel points whose depth values are less than the first detection distance lower limit and greater than the second detection distance lower limit, the depth values of the second resolution depth value are used to obtain a corrected depth map.
6. The multi-scale information fusion structured light camera according to claim 5, characterized in that: The steps between step S41 and step S42 also include: Step S43: for pixel points whose depth values are greater than the first detection distance lower limit, weighted fusion is performed on pixel points corresponding to the first depth map and the second depth map.
7. The multi-scale information fusion structured light camera according to claim 2, characterized in that: In step S4, the feature encoder includes multiple convolution layers and pooling layers, the convolution layers are used to extract features of the first structured light image and the corrected depth map, and the pooling layers are used to downsample the extracted features to generate the fused features.
8. The multi-scale information fusion structured light camera according to claim 2, characterized in that: In step S5, the feature decoder includes multiple upsampling layers and convolution layers, the upsampling layers are used to upsample the fused features, and the convolution layers are used to refine the upsampled features, and finally obtain the high-resolution depth map after super-resolution.
9. A multi-scale information fusion structured light camera, characterized in that: include: A structured light projector, used for projecting structured light; A structured light receiver, used for receiving a reflection signal of the structured light; A processor is used to generate a first structured light image according to the reflection signal, downsample the first structured light image to obtain a second structured light image, downsample the second structured light image to obtain a third structured light image; calculate the parallax between the first structured light image and the first reference image to obtain a first depth map, calculate the parallax between the second structured light image and the second reference image to obtain a second depth map, and calculate the parallax between the third structured light image and the third reference image to obtain a third depth map; and fuse the first depth map, the second depth map and the third depth map according to a lower limit of the detection distance to obtain a corrected depth map.
10. An electronic device, characterized in that: A multi-scale information fusion structured light camera comprising any one of claims 1 to 9.