A depth map generation system, method, and autonomous mobile device

Through the baseline fusion method of multi-projector and sensor, the problems of high system complexity and cost in the prior art are solved, and the accuracy and real-time improvement of deep calculations are achieved.

CN114627174BActive Publication Date: 2025-08-01HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210328464.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-08-01
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

In the prior art, in order to improve the depth calculation accuracy, multiple sensors and projectors need to be arranged, resulting in high system complexity and high chip requirements and increased costs.

Method used

Multiple projectors and sensors are used to calculate and fusion the depth map through different baselines. Multiple baselines with different distances between multiple projectors and sensors are used to generate and fusion the depth map respectively to improve the accuracy of depth calculation.

Benefits of technology

Through the baseline fusion method of multi-projector and sensor, the system structure is simplified, the chip requirements are reduced, the accuracy of depth calculations and the real-time system are improved, and the cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627174B_ABST
    Figure CN114627174B_ABST
Patent Text Reader

Abstract

A depth map generation system, method and autonomous mobile device provided by an embodiment of the present invention are applied to the field of information technology. Multiple projectors project light; a sensor collects an image to be processed; a processor compares the image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed; for any baseline, according to the baseline and the disparity corresponding to each pixel at each speckle position, calculate the depth value of each pixel at each speckle position corresponding to the baseline; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each baseline to obtain a target depth map. Thus, multiple baselines with different distances between multiple projectors and sensors can be used to calculate the corresponding depth maps respectively, and the calculated depth maps are fused, thereby improving the accuracy of depth calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a depth map generation system, method, and autonomous mobile device. Background Art

[0002] Currently, with the popularization of intelligent devices, the application of automatic obstacle avoidance devices has become increasingly widespread. For example, through a sweeping robot, obstacles around can be identified and avoided. When performing obstacle recognition, it is necessary to project speckles through a projector, then collect an image including the speckles through a sensor, and finally calculate the depth using the collected image and generate a depth map, so as to identify and avoid obstacles through the depth map.

[0003] In the prior art, in order to improve the accuracy of depth calculation, it is generally necessary to arrange multiple sensors for collection, utilize the baseline between different sensors and projectors, and measure using a three-dimensional distance measurement method. However, this method not only has high requirements for the chip, but also results in a high system complexity. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a depth map generation system, method, and autonomous mobile device to improve the accuracy of depth calculation. The specific technical solutions are as follows:

[0005] In the first aspect of the embodiments of the present application, first, a depth map generation system is provided, including a sensor, a processor, and multiple projectors;

[0006] The multiple projectors are used to project light;

[0007] The sensor is used to collect an image to be processed, where the image to be processed includes speckles after the light projected by the multiple projectors is reflected;

[0008] The processor is used to compare the image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference position of each pixel at each speckle position; for any baseline, calculate the depth value of each pixel at each speckle position corresponding to the baseline according to the baseline and the disparity corresponding to each pixel at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each baseline to obtain a target depth map.

[0009] Optionally, the processor is specifically configured to calculate the depth values of the pixels at each speckle position for any baseline according to the baseline and the disparities corresponding to the pixels at each speckle position; for any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline according to the preset depth range corresponding to the baseline, so as to obtain the depth values of the pixels at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

[0010] Optionally, the processor is specifically configured to, for any baseline, select the plane equation corresponding to the baseline; calculate the depth of the pixels at each speckle position according to the selected plane equation and the disparities corresponding to the pixels at each speckle position, so as to obtain the depth values of the pixels at each speckle position.

[0011] Optionally, the multiple projectors are specifically configured to alternately project light through each of the projectors;

[0012] The sensor is specifically configured to collect the image to be processed once every preset time period.

[0013] Optionally, the exposure duration of the projector closest to the sensor among the multiple projectors is less than the exposure duration of the projector farthest from the sensor.

[0014] In a second aspect of the embodiments of the present application, a depth map generation method is provided, which is applied to a processor in a depth map generation system. The depth map generation system includes: a sensor, a processor, and multiple projectors. The method includes:

[0015] Obtain an image to be processed, where the image to be processed includes speckles after the light projected by the multiple projectors is reflected;

[0016] Compare the image to be processed with a preset speckle image to obtain the disparities corresponding to the pixels at each speckle position in the image to be processed, where the preset speckle image includes the initial reference positions of the pixels at each speckle position;

[0017] For any baseline, calculate the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the disparities corresponding to the pixels at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different;

[0018] For any baseline, generate a depth map corresponding to the baseline according to the depth values of the pixels at each speckle position passing through the baseline;

[0019] Fuse the depth maps corresponding to each of the baselines to obtain a target depth map.

[0020] Optionally, for any baseline, calculating the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position includes:

[0021] For any baseline, calculating the depth values of the pixels at each of multiple speckle positions according to the baseline and the parallax corresponding to each pixel at each speckle position;

[0022] For any baseline, extracting the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline according to the preset depth range corresponding to the baseline, to obtain the depth values of the pixels at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

[0023] Optionally, for any baseline, calculating the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position includes:

[0024] For any baseline, selecting the plane equation corresponding to the baseline;

[0025] According to the selected plane equation and the parallax corresponding to each pixel at each speckle position, calculating the depth of each pixel at each speckle position to obtain the depth values of the pixels at each speckle position.

[0026] In a third aspect of the embodiments of the present application, an autonomous mobile device is provided, including a depth map generation system, and the depth map generation system includes: a sensor, a processor, and multiple projectors;

[0027] The multiple projectors are used to project light;

[0028] The sensor is used to collect an image to be processed, where the image to be processed includes speckles after the light projected by the multiple projectors is reflected;

[0029] The processor is used to compare the image to be processed with a preset speckle image to obtain the parallax corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference positions of the pixels at each speckle position; for any baseline, calculating the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different; for any baseline, generating a depth map corresponding to the baseline according to the depth values of the pixels at each speckle position corresponding to the baseline; and fusing the depth maps corresponding to each baseline to obtain a target depth map.

[0030] Optionally, the processor is specifically configured to, for any baseline, calculate the depth values of the pixels at multiple speckle positions according to the baseline and the parallax corresponding to each pixel at each speckle position; for any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline according to the preset depth range corresponding to the baseline, so as to obtain the depth values of the pixels at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

[0031] Optionally, the processor is specifically configured to, for any baseline, select the plane equation corresponding to the baseline; calculate the depth of the pixels at each speckle position according to the selected plane equation and the parallax corresponding to each pixel at each speckle position, so as to obtain the depth values of the pixels at each speckle position.

[0032] Optionally, the multiple projectors are specifically configured to alternately project light through each of the projectors.

[0033] The sensor is specifically configured to collect the image to be processed once every preset time period.

[0034] Optionally, the exposure duration of the projector closest to the sensor among the multiple projectors is less than the exposure duration of the projector farthest from the sensor.

[0035] On the other hand, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the depth map generation method described in any one of the above is implemented.

[0036] An embodiment of the present invention further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the depth map generation method described in any one of the above.

[0037] Beneficial effects of the embodiments of the present invention:

[0038] A depth map generation system, method and autonomous mobile device provided by an embodiment of the present invention. The depth map generation system includes a sensor, a processor and a plurality of projectors; the plurality of projectors are used for projecting light; the sensor is used for collecting an image to be processed, wherein the image to be processed includes speckles after the light projected by the plurality of projectors is reflected; the processor is used for comparing the image to be processed with a preset speckle image to obtain the parallax corresponding to each pixel at each speckle position in the image to be processed, wherein the preset speckle image includes the initial reference position of each pixel at each speckle position; for any baseline, according to the baseline and the parallax corresponding to each pixel at each speckle position, calculate the depth value of each pixel at each speckle position corresponding to the baseline, wherein the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each baseline to obtain a target depth map. Through the depth map generation system of the embodiments of the present application, multiple baselines with different distances between the projector and the sensor can be used to generate corresponding depth maps respectively, and the generated depth maps are fused, thereby improving the accuracy of depth calculation.

[0039] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0041] Figure 1 It is a schematic structural diagram of a depth map generation system provided by an embodiment of the present application;

[0042] Figure 2 It is a schematic diagram of a baseline provided by an embodiment of the present application

[0043] Figure 3 It is the corresponding relationship between the depth ranges of different baselines provided by an embodiment of the present application;

[0044] Figure 4 It is another schematic structural diagram of a depth map generation system provided by an embodiment of the present application;

[0045] Figure 5 It is a schematic flowchart of a depth map generation method provided by an embodiment of the present application;

[0046] Figure 6 Another flowchart of the depth map generation method provided by the embodiment of the present application;

[0047] Figure 7 A schematic structural diagram of an autonomous mobile device provided by the embodiment of the present application;

[0048] Figure 8 A schematic structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the protection scope of the present invention.

[0050] First, the professional terms that may be used in the embodiments of the present application are explained:

[0051] Baseline: The horizontal distance between the centers of the apertures of two cameras or between the center of the camera aperture and the center of the speckle projector is called the baseline of the depth camera.

[0052] Disparity: Binocular disparity is also called stereoscopic disparity. The closer an object is to the image sensor, the greater the difference in the objects captured by the two sensors, which forms binocular disparity. Subsequent algorithms can use the measurement of this disparity to estimate the distance from the object to the camera.

[0053] Speckle projector: A device for projecting speckles on the object to be photographed.

[0054] In the first aspect of the embodiments of the present application, first, a depth map generation system is provided. Refer to Figure 1 , which includes a sensor 102, a processor 101, and a plurality of projectors 103;

[0055] The plurality of projectors 103 are used for projecting light;

[0056] The sensor 102 is used for collecting the image to be processed, where the image to be processed includes the speckles after the light projected by the plurality of projectors is reflected;

[0057] A processor 101 is configured to compare an image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference positions of each pixel at each speckle position; for any baseline, calculate the depth value of each pixel at each speckle position corresponding to the baseline according to the baseline and the disparity corresponding to each pixel at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and sensors are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; and fuse the depth maps corresponding to each baseline to obtain a target depth map.

[0058] Among them, the multiple projectors in the embodiments of the present application may be multiple physically independent projectors. During actual use, the multiple projectors should be installed on the same horizontal line as much as possible to reduce the time loss during calculation. When the projectors are installed on an intelligent mobile device, the multiple projectors and the sensor may be installed on the same side or different sides of the intelligent mobile device. Among them, for each projector among the multiple projectors, the distance from the projector to the sensor is different, so the formed baselines are also different. For example, referring to Figure 2 , if the distances between projector 1, projector 2, projector 3 and the sensor (Sensor) are s1, s2, s3 respectively, then three baselines s1, s2, s3 can be formed.

[0059] After the sensor acquires the image to be processed, the acquired image to be processed can also be preprocessed. Specifically, image features can be increased and brightness equalization can be performed by means such as contrast enhancement, histogram equalization, and binarization.

[0060] Among them, to compare the image to be processed with the preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed, the pixels at each speckle position in the image to be processed acquired by the sensor can be compared with the preset speckle image to obtain the disparity corresponding to each pixel at each speckle position, and a corresponding disparity map can be generated. When calculating the depth of each speckle according to the baseline and each pixel at each speckle position for any baseline to obtain the depth value of each speckle, the corresponding depth value can be calculated according to the generated disparity map and the preset plane equation to generate a depth map. For example, an image matching calculation is performed on the speckle image acquired by the sensor and the reference plane speckle image recorded in the calibration stage to obtain a disparity map; using the above-mentioned disparity map, the plane equation, the camera focal length, and the distance between the speckle and the projector, a depth map is calculated.

[0061] Among them, for any baseline, according to the parallax corresponding to each pixel at the position of each speckle with respect to this baseline, the depth value of each pixel at the position of each speckle corresponding to this baseline can be calculated, and it can be calculated through a variety of preset algorithms. For example, it can be calculated through a monocular depth algorithm. For example, depth calculations are performed pairwise between the Sensor and the projectors; the Sensor and projector 1 use the monocular depth calculation method to calculate the depth map m1 with S1 as the baseline; the Sensor and projector 2 use the monocular depth calculation method to calculate the depth map m2 with S2 as the baseline; the Sensor and projector 3 use the monocular depth calculation method to calculate the depth map m3 with S3 as the baseline. Specifically, the calculation process can be referred to the subsequent embodiments.

[0062] In the embodiments of the present application, the inventors have found through research that when calculating depth values according to different baselines, the best accuracy corresponding to different baselines is often within a certain depth range, as shown in Figure 3 . Therefore, when calculating and generating depth maps according to each baseline, only the depths within a certain distance range are calculated for each baseline, thereby improving the calculation accuracy. Then, by fusing the depth maps generated by each baseline, the target depth map is obtained, so as to obtain a complete depth map through fusion.

[0063] Compared with the prior art, in order to improve the accuracy of depth calculation, by setting multiple sensors, it not only leads to the need to access a large number of image sensors, which places high requirements on the SOC (System on Chip), but also involves multi-SOC synchronization problems when using multiple SOCs, resulting in high system complexity, which is not conducive to the real-time performance and problem-solving ability of the system, and there is also an issue of increased cost. Through the method of the embodiments of the present application, the problem of improving the accuracy of depth calculation can be well solved by using multiple projectors, and generally, the projectors only need PWM (Pulse width modulation) control, with a simple circuit and easy to control.

[0064] It can be seen that through the depth map generation system of the embodiments of the present application, multiple baselines with different distances between multiple projectors and sensors can be utilized to calculate the corresponding depth maps respectively, and the calculated depth maps are fused, thereby improving the accuracy rate of depth calculation.

[0065] Optionally, the processor is specifically configured to, for any baseline, select the plane equation corresponding to this baseline; according to the selected plane equation and the parallax corresponding to each pixel at the position of each speckle, calculate the depth of each pixel at the position of each speckle, and obtain the depth value of each pixel at the position of each speckle.

[0066] During actual use, the internal parameters of the sensor can be pre-calibrated for the correction image. For example, select a reference plane, and determine the plane equation of the reference plane in the camera coordinate system, as well as the baseline distances s1, s2, and s3 between the projector and the sensor through existing camera calibration algorithms. The projector projects a speckle image onto the reference plane, and for each projector, record the reference images R1, R2, and R3 respectively. Calculate the depth intervals [0, d1), [d1, d2), [d2, d3), [d3, +∞) according to the maximum parallax calculation value dmax set by the algorithm. where f is the equivalent focal length after camera correction, thereby realizing the calibration between the sensor and the projector pairwise.

[0067] Among them, when calculating the depth of each speckle through the plane equation and parallax, it is necessary to obtain and calculate the depth corresponding to different parallaxes according to the corresponding relationships between the baseline, parallax, field of view, depth, etc. Therefore, during actual use, a plane equation can be created in advance for different baselines, and this plane equation can represent the corresponding relationship between the parallax and depth corresponding to this baseline. Thus, when calculating the depth, the corresponding plane equation can be selected according to the baseline, and then the corresponding depth can be calculated according to the parallax, thereby improving the calculation efficiency and accuracy.

[0068] Optionally, the processor is specifically configured to, for any baseline, calculate the depth values of each pixel at multiple speckle positions according to the baseline and the parallax corresponding to each pixel at the respective speckle positions; for any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to this baseline, and obtain the depth values of each pixel at the respective speckle positions corresponding to this baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

[0069] For example, depth calculations are performed pairwise between the Sensor and the projector: The Sensor and projector 1 use the monocular depth calculation method to calculate the depth map m1 with S1 as the baseline; The Sensor and projector 2 use the monocular depth calculation method to calculate the depth map m2 with S2 as the baseline; The Sensor and projector 3 use the monocular depth calculation method to calculate the depth map m3 with S3 as the baseline. When performing fusion, the depth map m1 calculated with the minimum baseline s1 can be used as a reference. The final depth map m determines which depth map to select the corresponding value from according to the range of the depth intervals [0, d1), [d1, d2), [d2, d3), [d3, +∞) determined during calibration for each point in m1. The formula is as follows:

[0070]

[0071] Finally, for any baseline, generate a depth map according to the extracted luminance values.

[0072] Optionally, multiple projectors are specifically configured to alternately project light through each projector;

[0073] A sensor is specifically configured to collect a to-be-processed image once every preset time interval.

[0074] Optionally, among the multiple projectors, the exposure duration of the projector closest to the sensor is less than the exposure duration of the projector farthest from the sensor.

[0075] During actual use, the exposure durations of different projectors can be proportional to the corresponding baselines. For example, taking 3 projectors as an example, the corresponding baselines are d1, d2, and d3 respectively. Then, the exposure times t1, t2, and t3 corresponding to the three projectors can be calculated through the following steps:

[0076] If the time of the synthesized depth Figure 1 frame is T, then t1 + t2 + t3 = T;

[0077] t1 / t2 / t3 = d1 / d2 / d3;

[0078] Then, t1 = d1*T / (d1 + d2 + d3);

[0079] t2 = d2*T / (d1 + d2 + d3);

[0080] t3 = d3*T / (d1 + d2 + d3).

[0081] See Figure 4 , the depth map generation system includes an image sensor, multiple projectors, an image signal processing module, a depth image processing module, a digital signal processing module, and a projector control module.

[0082] The projector control module is used to control multiple projectors to alternately encode the space. Each projector can respectively correspond to an image signal processor. The image encoded by projector 1 is input to image signal processing module 1, the image encoded by projector 2 is input to image signal processing module 2, and the image encoded by projector 3 is input to image signal processing module 3. Image signal processing module 1, image signal processing module 2, and image signal processing module 3 can be 3 physically independent modules or the same module for time-division multiplexing. Similarly, DPU (depth image processing module) 1, DPU2, DPU3, digital signal processor 1, and digital signal processor 2 can also be multiple physically independent modules or the same module for time-division multiplexing.

[0083] For the process of 3-projector image acquisition, taking the output frame rate of 30 FPS as an example, the total duration of obtaining 3 projector image frames is 1 / 30 second, and the average acquisition time for one frame is 1 / 90 second. Considering that the projection energy of the projector becomes smaller as the distance increases, in order to make the energy more uniform at different distances and obtain a higher-quality depth map, the exposure duration of the short-baseline projector image frame is less than that of the long-baseline projector image frame. For example, the exposure duration of projector image frame 1 is 5 ms, the exposure duration of projector image frame 2 is 10 ms, and the exposure duration of projector image frame 3 is 15 ms.

[0084] Among them, in the embodiments of the present application, a multi-view system can be formed by multiple projectors and sensors. Using this multi-view system, in the case of lack of texture, the projector can illuminate the object with infrared structured light, and project an artificial texture on the surface of the object. Then, a dual-mode depth acquisition method is used to obtain a depth map with higher quality and stronger reliability. Generally, the distances s1 and s2 between the projector and the camera are less than the distance s3 between the two cameras. In an outdoor high-brightness environment, at a short distance, the structured light intensity is high, and the monocular mode can also be used for acquisition. At a long distance, due to the attenuation of the structured light, the binocular depth calculation method is directly used to obtain a fused depth map, and the quality is also higher compared with the traditional structured light or binocular scheme.

[0085] To illustrate the solution of the present application, see Figure 5 The following is an illustration with specific embodiments:

[0086] Depth map generation method:

[0087] I. Offline processing:

[0088] Offline calibration is performed pairwise between the sensor and the projector: Select a reference plane, and determine the plane equation F of the reference plane in the camera coordinate system, as well as the baseline distances s1, s2, and s3 between the projector and the sensor through existing camera calibration algorithms.

[0089] The projector projects a speckle image onto the reference plane, and for each projector, the reference images R1, R2, and R3 are respectively recorded; according to the maximum disparity calculation value dmax set by the algorithm, the depth intervals [0, d1), [d1, d2), [d2, d3), [d3, +∞) are calculated, where f is the equivalent focal length after camera calibration.

[0090] II. Online processing:

[0091] Step 1: The laser speckle projector performs spatial encoding.

[0092] Step 2: The sensor collects and preprocesses the image. The preprocessing methods include contrast enhancement, histogram equalization, binarization, etc. to increase image features and perform brightness equalization.

[0093] Step 3: Perform depth calculation pairwise between the sensor and the projectors: The sensor and projector 1 use the monocular depth calculation method to calculate the depth map m1 with S1 as the baseline, the sensor and projector 2 use the monocular depth calculation method to calculate the depth map m2 with S2 as the baseline, and the sensor and projector 3 use the monocular depth calculation method to calculate the depth map m3 with S3 as the baseline.

[0094] The monocular depth calculation method is as follows: Calibrate the image using the offline calibrated camera internal parameters; Perform image matching calculation on the speckle image obtained by the camera and the reference plane speckle image recorded in the calibration stage to obtain the disparity map; Use the above disparity map, the reference plane equation, the camera focal length, and the distance between the speckle and the projector to calculate the depth map.

[0095] Step 4: Perform a fusion operation on the three depth maps obtained in the previous step: Taking the depth map m1 calculated with the minimum baseline s1 as the reference, the final depth map m determines which depth map to select the corresponding value from according to the range of the depth intervals [0, d1), [d1, d2), [d2, d3), [d3, +∞) determined for each point in m1 during calibration.

[0096] In the second aspect of the embodiments of the present application, a depth map generation method is provided, which is applied to a processor in a depth map generation system, and the depth map generation system includes: a sensor, a processor, and multiple projectors. Refer to Figure 6 , and the above method includes:

[0097] Step S61: Obtain the image to be processed, where the image to be processed includes the speckles after the light projected by multiple projectors is reflected;

[0098] Step S62: Compare the image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference position of each pixel at each speckle position;

[0099] Step S63: For any baseline, calculate the depth value of each pixel at each speckle position corresponding to the baseline according to the baseline and the disparity corresponding to each pixel at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different;

[0100] Step S64: For any baseline, generate the depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline;

[0101] Step S65: Fuse the depth maps corresponding to each baseline to obtain the target depth map.

[0102] Optionally, for any baseline, calculate the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position, including:

[0103] For any baseline, calculate the depth values of the pixels at each of multiple speckle positions according to the baseline and the parallax corresponding to each pixel at each speckle position;

[0104] For any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline, to obtain the depth values of the pixels at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

[0105] Optionally, for any baseline, calculate the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position, including:

[0106] For any baseline, select the plane equation corresponding to the baseline;

[0107] According to the selected plane equation and the parallax corresponding to each pixel at each speckle position, calculate the depth of each pixel at each speckle position, to obtain the depth values of the pixels at each speckle position.

[0108] It can be seen that through the depth map generation method of the embodiments of the present application, multiple baselines with different distances between multiple projectors and sensors can be used to generate corresponding depth maps respectively, and the generated depth maps are fused, thereby improving the accuracy of depth calculation.

[0109] In the third aspect of the embodiments of the present application, an autonomous mobile device is provided, including a depth map generation system, and the depth map generation system includes: a sensor, a processor, and multiple projectors;

[0110] Multiple projectors, configured to project light;

[0111] A sensor, configured to collect an image to be processed, where the image to be processed includes speckles after the light projected by multiple projectors is reflected;

[0112] A processor is configured to compare an image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference position of each pixel at each speckle position; for any baseline, calculate the depth value of each pixel at each speckle position corresponding to the baseline according to the baseline and the disparity corresponding to each pixel at each speckle position, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and sensors are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each baseline to obtain a target depth map.

[0113] Optionally, the processor is specifically configured to, for any baseline, calculate the depth values of each pixel at multiple speckle positions according to the baseline and the disparity corresponding to each pixel at each speckle position; for any baseline, extract the depth values of the speckles whose depth values are within the preset depth range according to the preset depth range corresponding to the baseline to obtain the depth values of each pixel at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges of different baselines are different.

[0114] Optionally, the processor is specifically configured to, for any baseline, select the plane equation corresponding to the baseline; calculate the depth of each pixel at each speckle position according to the selected plane equation and the disparity corresponding to each pixel at each speckle position to obtain the depth values of each pixel at each speckle position.

[0115] Optionally, multiple projectors are specifically configured to alternately project light through each projector;

[0116] The sensor is specifically configured to collect the image to be processed once every preset time interval.

[0117] Optionally, the exposure duration of the projector closest to the sensor among the multiple projectors is less than the exposure duration of the projector farthest from the sensor.

[0118] Specifically, the above autonomous mobile device can be an automatically movable device such as a mobile robot or a floor cleaning robot. The automatically movable device can perform depth measurement and depth map generation according to the depth map generation system, and then perform autonomous movement or stop according to the detection result. For example, see Figure 7, a plurality of projectors and a detector can be installed at the front and rear ends of the autonomous mobile device respectively. The front and rear projectors can be controlled by a processor to project speckles forward and backward, and the speckle images can be collected by the front and rear sensors. Then, the collected images are input into the processor for calculation to obtain corresponding depth calculation results, such as the clustering of obstacles in front of or behind, so as to stop when an obstacle is detected in front through depth detection, or continue to move forward after the obstacle in front is removed through depth detection.

[0119] It can be seen that through the processor of the embodiment of the present application, multiple baselines with different distances between the projectors and sensors can be utilized to generate corresponding depth maps respectively, and the generated depth maps are fused, thereby improving the accuracy of depth calculation.

[0120] The embodiment of the present invention also provides an electronic device, such as Figure 8 shown, including a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804.

[0121] The memory 803 is used to store computer programs.

[0122] When the processor 801 is used to execute the program stored on the memory 803, the following steps are implemented:

[0123] Obtain an image to be processed, where the image to be processed includes speckles after the light projected by a plurality of projectors is reflected.

[0124] Compare the image to be processed with a preset speckle image to obtain the disparity corresponding to each pixel at the position of each speckle in the image to be processed, where the preset speckle image includes the initial reference position of each pixel at the position of each speckle.

[0125] For any baseline, calculate the depth value of each pixel at the position of each speckle corresponding to the baseline according to the baseline and the disparity corresponding to each pixel at the position of each speckle, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and sensors are different.

[0126] For any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at the position of each speckle corresponding to the baseline.

[0127] Fuse the depth maps corresponding to each baseline to obtain a target depth map.

[0128] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0129] The communication interface is used for communication between the above electronic device and other devices.

[0130] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0131] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0132] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above depth map generation methods are implemented.

[0133] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the depth map generation methods in the above embodiments.

[0134] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0135] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0136] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the method, autonomous mobile device, computer-readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.

Claims

1. A depth map generation system, characterized in that, including a sensor, a processor, and a plurality of projectors; the plurality of projectors are configured to project light, and an exposure duration of a projector closest to the sensor among the plurality of projectors is less than an exposure duration of a projector farthest from the sensor; the sensor is configured to acquire an image to be processed, wherein the image to be processed includes speckles after reflection of light projected by the plurality of projectors; the processor is configured to compare the image to be processed with a preset speckle image to obtain a parallax corresponding to each pixel at each speckle position in the image to be processed, wherein the preset speckle image includes an initial reference position of each pixel at each speckle position; for any baseline, calculate a depth value of each pixel at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each speckle position, wherein the baseline represents a distance between a projector and the sensor, and distances between different projectors and the sensor are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of each pixel at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each of the baselines to obtain a target depth map.

2. The system according to claim 1, wherein the processor is specifically configured to, for any baseline, calculate depth values of each pixel at a plurality of speckle positions according to the baseline and the parallax corresponding to each pixel at each speckle position; for any baseline, extract depth values of speckles whose depth values are within a preset depth range corresponding to the baseline to obtain depth values of each pixel at each speckle position corresponding to the baseline, wherein each baseline corresponds to a preset depth range, and preset depth ranges for different baselines are different.

3. The system according to claim 1, wherein the processor is specifically configured to, for any baseline, select a plane equation corresponding to the baseline; calculate the depth of each pixel at each speckle position according to the selected plane equation and the parallax corresponding to each pixel at each speckle position to obtain depth values of each pixel at each speckle position.

4. The system according to claim 1, wherein the plurality of projectors are specifically configured to project light through each of the projectors alternately; the sensor is specifically configured to acquire an image to be processed once every preset duration.

5. A method for generating a depth map, characterized in that, A processor applied to a depth map generation system, the depth map generation system including: a sensor, a processor, and a plurality of projectors, the method including: acquiring an image to be processed, wherein the image to be processed includes speckles after reflection of light projected by the plurality of projectors, and an exposure duration of a projector closest to the sensor among the plurality of projectors is less than an exposure duration of a projector farthest from the sensor; comparing the image to be processed with a preset speckle image to obtain a parallax corresponding to each pixel at each speckle position in the image to be processed, wherein the preset speckle image includes an initial reference position of each pixel at each speckle position; For any baseline, calculate the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each of the speckle positions, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different; For any baseline, generate a depth map corresponding to the baseline according to the depth values of the pixels at each speckle position corresponding to the baseline; Fuse the depth maps corresponding to each of the baselines to obtain a target depth map.

6. The method according to claim 5, characterized in that, The step of calculating the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each of the speckle positions for any baseline includes: For any baseline, calculate the depth values of the pixels at multiple speckle positions according to the baseline and the parallax corresponding to each pixel at each of the speckle positions; For any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline according to the preset depth range corresponding to the baseline, to obtain the depth values of the pixels at each speckle position corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges of different baselines are different.

7. The method according to claim 5, characterized in that The step of calculating the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each of the speckle positions for any baseline includes: For any baseline, select the plane equation corresponding to the baseline; According to the selected plane equation and the parallax corresponding to each pixel at each of the speckle positions, calculate the depth of each pixel at each speckle position to obtain the depth values of the pixels at each speckle position.

8. An autonomous mobile device, characterized in that, It includes a depth map generation system, and the depth map generation system includes: a sensor, a processor, and multiple projectors; The multiple projectors are used to project light, and the exposure duration of the projector closest to the sensor among the multiple projectors is less than the exposure duration of the projector farthest from the sensor; The sensor is used to collect the image to be processed, where the image to be processed includes speckles after the light projected by the multiple projectors is reflected; The processor is used to compare the image to be processed with a preset speckle image to obtain the parallax corresponding to each pixel at each speckle position in the image to be processed, where the preset speckle image includes the initial reference positions of the pixels at each speckle position; for any baseline, calculate the depth values of the pixels at each speckle position corresponding to the baseline according to the baseline and the parallax corresponding to each pixel at each of the speckle positions, where the baseline represents the distance between the projector and the sensor, and the distances between different projectors and the sensor are different; for any baseline, generate a depth map corresponding to the baseline according to the depth values of the pixels at each speckle position corresponding to the baseline; fuse the depth maps corresponding to each of the baselines to obtain a target depth map.

9. The autonomous mobile device according to claim 8, wherein The processor is specifically configured to calculate the depth values of the pixels at multiple speckle positions for any baseline according to the baseline and the parallax corresponding to each pixel at the respective speckle positions; for any baseline, extract the depth values of the speckles whose depth values are within the preset depth range corresponding to the baseline according to the preset depth range corresponding to the baseline, so as to obtain the depth values of the pixels at the respective speckle positions corresponding to the baseline, where each baseline corresponds to a preset depth range, and the preset depth ranges for different baselines are different.

10. The autonomous mobile device according to claim 8, wherein The processor is specifically configured to select the plane equation corresponding to any baseline; calculate the depth of the pixels at each speckle position according to the selected plane equation and the parallax corresponding to each pixel at the respective speckle positions, so as to obtain the depth values of the pixels at each speckle position.

11. The autonomous mobile device according to claim 8, wherein The multiple projectors are specifically configured to alternately project light through each of the projectors; The sensor is specifically configured to collect the image to be processed once every preset time interval.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 5-7 are implemented.

Citation Information

Patent Citations

  • Depth image acquisition method and device and monocular speckle structured light system

    CN112927280A

  • Machine vision depth estimation method, device and system

    CN113034568A