Depth Map Generation Method, Apparatus, Computer-Readable Storage Medium, and Electronic Device

By generating and processing mapped point sets of point cloud data under the camera coordinate system, the noise problem in depth estimation is solved, and efficient and low-cost high-quality depth map generation is achieved.

CN116645406BActive Publication Date: 2025-07-25SHANGHAI ANTING HORIZON INTELLIGENT TRANSP TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310585457.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-07-25
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

In the prior art, the depth estimation method based on a supervised deep learning model has a problem of high noise in image depth estimation, especially due to the difference in the installation position of the camera and the lidar and the sparse point clouds, it is difficult to effectively remove the depth truth noise caused by the camera and lidar installation position and the sparse point clouds.

Method used

By generating the first depth map, mapped to the camera coordinate system of the target camera, the mapped point set of point cloud data is determined, and invalid points are deleted based on the coordinate order of point cloud data to generate a low-noise depth map.

Benefits of technology

It effectively reduces noise in depth maps, and realizes efficient and low-cost generation of high-quality depth maps, avoiding the needs of complex image processing and multi-sensor consistency verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645406B_ABST
    Figure CN116645406B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, computer-readable storage medium, and electronic device for generating a depth map. The method includes: generating a first depth map based on a point cloud data set collected from a target scene; mapping the first depth map to a camera coordinate system of a target camera based on parameters of the target camera to obtain a second depth map; determining a mapping point set composed of mapping points corresponding to the point cloud data included in the point cloud data set in the second depth map; determining invalid points from the mapping point set based on the coordinate order of the point cloud data in the point cloud data set, and deleting the invalid points from the mapping point set; determining a third depth map representing depth ground truth in the camera coordinate system based on the mapping point set after deleting the invalid points. Embodiments of the present disclosure can not only effectively reduce the noise of the depth ground truth in the depth map, but also efficiently and low-costly generate a high-quality depth map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a method and apparatus for generating a depth map, a computer-readable storage medium, and an electronic device. Background Art

[0002] Image depth estimation refers to estimating the depth of a scene in an image, that is, the distance from each pixel point in the image to the imaging plane of the camera. Training a depth estimation model based on a supervised deep learning model is an important method for image depth estimation currently.

[0003] In a general depth estimation model, before training, it is necessary to collect depth ground truth using a ranging sensor, and the commonly used sensor is a lidar. However, due to the difference in the installation positions of the camera and the lidar and the sparse point cloud collected by the lidar, etc., directly projecting the three-dimensional point cloud onto the image will have a problem that the points projected onto the image do not match the actual scene. For example, the point cloud behind an obstacle will also appear on the image plane, resulting in a large amount of incorrect data (i.e., noise) in the depth ground truth. Therefore, how to reduce the noise in the depth ground truth is a problem that needs to be solved currently. Summary of the Invention

[0004] In order to solve the above technical problems, embodiments of the present disclosure provide a method and apparatus for generating a depth map, a computer-readable storage medium, and an electronic device to solve the problem of how to efficiently reduce the noise included in the depth ground truth of a depth image.

[0005] Embodiments of the present disclosure provide a method for generating a depth map, the method including: generating a first depth map based on a set of point cloud data collected from a target scene; mapping the first depth map to a camera coordinate system of a target camera based on parameters of the target camera to obtain a second depth map; determining a set of mapped points composed of mapped points corresponding to the point cloud data included in the set of point cloud data in the second depth map; determining invalid points from the set of mapped points based on the coordinate order of the point cloud data in the set of point cloud data, and deleting the invalid points from the set of mapped points; determining a third depth map representing the depth ground truth in the camera coordinate system based on the set of mapped points after deleting the invalid points.

[0006] According to another aspect of the embodiments of the present disclosure, a depth map generation device is provided. The device includes: a generation module configured to generate a first depth map based on a set of point cloud data collected from a target scene; a mapping module configured to map the first depth map to a camera coordinate system of a target camera based on parameters of the target camera to obtain a second depth map; a first determination module configured to determine, in the second depth map, a set of mapped points composed of the mapped points corresponding to the point cloud data included in the set of point cloud data; a second determination module configured to determine invalid points from the set of mapped points based on the coordinate order of the point cloud data in the set of point cloud data, and delete the invalid points from the set of mapped points; and a third determination module configured to determine, based on the set of mapped points after deleting the invalid points, a third depth map representing the depth ground truth in the camera coordinate system.

[0007] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the above-mentioned depth map generation method.

[0008] According to another aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes: a processor; a memory configured to store executable instructions of the processor; and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-mentioned depth map generation method.

[0009] According to another aspect of the embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer program instructions, and when the computer program instructions are executed by an instruction processor, the depth map generation method proposed by the present disclosure is executed.

[0010] Based on the depth map generation method, device, computer-readable storage medium, and electronic device provided in the above embodiments of the present disclosure, by determining, in the depth map in the camera coordinate system, a set of mapped points corresponding to the set of point cloud data, determining invalid points from the set of mapped points based on the coordinate order of the point cloud data in the set of point cloud data, and deleting the invalid points from the set of mapped points, a depth map with low noise is obtained. This method does not require complex processing of images, nor does it require the use of multiple sensors to verify the consistency of the mapped points of the point cloud in the depth map. It only needs to use the coordinate order of the point cloud data to judge the validity of the spatial distribution of the point cloud, and obtain invalid points that do not match the shooting angle of the camera in the camera coordinate system, thereby effectively reducing the noise of the depth ground truth in the depth map and generating a high-quality depth map efficiently and at low cost.

[0011] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0012] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation on the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps;

[0013] Figure 1 is a system diagram applicable to the present disclosure;

[0014] Figure 2 is a schematic flowchart of a depth map generation method provided by an exemplary embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram of determining invalid mapping points in the camera coordinate system according to an embodiment of the present disclosure;

[0016] Figure 4 is a schematic flowchart of a depth map generation method provided by another exemplary embodiment of the present disclosure;

[0017] Figure 5 is a schematic flowchart of a depth map generation method provided by another exemplary embodiment of the present disclosure;

[0018] Figure 6A is a schematic diagram of dividing the azimuth range of a point cloud acquisition device into multiple angular regions according to an embodiment of the present disclosure;

[0019] Figure 6B is a schematic diagram of the laser beam lines within the depression angle range of a point cloud acquisition device according to an embodiment of the present disclosure;

[0020] Figure 7 is a schematic flowchart of a depth map generation method provided by another exemplary embodiment of the present disclosure;

[0021] Figure 8 is a schematic flowchart of a depth map generation method provided by another exemplary embodiment of the present disclosure;

[0022] Figure 9 is a schematic flowchart of a depth map generation method provided by another exemplary embodiment of the present disclosure;

[0023] Figure 10 is a schematic structural diagram of a depth map generation device provided by an exemplary embodiment of the present disclosure;

[0024] Figure 11 is a schematic structural diagram of a depth map generation device provided by another exemplary embodiment of the present disclosure;

[0025] Figure 12 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0026] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0027] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0028] Overview of the Application

[0029] To reduce the noise of the depth ground truth in the depth map, commonly used methods currently include using pooling operations in the image, or morphological methods to remove point clouds that are visually inconsistent, etc. The effect of removing noise by such methods is poor, and it is impossible to comprehensively detect and remove the depth value noise. It is also possible to use multiple sensors for consistency verification to reduce the noise of the depth ground truth in the depth map, but the cost of using this method is relatively high, and it usually requires coupling the characteristics of the camera. When using multiple cameras for consistency verification, due to the different parameters between multiple cameras and the non-uniform distribution of the point cloud itself, the difficulty of parameter adjustment is greatly increased.

[0030] Exemplary System

[0031] Figure 1 An exemplary system architecture 100 of a depth map generation method or a depth map generation device to which the embodiments of the present disclosure can be applied is shown.

[0032] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, a server 103, a point cloud acquisition device 104, and a camera 105.

[0033] The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0034] The point cloud acquisition device 104 is used to perform point cloud acquisition on a target scene to obtain a set of point cloud data representing the positions of objects in the target scene. The point cloud acquisition device 104 may include, but is not limited to, at least one of the following devices: lidar, binocular stereo camera, etc. The point cloud acquisition device 104 may be set at any position. For example, the point cloud acquisition device 104 may be set on a vehicle, and the target scene may be a road where the vehicle travels, a parking lot, or other scenes. For another example, the point cloud acquisition device 104 may be set on an aircraft, and the target scene may be the scene where the aircraft is flying.

[0035] The camera 105 is used to capture images of the target scene, and the set of point cloud data collected by the point cloud acquisition device 104 can be mapped into the images captured by the camera 105.

[0036] The terminal device 101 can interact with the server 103 through the network 102 to receive or send messages, etc. Various applications can be installed on the terminal device 101, such as monitoring applications, navigation applications, etc.

[0037] The terminal device 101 can be various electronic devices, including but not limited to mobile terminals such as in-vehicle terminals, mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (tablet computers), etc. and fixed terminals such as digital TVs, desktop computers, etc.

[0038] The server 103 can be a server that provides various services, such as a background server that receives the set of point cloud data, two-dimensional images, etc. uploaded by the terminal device 101 to generate a depth map.

[0039] It should be noted that the depth map generation method provided by the embodiments of the present disclosure can be executed by the server 103 or by the terminal device 101. Correspondingly, the depth map generation device can be set in the server 103 or in the terminal device 101.

[0040] It should be understood that Figure 1 the number of terminal devices, networks, servers, point cloud acquisition devices, and cameras in

[0041] Exemplary Method

[0042] Figure 2 is a schematic flowchart of the depth map generation method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device (such as Figure 1 the terminal device 101 or the server 103 shown) as Figure 2 shown, and the method includes the following steps:

[0043] Step 201, generate a first depth map based on the set of point cloud data collected from the target scene.

[0044] Among them, each piece of point cloud data included in the set of point cloud data usually includes a coordinate value, and this coordinate value represents a point in the target scene in Figure 1The position in the coordinate system of the point cloud acquisition device 104 shown (i.e., the coordinate system established with the position of the point cloud acquisition device 104 as the origin). The three-dimensional coordinate values corresponding to the pixel points in the first depth map can be obtained from the point cloud data corresponding to the pixel points. For example, in a rectangular coordinate system with the position of the point cloud acquisition device 104 as the origin, it includes three coordinate axes, x, y, and z. The z-axis is the vertical coordinate axis, the x-axis is the coordinate value in the direction of the optical axis of the point cloud acquisition device, and the y-axis is the horizontal coordinate axis. The three-dimensional coordinate values corresponding to the pixel points in the first depth map are the (x, y, z) coordinate values, where the x coordinate value can be used as the depth value, representing the distance between the three-dimensional space point corresponding to the pixel point and the point cloud acquisition device.

[0045] Step 202, based on the parameters of the target camera, map the first depth map to the camera coordinate system of the target camera to obtain a second depth map.

[0046] Among them, the target camera can be the camera 105 as shown in Figure 1 Since the positions of the camera 105 and the point cloud acquisition device 104 are different, it is necessary to map the first depth map to the camera coordinate system. Generally, the parameters of the camera can include external parameters, which refer to the parameters representing the mapping relationship between the points in the world coordinate system and the points in the camera coordinate system. As an example, in this embodiment, the external parameters of the point cloud acquisition device 104 can be used to map the pixel points in the first depth map to the world coordinate system, and then the external parameters of the target camera can be used to map the points in the world coordinate system to the camera coordinate system of the target camera to obtain a second depth map.

[0047] Step 203, determine a set of mapped points composed of the mapped points corresponding to the point cloud data included in the point cloud data set in the second depth map.

[0048] Since the mapping relationship between the pixel points in the first depth map and the pixel points in the second depth map has been obtained in step 202, and the mapping relationship between the point cloud data included in the point cloud data set and the pixel points in the first depth map is known, a set of mapped points composed of the mapped points corresponding to the point cloud data included in the point cloud data set can be determined in the second depth map.

[0049] Step 204, based on the coordinate order of the point cloud data in the point cloud data set, determine invalid points from the set of mapped points and delete the invalid points from the set of mapped points.

[0050] Specifically, the coordinate order of the point cloud data represents the arrangement order of the spatial point set indicated by the point cloud data set in space. Since the spatial points represented by the point cloud data are located on the surface of the object, in the camera coordinate system, if there is no occlusion of the spatial points by the object, the coordinate order of the mapped points in the mapped point set in the camera coordinate system is consistent with the coordinate order of the point cloud data in the point cloud data set. If occlusion occurs, the arrangement order of the coordinates of the occluded points and the unoccluded points may be disordered. At this time, the mapped points with disordered order can be determined as invalid points.

[0051] As Figure 3 shown, O1 is the coordinate origin of the point cloud acquisition device, and O2 is the coordinate origin of the target camera. 301 is an object in the target scene. Points A, B, and C are located on the surface of the object, and points A, B, and C respectively correspond to a point cloud data. It can be seen from the figure that in the coordinate system of the point cloud acquisition device, the order of the elevation angles of the lines connecting points A, B, and C to O1 from small to large is C, B, A. In the camera coordinate system of the target camera, since the line connecting point A to O2 passes through the object 301, that is, the object 301 occludes point A, resulting in a change in the arrangement order of points A, B, and C. That is, in the camera coordinate system, the order of the elevation angles of the lines connecting points A, B, and C to O2 from small to large is C, A, B, no longer the order of C, B, A. And the occlusion of point A by the object 301 is the reason for the order change. Therefore, point A can be determined as an invalid point.

[0052] Step 205, based on the mapped point set after deleting the invalid points, determine the third depth map representing the depth ground truth in the camera coordinate system.

[0053] Specifically, the depth values of the pixels corresponding to the invalid points in the second depth map can be deleted or set to a preset depth value to obtain the third depth map. The depth value corresponding to the pixel point in the third depth map is the depth ground truth, and the depth ground truth can accurately reflect the distance between the object corresponding to the pixel point in the image captured by the target camera and the target camera.

[0054] The method provided by the above embodiments of the present disclosure determines the mapped point set corresponding to the point cloud data set in the depth map in the camera coordinate system, determines the invalid points from the mapped point set based on the coordinate order of the point cloud data in the point cloud data set, and deletes the invalid points from the mapped point set, thereby obtaining a depth map with low noise. This method does not require complex processing of the image, nor does it require using multiple sensors to verify the consistency of the mapped points of the point cloud in the depth map. It only needs to use the coordinate order of the point cloud data to judge the validity of the spatial distribution of the point cloud, and obtain the invalid points that do not match the shooting angle of the camera in the camera coordinate system, thereby realizing both effectively reducing the noise of the depth ground truth in the depth map and efficiently and low-costly generating a high-quality depth map.

[0055] In some alternative implementations, such as Figure 4 shown, step 204 includes:

[0056] Step 2041, dividing the point cloud data set into at least two point cloud data subsets.

[0057] Among them, there are various ways to divide the point cloud data set. For example, the azimuth range of the point cloud acquisition device (i.e., the range of horizontal swing) can be divided into at least two angular regions, and the point cloud data corresponding to the points in each angular region is a point cloud data subset. Or, the pitch angle range of the point cloud acquisition device (i.e., the range of vertical swing) can be divided into at least two angular regions, and the point cloud data corresponding to the points in each angular region is a point cloud data subset.

[0058] Step 2042, for each point cloud data subset among the at least two point cloud data subsets, based on the preset arrangement direction, determine the serial number of each point cloud data in this point cloud data subset.

[0059] It should be understood that step 2042 - step 2044 are for each point cloud data subset among the at least two point cloud data subsets, that is, when processing the data, it is described for one of the point cloud data subsets, and the processing methods for other point cloud data subsets are the same.

[0060] The above preset arrangement direction is the arrangement direction within the spatial points corresponding to the point cloud data included in the point cloud data subset. For example, if the point cloud data subset is divided according to the azimuth angle, the preset arrangement direction can be the change direction of the pitch angle (or the vertical direction); if the point cloud data subset is divided according to the pitch angle, the preset arrangement direction can be the change direction of the azimuth angle (or the horizontal direction).

[0061] The serial number of the point cloud data can be generated based on the coordinates of the point cloud data. For example, for a point cloud data set, the serial numbers are set in ascending order of the pitch angle of each point cloud data (e.g., elevation_0, elevation_1...).

[0062] Step 2043, in the mapping point set, determine the mapping point subset composed of the mapping points corresponding to each point cloud data in this point cloud data subset.

[0063] Since step 2042 has determined the correspondence between the point cloud data set and the mapping point set, therefore, this step can determine the mapping point subset corresponding to each point cloud data subset.

[0064] Step 2044, from the mapping point subset, according to the preset arrangement direction, determine the mapping points whose corresponding serial numbers of the point cloud data do not conform to the preset arrangement order as invalid points.

[0065] As an example, as Figure 3 shown, if the preset arrangement direction is the changing direction of the pitch angle, in the coordinate system of the point cloud acquisition device, the sequence numbers of the three point cloud data included in the point cloud data subset are C, B, A in the initial order from the smallest pitch angle to the largest. In the camera coordinate system, the three point cloud data are also in the order of C, A, B from the smallest pitch angle to the largest. Therefore, the arrangement position of A makes the arrangement order of the three point cloud data no longer consistent with the initial order, and the mapping point corresponding to A is determined as an invalid point. The reason for determining the mapping point corresponding to A as an invalid point is that in the camera coordinate system, the object 301 blocks A.

[0066] In this embodiment, by dividing the point cloud data set into at least two point cloud data subsets, in each point cloud data subset, invalid mapping points are determined based on the arrangement order of the point cloud data. Since the spatial range represented by the point cloud data subset is relatively small compared to the entire point cloud data set, it is possible to avoid the difficulty in judging the distribution law of the point cloud data when processing the point cloud data due to a large number of point cloud data in a large space. The distribution law of the point cloud data in a small range can be judged more accurately, thereby realizing a higher-precision validity judgment of the point cloud data, further improving the recognition efficiency of the valid point cloud data, and ensuring the accuracy of generating the depth ground truth at the same time.

[0067] In some alternative implementation manners, as Figure 5 shown, step 2041 includes:

[0068] Step 20411, divide the space into at least two regions according to the first dimension of the space where the point cloud data set is distributed.

[0069] Among them, the first dimension can be a dimension in any preset direction. As an example, the first dimension can be the azimuth angle dimension (or the horizontal field of view angle dimension). The azimuth angle refers to the horizontal scanning angle range of the point cloud acquisition device with the position of the point cloud acquisition device as the origin. Therefore, the azimuth angle of the space where the point cloud data set is distributed can be divided into at least two angular regions. As Figure 6A shown, it shows a top view of the acquisition range of the point cloud acquisition device. The range of the azimuth angle is divided into multiple angular regions, and the angle of each angular region is α. The spatial region corresponding to each angular region ( Figure 6A the region divided by the vertical plane where the straight line shown is the boundary surface) is the region after dividing the space where the point cloud data set is distributed.

[0070] Optionally, the first dimension can also be the pitch angle dimension (or the vertical field of view angle dimension). According to a method similar to the above-mentioned region division based on the azimuth angle dimension, the space can be divided into at least two regions.

[0071] Step 20412: Determine the point cloud data in the same area among at least two areas as a point cloud data subset, obtaining at least two point cloud data subsets.

[0072] In this embodiment, by partitioning the space of the point cloud data set distribution based on the first dimension, the partitioned areas can be matched with attributes such as the scanning direction of the point cloud acquisition device and the offset of the scanning direction, enabling each point cloud data subset to accurately reflect the characteristics of the actual three-dimensional space scanned by the point cloud acquisition device, which helps to more accurately judge the validity of the point cloud data.

[0073] In some alternative implementation manners, such as Figure 7 shown, step 2042 includes:

[0074] Step 20421: In response to determining that the point cloud data subset contains first point cloud data with a corresponding laser beam line number, determine the laser beam line number as the number of the first point cloud data.

[0075] Among them, the laser beam line numbers are distributed in a preset arrangement direction. Specifically, when the point cloud acquisition device is a lidar, the lidar scans according to a preset scanning angle resolution. Therefore, the scanned point cloud data can include laser beam line numbers. As Figure 6B shown, it shows the elevation angle range (or vertical field of view angle) scanned by the lidar in the vertical direction. The vertical field of view angle range shown in FIG. 6 is [-25.0°, +25°], the number of laser beam lines is 128, and the laser beam line numbers are elevation_0, elevation_1,..., elevation_127. Therefore, the order of the laser beam line numbers can reflect the arrangement direction of the point cloud data.

[0076] Step 20422: In response to determining that the point cloud data subset contains second point cloud data without a corresponding laser beam line number, determine the number of the second point cloud data in the point cloud data subset based on the magnitude of the coordinate values of the second point cloud data in the point cloud data subset.

[0077] Specifically, in some application scenarios, if the lidar does not have the function of generating laser beam line numbers in the point cloud data, or if the laser beam line numbers are not added to the point cloud data due to acquisition errors, etc., the collected point cloud data set includes second point cloud data without corresponding laser beam line numbers. If a certain point cloud data subset contains the second point cloud data, the number of the second point cloud data can be generated according to the coordinate values of the second point cloud data.

[0078] As an example, the coordinate values of the second point cloud data include coordinates representing height, and the height coordinate can be determined as the serial number of the second point cloud data. Alternatively, based on the coordinate values of the second point cloud data, the pitch angle can be calculated, and then according to Figure 6B the pitch angle range and the range of the laser beam line numbers shown, the laser beam line number of the second point cloud data can be calculated.

[0079] It should be understood that when the point cloud data subset includes both the first point cloud data and the second point cloud data, the serial number of the second point cloud data should have the same dimension as the serial number of the first point cloud data, that is, the serial numbers of the first point cloud data and the second point cloud data are comparable. For example, the serial numbers of the first point cloud data and the second point cloud data can both be pitch angle values.

[0080] In this embodiment, by directly reading the laser beam line number as the serial number of the point cloud data when the point cloud data contains the laser beam line number, and generating the serial number of the point cloud data according to the coordinate values when the point cloud data does not contain the laser beam line number, the serial numbers of all the point cloud data included in the point cloud data subset can be made comparable, thereby improving the universality and accuracy of the validity judgment of the point cloud data.

[0081] In some alternative implementation manners, step 20422 can be executed according to the following steps:

[0082] First, determine a preset arrangement direction based on the second dimension of the space in which the point cloud data set is distributed.

[0083] As an example, when the first dimension direction is the azimuth angle dimension, the second dimension can be the pitch angle dimension, and the preset arrangement direction can be the direction in which the pitch angle increases from small to large; when the first dimension direction is the pitch angle dimension, the second dimension can be the azimuth angle dimension, and the preset arrangement direction can be the direction in which the azimuth angle increases from small to large.

[0084] Then, based on the magnitude of the coordinate values of the second point cloud data in the point cloud data subset, determine the angle of the second point cloud data in the point cloud data subset relative to the preset arrangement direction.

[0085] Generally, the coordinate values of the second point cloud data are coordinates in a rectangular coordinate system, that is, including three coordinate values of x, y, and z. Based on x, y, and z, the pitch angle can be calculated, and the pitch angle can be determined as the preset arrangement direction angle. Alternatively, based on x, y, and z, the azimuth angle can be calculated, and the azimuth angle can be determined as the preset arrangement direction angle.

[0086] Finally, based on the angle and the preset angle resolution, determine the serial number of the second point cloud data in the point cloud data subset.

[0087] As an example, when the above angle is the pitch angle, the pitch angle can be divided by the pitch angle resolution to obtain the serial number of the second point cloud data. Or, according to the laser beam serial number range shown in Figure 6B the numerical value obtained by dividing the pitch angle by the pitch angle resolution is mapped to the laser beam serial number range, so as to obtain the serial number of the second point cloud data.

[0088] Similarly, when the above angle is the azimuth angle, the serial number of the second point cloud data can be determined based on the azimuth angle and the azimuth angle resolution in a similar manner.

[0089] In this embodiment, by determining the angle of the second point cloud data in the point cloud data subset and determining the serial number of the second point cloud data according to the angle, the obtained serial number can accurately represent the distribution order of the point cloud data in the point cloud data subset, and the serial number of the second point cloud data and the serial number of the first point cloud data can be arranged in the same dimension, which is beneficial to making all the point cloud data in the point cloud data subset comparable, thereby helping to improve the universality and accuracy of determining invalid points in the second depth map.

[0090] In some alternative implementation manners, as shown in Figure 8 step 2044 includes:

[0091] Step 20441, traverse the mapping points in the mapping point subset according to the preset arrangement order of the serial numbers of the corresponding point cloud data, and determine whether the coordinates corresponding to the currently traversed mapping point satisfy the target extreme value condition corresponding to the preset arrangement direction compared with the coordinates corresponding to the already traversed mapping points.

[0092] The target extreme value condition refers to whether the coordinates corresponding to the currently traversed mapping point are the maximum or minimum compared with the coordinates corresponding to the already traversed mapping points, that is, to determine whether the monotonicity of the coordinates of the mapping point is consistent with the monotonicity of the corresponding serial number. Optionally, when judging whether the target extreme value condition is satisfied based on the coordinates, the x, y, z coordinate values (such as the z coordinate value representing the height in the vertical direction) can be used for judgment, or the angle (such as the pitch angle or the azimuth angle) can be used for judgment.

[0093] For example, when the preset arrangement order is from small to large, traverse the mapping points in the order of the serial numbers from small to large. If the angle (such as the pitch angle) corresponding to the currently traversed mapping point is greater than the angle corresponding to the already traversed mapping point, it is determined that the currently traversed mapping point satisfies the first target extreme value condition corresponding to the order from small to large.

[0094] Or, when the preset arrangement order is from large to small, traverse the mapping points in the order of the serial numbers from large to small. If the angle corresponding to the currently traversed mapping point is less than the angle corresponding to the already traversed mapping point, it is determined that the currently traversed mapping point satisfies the second target extreme value condition corresponding to the order from large to small.

[0095] Step 20442: If the target extreme value condition is not satisfied, determine that the currently traversed mapping point is an invalid point.

[0096] As Figure 3 shown, the sequence numbers of the point cloud data A, B, and C are 3, 2, and 1 respectively. In the second depth map, traverse the mapping points corresponding to C, B, and A in ascending order of the sequence numbers.

[0097] When traversing the mapping point c corresponding to C for the first time, the pitch angle of this mapping point c is the smallest, and it is determined that this mapping point c satisfies the target extreme value condition.

[0098] When traversing the mapping point b corresponding to B for the second time, the pitch angle of this mapping point b is the maximum pitch angle compared to the pitch angle of c, and it is determined that this mapping point b satisfies the target extreme value condition.

[0099] When traversing the mapping point a corresponding to A for the third time, the pitch angle of this mapping point a is less than the pitch angle of b, that is, the pitch angle of a is not the maximum pitch angle compared to the pitch angles of c and b, and it is determined that this mapping point a does not satisfy the target extreme value condition, and a is an invalid point.

[0100] This embodiment realizes determining whether the monotonicity of the coordinate change of the mapping point is consistent with the monotonicity of the sequence number corresponding to the mapping point according to the preset arrangement direction of the mapping points, and further determines the mapping points with inconsistent monotonicity as invalid points. Since the sequence number of the mapping point corresponds to the coordinates of the mapping point, and the sequence number of the mapping point is easier to obtain compared to the coordinate change, the method for determining whether the mapping point is an invalid point based on the monotonicity of the sequence number change is simple and efficient, which helps to improve the efficiency of determining invalid points from the mapping point set.

[0101] In some alternative implementation manners, as Figure 9 shown, before step 201, the method further includes:

[0102] Step 901: Obtain the initial point cloud data set collected by the point cloud acquisition device for the target scene.

[0103] Among them, the initial point cloud data set is the set of unprocessed point cloud data directly collected by the point cloud acquisition device.

[0104] Step 902: Based on the spatial information of the target object in the target scene determined in advance, delete the point cloud data corresponding to the target object from the initial point cloud data set to obtain the point cloud data set.

[0105] Among them, the target object can be a pre-specified object. For example, when the point cloud acquisition device and the camera are set on a vehicle, the target object is the vehicle. The electronic device can use the target detection method based on the point cloud to determine information such as the position and size of the vehicle, determine the spatial points located on the vehicle from the spatial point set represented by the initial point cloud data set, and delete the point cloud data corresponding to the spatial points located on the vehicle to obtain the above-mentioned point cloud data set.

[0106] In this embodiment, by detecting the target object in the target scene and deleting the point cloud data corresponding to the target object, it is possible to specifically extract part of the point cloud data as the basis for generating the depth map, avoid the interference during data processing caused by including the point cloud data that does not need to be processed in the actual scene in the point cloud data set, improve the accuracy of generating the depth map, and at the same time help reduce the data processing amount for generating the depth map and improve the efficiency of generating the depth map.

[0107] Exemplary Device

[0108] Figure 10 It is a schematic structural diagram of a depth map generation device provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 10 As shown, the depth map generation device includes: a generation module 1001 for generating a first depth map based on the point cloud data set collected from the target scene; a mapping module 1002 for mapping the first depth map to the camera coordinate system of the target camera based on the parameters of the target camera to obtain a second depth map; a first determination module 1003 for determining a mapping point set composed of mapping points corresponding to the point cloud data included in the point cloud data set in the second depth map; a second determination module 1004 for determining invalid points from the mapping point set based on the coordinate order of the point cloud data in the point cloud data set and deleting the invalid points from the mapping point set; a third determination module 1005 for determining a third depth map representing the depth truth value in the camera coordinate system based on the mapping point set after deleting the invalid points.

[0109] In this embodiment, the generation module 1001 can generate a first depth map based on the point cloud data set collected from the target scene. Among them, each point cloud data included in the point cloud data set usually includes a coordinate value, and this coordinate value represents a point in the target scene Figure 1The position in the coordinate system of the point cloud acquisition device 104 shown (i.e., the coordinate system established with the position of the point cloud acquisition device 104 as the origin). The depth value corresponding to the pixel point in the first depth map can be obtained according to the point cloud data corresponding to the pixel point. For example, in a rectangular coordinate system with the position of the point cloud acquisition device 104 as the origin, it includes three coordinate axes: x, y, and z. The z-axis is the vertical coordinate axis, the x-axis is the coordinate value in the optical axis direction of the point cloud acquisition device, and the y-axis is the horizontal coordinate axis. The depth value corresponding to the pixel point in the first depth map is the x coordinate value.

[0110] In this embodiment, the mapping module 1002 can map the first depth map to the camera coordinate system of the target camera based on the parameters of the target camera to obtain a second depth map. Among them, the target camera can be the camera 105 shown as follows. The position of the camera 105 is different from that of the point cloud acquisition device 104. Therefore, it is necessary to map the first depth map to the camera coordinate system. The parameters of the camera can include internal parameters and external parameters. The internal parameters refer to the parameters representing the mapping relationship between the points in the camera coordinate system and the points in the image plane coordinate system, and the external parameters refer to the parameters representing the mapping relationship between the points in the world coordinate system and the points in the camera coordinate system. As an example, in this embodiment, the parameters of the point cloud acquisition device 104 can be used to map the pixel points in the first depth map to the world coordinate system, and then map the points in the world coordinate system to the camera coordinate system of the target camera. Using the parameters of the target camera, map the points in the camera coordinate system to the image plane to obtain a second depth map. Figure 1 In this embodiment, the first determination module 1003 can determine, in the second depth map, a set of mapping points composed of the mapping points corresponding to the point cloud data included in the point cloud data set. Since the mapping relationship between the pixel points in the first depth map and the pixel points in the second depth map has been obtained in step 202, and the mapping relationship between the point cloud data included in the point cloud data set and the pixel points in the first depth map is known, it is possible to determine, in the second depth map, a set of mapping points composed of the mapping points corresponding to the point cloud data included in the point cloud data set.

[0111] In this embodiment, the first determination module 1003 can determine, in the second depth map, a set of mapping points composed of the mapping points corresponding to the point cloud data included in the point cloud data set. Since the mapping relationship between the pixel points in the first depth map and the pixel points in the second depth map has been obtained in step 202, and the mapping relationship between the point cloud data included in the point cloud data set and the pixel points in the first depth map is known, it is possible to determine, in the second depth map, a set of mapping points composed of the mapping points corresponding to the point cloud data included in the point cloud data set.

[0112] In this embodiment, the second determination module 1004 may determine invalid points from the mapped point set based on the coordinate order of the point cloud data in the point cloud data set, and delete the invalid points from the mapped point set. Specifically, the coordinate order of the point cloud data represents the arrangement order of the spatial point set indicated by the point cloud data set in space. Since the spatial points represented by the point cloud data are located on the object surface, in the camera coordinate system, if there is no object occlusion of the spatial points, the coordinate order of the mapped points in the mapped point set in the camera coordinate system is consistent with the coordinate order of the point cloud data in the point cloud data set. If occlusion occurs, the arrangement order of the coordinates of the occluded points and the unoccluded points may be disordered. At this time, the mapped points with disordered order can be determined as invalid points.

[0113] In this embodiment, the third determination module 1005 may determine a third depth map representing the depth ground truth in the camera coordinate system based on the mapped point set after deleting the invalid points. Specifically, the depth values of the pixels corresponding to the invalid points in the second depth map may be deleted or set to a preset depth value to obtain the third depth map. The depth value corresponding to the pixel point in the third depth map is the depth ground truth, and the depth ground truth can accurately reflect the distance between the object corresponding to the pixel point in the image captured by the target camera and the target camera.

[0114] Refer to Figure 11 , Figure 11 is a schematic structural diagram of a depth map generation device provided by another exemplary embodiment of the present disclosure.

[0115] In some optional implementation manners, the second determination module 1004 includes: a partitioning unit 10041, configured to partition the point cloud data set into at least two point cloud data subsets; a first determination unit 10042, configured to, for each point cloud data subset in the at least two point cloud data subsets, determine the serial number of each point cloud data in the point cloud data subset based on a preset arrangement direction; a second determination unit 10043, configured to determine a mapped point subset composed of the mapped points corresponding to each point cloud data in the point cloud data subset in the mapped point set; and a third determination unit 10044, configured to determine, from the mapped point subset, the mapped points whose serial numbers of the corresponding point cloud data do not conform to the preset arrangement order as invalid points according to the preset arrangement direction.

[0116] In some optional implementation manners, the partitioning unit 10041 includes: a partitioning subunit 100411, configured to partition the space into at least two regions according to a first dimension of the space in which the point cloud data set is distributed; and a first determination subunit 100412, configured to determine the point cloud data in the same region among the at least two regions as a point cloud data subset, so as to obtain at least two point cloud data subsets.

[0117] In some alternative implementation manners, the first determination unit 10042 includes: a second determination subunit 100421, configured to, in response to determining that the point cloud data subset includes first point cloud data having a corresponding laser beam line number, determine that the laser beam line number is the number of the first point cloud data, where the laser beam line numbers are distributed according to a preset arrangement direction; and a third determination subunit 100422, configured to, in response to determining that the point cloud data subset includes second point cloud data not having a corresponding laser beam line number, determine the number of the second point cloud data in the point cloud data subset based on the magnitude of the coordinate values of the second point cloud data in the point cloud data subset.

[0118] In some alternative implementation manners, the third determination subunit is further configured to: determine a preset arrangement direction based on a second dimension of the space in which the point cloud data set is distributed; determine an angle of the second point cloud data in the point cloud data subset relative to the preset arrangement direction based on the magnitude of the coordinate values of the second point cloud data in the point cloud data subset; and determine the number of the second point cloud data in the point cloud data subset based on the angle and a preset angle resolution.

[0119] In some alternative implementation manners, the third determination unit 10044 includes: a fourth determination subunit 100441, configured to traverse the mapped points in the mapped point subset in a preset arrangement order according to the numbers of the corresponding point cloud data, and determine whether the coordinates of the currently traversed mapped point satisfy a target extreme value condition corresponding to the preset arrangement direction as compared with the coordinates of the already traversed mapped points; and a fifth determination subunit 100442, configured to, if the target extreme value condition is not satisfied, determine that the currently traversed mapped point is an invalid point.

[0120] In some alternative implementation manners, the apparatus further includes: an acquisition module 1006, configured to acquire an initial point cloud data set collected by a point cloud acquisition device for a target scene; and a deletion module 1007, configured to delete, from the initial point cloud data set, the point cloud data corresponding to a target object based on the spatial information of the target object in the target scene determined in advance, so as to obtain a point cloud data set.

[0121] The depth map generation device provided in the above embodiments of the present disclosure determines a set of mapped points corresponding to a point cloud data set in the depth map in the camera coordinate system, determines invalid points from the set of mapped points based on the coordinate order of the point cloud data in the point cloud data set, and deletes the invalid points from the set of mapped points, thereby obtaining a depth map with low noise. This method does not require complex processing of the image, nor does it require using multiple sensors to verify the consistency of the mapped points of the point cloud in the depth map. It only needs to utilize the coordinate order of the point cloud data to judge the validity of the spatial distribution of the point cloud, and obtain invalid points that do not match the shooting angle of the camera in the camera coordinate system, thereby achieving both effectively reducing the noise of the depth truth value in the depth map and efficiently and low-costly generating a high-quality depth map.

[0122] Exemplary Electronic Device

[0123] Next, refer to Figure 12 to describe an electronic device according to an embodiment of the present disclosure. The electronic device may be any one or both of the terminal device 101 and the server 103 as shown in Figure 1 or a stand-alone device independent of them, and the stand-alone device may communicate with the terminal device 101 and the server 103 to receive the input signals collected from them.

[0124] Figure 12 FIG. shows a block diagram of an electronic device according to an embodiment of the present disclosure.

[0125] As shown in Figure 12 the electronic device 1200 includes one or more processors 1201 and a memory 1202.

[0126] The processor 1201 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1200 to perform desired functions.

[0127] The memory 1202 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1201 may run the program instructions to implement the depth map generation method of the above embodiments of the present disclosure and / or other desired functions. Various contents such as a set of point cloud data and two-dimensional images may also be stored in the computer-readable storage medium.

[0128] In one example, the electronic device 1200 may further include: an input device 1203 and an output device 1204, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0129] For example, when the electronic device is the terminal device 101 or the server 103, the input device 1203 may be a point cloud acquisition device, a camera, a mouse, a keyboard, etc., for inputting a point cloud data set, a two-dimensional image, etc. When the electronic device is a stand-alone device, the input device 1203 may be a communication network connector for receiving the input point cloud data set, two-dimensional image, etc. from the terminal device 101 and the server 103.

[0130] The output device 1204 may output various information to the outside, including the generated depth map, etc. The output device 1204 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0131] Of course, for simplicity, Figure 12 only some of the components related to the present disclosure in the electronic device 1200 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1200 may further include any other appropriate components.

[0132] Exemplary Computer Program Product and Computer Readable Storage Medium

[0133] In addition to the above methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which when run by a processor cause the processor to execute the steps in the depth map generation methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0134] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0135] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the depth map generation method of various embodiments of the present disclosure described in the "Exemplary Method" section above.

[0136] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium includes, for example but not limited to, a system, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0137] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-described specific details are only for the purposes of illustration and facilitation of understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0138] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.

Claims

1. A method for generating a depth map, comprising: Generating a first depth map based on a set of point cloud data collected from a target scene; Mapping the first depth map to the camera coordinate system of the target camera based on the parameters of the target camera to obtain a second depth map; Determining, in the second depth map, a set of mapped points composed of the mapped points corresponding to the point cloud data included in the set of point cloud data; Determining invalid points from the set of mapped points based on the coordinate order of the point cloud data in the set of point cloud data, and deleting the invalid points from the set of mapped points; Determining a third depth map representing the depth ground truth in the camera coordinate system based on the set of mapped points after deleting the invalid points.

2. The method according to claim 1, wherein The determining invalid points from the set of mapped points based on the coordinate order of the point cloud data in the set of point cloud data includes: Dividing the set of point cloud data into at least two subsets of point cloud data; For each subset of point cloud data in the at least two subsets of point cloud data, determining the serial number of each point cloud data in the subset of point cloud data based on a preset arrangement direction; Determining, in the set of mapped points, a subset of mapped points composed of the mapped points corresponding to each point cloud data in the subset of point cloud data; Determining, from the subset of mapped points, the mapped points whose serial numbers of the corresponding point cloud data do not conform to the preset arrangement order as invalid points according to the preset arrangement direction.

3. The method according to claim 2, wherein The dividing the set of point cloud data into at least two subsets of point cloud data includes: Dividing the space according to a first dimension of the space in which the set of point cloud data is distributed into at least two regions; Determining the point cloud data in the same region among the at least two regions as a subset of point cloud data to obtain the at least two subsets of point cloud data.

4. The method according to claim 2, wherein, The determining the serial number of each point cloud data in the subset of point cloud data based on a preset arrangement direction includes: In response to determining that the subset of point cloud data includes first point cloud data having a corresponding laser beam line serial number, determining the laser beam line serial number as the serial number of the first point cloud data, wherein the laser beam line serial numbers are distributed according to the preset arrangement direction; In response to determining that the subset of point cloud data includes second point cloud data without a corresponding laser beam line serial number, determining the serial number of the second point cloud data in the subset of point cloud data based on the magnitude of the coordinate values of the second point cloud data in the subset of point cloud data.

5. The method according to claim 4, wherein The determining the serial number of the second point cloud data in the subset of point cloud data based on the magnitude of the coordinate values of the second point cloud data in the subset of point cloud data includes: Determining a preset arrangement direction based on a second dimension of the space in which the set of point cloud data is distributed; determining the angle of the second point cloud data in the subset of point cloud data relative to the preset arrangement direction based on the magnitude of the coordinate values of the second point cloud data in the subset of point cloud data; Determining the serial number of the second point cloud data in the subset of point cloud data based on the angle and a preset angle resolution.

6. The method according to claim 2, wherein The determining, from the subset of mapped points, the mapped points whose serial numbers of the corresponding point cloud data do not conform to the preset arrangement order as invalid points according to the preset arrangement direction includes: Traverse the mapping points in the subset of mapping points according to the preset arrangement order of the serial numbers of the corresponding point cloud data, and determine whether the coordinates corresponding to the currently traversed mapping point satisfy the target extreme value condition corresponding to the preset arrangement direction compared with the coordinates corresponding to the already traversed mapping points; If the target extreme value condition is not satisfied, determine that the currently traversed mapping point is an invalid point.

7. According to the method of any one of claims 1-6, wherein Before generating the first depth map from the point cloud data set collected for the target scene, the method further includes: Obtain an initial point cloud data set collected by a point cloud acquisition device for the target scene; Based on the spatial information of the target object in the target scene determined in advance, delete the point cloud data corresponding to the target object from the initial point cloud data set to obtain the point cloud data set.

8. A depth map generation device, comprising: A generation module, configured to generate a first depth map based on a point cloud data set collected for a target scene; A mapping module, configured to map the first depth map to the camera coordinate system of the target camera based on the parameters of the target camera to obtain a second depth map; A first determination module, configured to determine a mapping point set composed of mapping points corresponding to the point cloud data included in the point cloud data set in the second depth map; A second determination module, configured to determine invalid points from the mapping point set based on the coordinate order of the point cloud data in the point cloud data set, and delete the invalid points from the mapping point set; A third determination module, configured to determine a third depth map representing the depth true value in the camera coordinate system based on the mapping point set after deleting the invalid points.

9. A computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to be executed by a processor to implement the method according to any one of claims 1-7 above.

10. An electronic device, the electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Image stabilization method of depth image, electronic equipment and storage medium

    CN115619855A

  • RGBD camera obstacle detection method, device system and moving tool

    CN115705671A