Point cloud data processing method and apparatus, device, medium and computer program product

By mapping point cloud data into depth images and interpolation processing, the problems of low resolution and high noise in the prior art are solved, and efficient point cloud data upsampling and object detection are achieved.

WO2025086839A9PCT designated stage expired Publication Date: 2025-06-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112653
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-26
Filing Date
2024-08-16
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, when upsampling three-dimensional point cloud data into more dense point cloud data through deep learning models, it is impossible to fully describe the shape of the object and there is noise, and the calculation is large, which requires a large amount of training data.

Method used

The original point cloud data is mapped into depth images, interpolated to obtain extended depth images, and high-resolution point cloud data are generated through spatial position transformation, avoiding the dependence of deep learning models.

Benefits of technology

It improves the resolution of point cloud data and the accuracy of object detection, reduces noise, reduces calculation complexity, and does not rely on a large amount of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112653_26062025_PF_FP_ABST
    Figure CN2024112653_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A point cloud data processing method, executed by an electronic device, comprising: acquiring original point cloud data at a first resolution, wherein the original point cloud data comprises a plurality of feature points obtained by sampling a three-dimensional scene and spatial position information of the feature points, and the spatial position information of the feature points represents spatial positions corresponding to the feature points in the three-dimensional scene (201); on the basis of the spatial position information of each feature point in the original point cloud data, mapping each feature point to a mapping pixel point in a depth image at the first resolution, wherein a pixel value of the mapping pixel point represents depth information of the mapped feature point (202); on the basis of the depth information represented by the pixel value of the mapping pixel point in the depth image, interpolating the depth image to obtain an extended depth image at a second resolution (203); on the basis of a pixel position of an interpolation pixel point in the extended depth image, performing spatial position transformation on the interpolation pixel point to obtain spatial position information of a target feature point corresponding to the interpolation pixel point in the three-dimensional scene (204); and on the basis of the spatial position information of the target feature point and the original point cloud data, determining target point cloud data at the second resolution (205).
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud data processing method, device, equipment, medium and computer program product

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023114187756, filed on October 26, 2023, entitled “Point cloud data processing method and related equipment,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to a point cloud data processing method, apparatus, device, medium, and computer program product. Background Art

[0004] Point cloud data is a collection of point data in a certain coordinate system obtained by scanning an object or scene using a measuring instrument. In an exemplary application scenario, a 3D lidar scanner can be used to scan an object or scene to obtain point cloud data, which can be used to model the corresponding object or scene, or to analyze and process the corresponding object or scene. In actual applications, due to the limitations of insufficient performance of the measuring instrument, the original point cloud data obtained by scanning with the measuring instrument may be relatively sparse, that is, the resolution is low. Therefore, it is often difficult to achieve ideal results when modeling or analyzing objects or scenes based on the original point cloud data. Therefore, it is necessary to increase the resolution of the point cloud data. Dense point cloud data can improve the accuracy of modeling or analysis.

[0005] In current related technologies, three-dimensional point cloud data is usually upsampled into denser point cloud data through deep learning models. However, the distribution of points obtained by this method cannot fully describe the shape of the object. At the same time, there are also many noise points around the target, which is not conducive to improving the accuracy of object detection. In addition, this method requires a large amount of training data to help the deep learning model to reconstruct the data, and the model has many parameters and a large amount of calculation.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide a point cloud data processing method, apparatus, device, medium, and computer program product.

[0008] An embodiment of the present application provides a point cloud data processing method, which is executed by an electronic device and includes:

[0009] Acquire raw point cloud data at a first resolution, the raw point cloud data including a plurality of feature points sampled from a three-dimensional scene and spatial position information of the feature points, wherein the spatial position information of the feature points represents a spatial position corresponding to the feature points in the three-dimensional scene;

[0010] Based on the spatial position information of each feature point in the original point cloud data, each feature point is mapped to a mapping pixel point in the depth image at the first resolution, where the pixel value of the mapping pixel point represents the depth information of the mapped feature point;

[0011] interpolating the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution;

[0012] Performing spatial position transformation on the interpolated pixel point according to the pixel position of the interpolated pixel point in the extended depth image to obtain spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; and

[0013] Target point cloud data at the second resolution is determined based on the spatial position information of the target feature point and the original point cloud data.

[0014] Accordingly, an embodiment of the present application provides a point cloud data processing device, comprising:

[0015] an acquisition unit, configured to acquire raw point cloud data at a first resolution, the raw point cloud data including a plurality of feature points sampled from a three-dimensional scene and spatial position information of the feature points, wherein the spatial position information of the feature points represents a spatial position corresponding to the feature points in the three-dimensional scene;

[0016] a mapping unit, configured to map each feature point in the original point cloud data to a mapping pixel point in the depth image at the first resolution based on the spatial position information of each feature point, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point;

[0017] an interpolation unit, configured to interpolate the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution;

[0018] a position transformation unit, configured to perform spatial position transformation on the interpolation pixel point according to the pixel position of the interpolation pixel point in the extended depth image, to obtain spatial position information of the target feature point corresponding to the interpolation pixel point in the three-dimensional scene; and

[0019] A determining unit is configured to determine target point cloud data at the second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0020] An electronic device provided in an embodiment of the present application includes a processor and a memory, wherein the memory stores a plurality of instructions, and the processor loads the instructions to execute the steps in the point cloud data processing method provided in the embodiment of the present application.

[0021] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the point cloud data processing method provided in the embodiment of the present application are implemented.

[0022] In addition, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the point cloud data processing method provided in the embodiment of the present application.

[0023] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0025] FIG1 is a schematic diagram of a scene of a point cloud data processing method provided in an embodiment of the present application;

[0026] FIG2 is a flow chart of a point cloud data processing method provided in an embodiment of the present application;

[0027] FIG3 is an illustration of a point cloud data processing method provided in an embodiment of the present application;

[0028] FIG4 is another diagram illustrating the point cloud data processing method provided by an embodiment of the present application;

[0029] FIG5 is another diagram illustrating the point cloud data processing method provided by an embodiment of the present application;

[0030] FIG6 is another diagram illustrating the point cloud data processing method provided by an embodiment of the present application;

[0031] FIG7 is another diagram illustrating the point cloud data processing method provided in an embodiment of the present application;

[0032] FIG8 is another flow chart of the point cloud data processing method provided in an embodiment of the present application;

[0033] FIG9 is another flow chart of the point cloud data processing method provided in an embodiment of the present application;

[0034] FIG10 is a schematic diagram of the structure of a point cloud data processing device provided in an embodiment of the present application;

[0035] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] The present invention provides a method, apparatus, device, medium, and computer program product for processing point cloud data. The point cloud data processing apparatus can be integrated into an electronic device, such as a terminal or a server.

[0038] It is understandable that the point cloud data processing method of this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting the present application.

[0039] [Corrected 13.09.2024 in accordance with Rule 91] As shown in FIG1 , a point cloud data processing method is performed jointly by a terminal and a server as an example. The point cloud data processing system provided in an embodiment of the present application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, wherein the point cloud data processing device can be integrated into the server.

[0040] The server 11 can be configured to: receive raw point cloud data at a first resolution sent by the terminal 10, the raw point cloud data including multiple feature points sampled for a three-dimensional scene and spatial position information of the feature points, wherein the spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene; map each feature point to a mapping pixel in a depth image at the first resolution based on the spatial position information of each feature point in the raw point cloud data, wherein the pixel value of the mapping pixel represents the depth information of the mapped feature point; interpolate the depth image based on the depth information represented by the pixel value of the mapping pixel in the depth image to obtain an extended depth image at a second resolution; perform spatial position transformation on the interpolated pixel according to the pixel position of the interpolated pixel in the extended depth image to obtain spatial position information of a target feature point corresponding to the interpolated pixel in the three-dimensional scene; and determine target point cloud data at the second resolution based on the spatial position information of the target feature point and the raw point cloud data. The server 11 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0041] The terminal 10 can be used to collect raw point cloud data at a first resolution and send the raw point cloud data to the server 11. The raw point cloud data includes multiple feature points sampled from a three-dimensional scene, as well as spatial position information of the feature points. The spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene. The terminal 10 can include a mobile phone, a vehicle-mounted terminal, an aircraft, a tablet computer, a laptop computer, or a personal computer (PC). A client can also be provided on the terminal 10, which can be an application client or a browser client, etc.

[0042] The point cloud data processing method provided in the embodiments of the present application relates to the fields of transportation and mapping.

[0043] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0044] This embodiment will be described from the perspective of a point cloud data processing device. The point cloud data processing device can be integrated into an electronic device, which can be a server, a terminal or other device.

[0045] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0046] As shown in FIG2 , the specific process of the point cloud data processing method can be as follows:

[0047] 201. Obtain original point cloud data at a first resolution, where the original point cloud data includes multiple feature points sampled from a three-dimensional scene and spatial position information of the feature points, where the spatial position information of the feature points represents corresponding spatial positions of the feature points in the three-dimensional scene.

[0048] Raw point cloud data is a collection of point data in a certain coordinate system obtained by scanning a three-dimensional scene using a measuring instrument, such as a lidar. Raw point cloud data consists of multiple sampling points, each of which is also a feature point. The three-dimensional scene can be the environment in which the sampling takes place, specifically the spatial range in a real-world environment. Specifically, point cloud data is a massive collection of points that expresses the spatial distribution and surface characteristics of a target object within a common spatial reference system.

[0049] Specifically, the feature points are obtained by rotating the lidar to simultaneously emit multiple groups of laser samples along different vertical angles to the three-dimensional scene according to the horizontal azimuth resolution. Different groups of lasers correspond to different vertical channels, that is, to different vertical angles.

[0050] In this embodiment, the first resolution may include the horizontal azimuth resolution and vertical angular resolution corresponding to the lidar hardware used. Azimuth refers to the angle of a geographic location or object relative to a reference point. It is usually measured in a clockwise direction with respect to the north direction, ranging from 0° to 360°. For example, when the azimuth of an object is 0°, it means it is in the north direction; when the azimuth is 90°, it means it is in the east direction; when the azimuth is 180°, it means it is in the south direction; and when the azimuth is 270°, it means it is in the west direction.

[0051] In some embodiments, the spatial position information of a feature point may specifically be three-dimensional (3D) position information, which may include the specific coordinate values ​​of the feature point on the x-axis, y-axis, and z-axis in a three-dimensional coordinate system. Specifically, the spatial position information of a feature point may also include information such as the horizontal and vertical scanning angles of a laser radar corresponding to the feature point.

[0052] 202. Based on the spatial position information of each feature point in the original point cloud data, map each feature point to a mapping pixel point in the depth image at a first resolution, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point.

[0053] The depth information of a feature point can be the distance from the feature point to a measuring instrument, which is specifically the distance from the feature point to a laser radar. The depth information of the feature point can be determined based on the spatial position information of the feature point.

[0054] In some embodiments, each feature point can be mapped to a preset image to obtain a depth image at the first resolution, wherein the preset image can be a blank image in which the pixel value of each pixel is zero. The rows of the preset image can represent horizontal azimuth angles (i.e., horizontal angles), and the columns can represent vertical angles.

[0055] For each feature point in the original point cloud data, the spatial position information of that feature point can be transformed, converting the 3D coordinates into 2D coordinates, which can then be mapped onto the preset image. This position transformation process allows the unordered 3D point cloud data to be projected onto the ordered depth image. Each feature point in the original point cloud data corresponds one-to-one to each mapped pixel in the depth image.

[0056] Among them, a depth image is an image that maps point cloud data in a three-dimensional space to an image on a two-dimensional plane. It can also be called a distance image. The pixel value of a pixel in the image represents the depth information of the corresponding point. A pixel in a depth image can have one or more image channels (Ch, Channels). An image channel can correspond to one type of information. The number of image channels is related to the number of types of information. In a specific embodiment, a pixel in a depth image can have a depth channel and an attribute channel. The depth channel is also a distance measurement channel. The attribute channel may include an intensity channel or a color channel. The pixel value corresponding to the pixel in the distance measurement channel represents the distance from the corresponding feature point to the laser radar. The pixel value corresponding to the pixel in the intensity channel represents the intensity value of the laser beam point corresponding to the corresponding feature point. The pixel value corresponding to the pixel in the color channel represents the color information corresponding to the corresponding feature point.

[0057] In a specific scenario, the raw point cloud data can be a list structure of three-dimensional positions and laser beam point intensity values ​​arranged in the order of laser diode scanning, and the list may vary depending on the lidar model. The dimension of the raw point cloud data can be represented by N*D, where N represents the total number of laser beam points and D represents the number of features of the laser beam points, such as three-dimensional Cartesian coordinates and intensity values. The raw point cloud data is in an unordered form in terms of spatial coordinates. Typically, the raw point cloud data is a three-dimensional or higher-dimensional space composed of sparse and unordered data points, which increases the complexity of interpolation. The present application can reduce the complexity of processing such unordered data by projecting the raw point cloud data into a depth image, which can be spatially sorted using 2D coordinates of horizontal and vertical scanning angles, as shown in Figure 3.

[0058] In some embodiments, the size of the depth image can be H×V×C, where H is the horizontal azimuth angle (α), V is the vertical angle of the lidar (specifically, it can also be the number of vertical channels of the lidar), and C can include the distance measurement channel and the intensity channel. Specifically, the number of 3D points obtained by a single scan of the rotating lidar is determined by the number of vertical channels, the vertical resolution (Vres), the horizontal field of view (FOV), and the horizontal angular resolution (Hres).

[0059] Specifically, in this embodiment, the three-dimensional point cloud data is projected onto the range image in terms of azimuth and vertical angles, and the position transformation processing can be performed using the transformation equations in spherical coordinates, as shown in the following equations (1), (2), and (3): α=arctan(y,x) (2)

[0060] As shown in Figure 3, the spatial rectangular coordinate system takes the laser radar as the center, x, y, and z are the coordinate values ​​of the feature point on the horizontal, vertical, and vertical axes of the spatial rectangular coordinate system, respectively. R represents the distance from the feature point to the laser radar, that is, the depth information of the feature point, α represents the horizontal azimuth, and ω represents the vertical angle.

[0061] Specifically, the 3D point cloud data generated by the LiDAR (HDL-64E) has a dimension of 2048 × 64 × 4 Ch (X-Ch, Y-Ch, Z-Ch, Intensity-Ch), which can be converted into a 2048 × 64 × 2 Ch (α-Ch, ω-Ch) range image. Ch represents the channel, X-Ch, Y-Ch, Z-Ch, and Intensity-Ch represent the X, Y, Z, and intensity value channels, respectively. α-Ch and ω-Ch are the channels of the converted range image, representing the horizontal and vertical angles, respectively.

[0062] It should be noted that the "depth image" in this application is different from the image in the everyday sense, and it does not correspond to a visually visible picture; specifically, the depth image in this application uses a two-dimensional pixel array and the pixel value of each pixel to represent the information in the original point cloud for subsequent processing.

[0063] 203. Interpolate the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution.

[0064] Among them, the second resolution is greater than the first resolution. The second resolution may include horizontal azimuth resolution and vertical angular resolution. Specifically, this embodiment may perform upsampling (i.e., interpolation) in the horizontal direction and the elevation direction, or may perform upsampling in only one of the directions. For example, upsampling may be performed only in the elevation direction to increase the density of the vertical angle of the lidar channel, so that compared to the first resolution, the second resolution only changes the vertical angular resolution, and the horizontal azimuth resolution remains unchanged. For another example, upsampling may be performed only in the horizontal direction, so that compared to the first resolution, the second resolution only changes the horizontal azimuth resolution, and the vertical angular resolution remains unchanged.

[0065] The extended depth image is the interpolated depth image.

[0066] Optionally, in this embodiment, the step of “interpolating the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at the second resolution” may include:

[0067] Determine a pixel position of a pixel to be interpolated in the depth image, and determine adjacent pixels of the pixel to be interpolated from the mapped pixel based on the mapped pixel and the pixel position of the pixel to be interpolated in the depth image;

[0068] The depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated, and an extended depth image at the second resolution is obtained.

[0069] After the depth information of the pixel to be interpolated is determined, the depth information may be interpolated to the pixel to be interpolated in the depth image, thereby obtaining an interpolated pixel.

[0070] There are many ways to fuse the depth information represented by the pixel values ​​of adjacent pixels, such as weighted fusion.

[0071] In some embodiments, mapped pixels whose pixel distance to the pixel to be interpolated is less than a preset distance can be determined as neighboring pixels of the pixel to be interpolated, where the pixel distance is the distance between two pixels on the depth image. In other embodiments, the neighboring pixels of a pixel to be interpolated can be the k mapped pixels with existing depth information that are closest to the pixel to be interpolated, where k can be set based on actual conditions, such as 6.

[0072] Specifically, the rows of a depth image represent horizontal azimuths, while the columns represent vertical angles. Pixels in the same row have different horizontal azimuths but the same vertical angle; pixels in the same column have the same horizontal azimuths but different vertical angles. If upsampling is performed only in the elevation direction, the extended depth image has an additional row count, while the number of columns remains unchanged. Conversely, if upsampling is performed only in the horizontal direction, the extended depth image has an additional column count, while the number of rows remains unchanged.

[0073] Optionally, in this embodiment, the step of “fusing depth information represented by pixel values ​​of adjacent pixels to determine depth information of a pixel to be interpolated” may include:

[0074] Determine the pixel distance between each adjacent pixel and the pixel to be interpolated according to the pixel positions of each adjacent pixel and the pixel to be interpolated in the depth image;

[0075] Determine the weight information of adjacent pixels based on pixel distance;

[0076] Based on the weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0077] Among them, the greater the pixel distance, the smaller the weight information of adjacent pixel points; conversely, the smaller the pixel distance, the greater the weight information of adjacent pixel points.

[0078] In a specific example, the depth image is a distance image, and the depth image can be upsampled by the pixel distance weighted interpolation method to obtain an interpolated extended depth image. Specifically, the method can use the six adjacent pixels around the anchor point (i.e., the pixel to be interpolated) in the distance image obtained by LiDAR (Light Detection and Ranging) to perform weighted interpolation on them according to their pixel distance relative to the anchor point. As shown in Figure 4, it is an illustration corresponding to the interpolation using the weighted sum of adjacent pixels. The six interpolated neighboring pixels are marked as P1 to P6 starting from the upper left corner pixel, where the blank grid represents a blank area without a lidar point. The depth information of the pixel point P′ to be interpolated (specifically, the distance between the corresponding feature point and the lidar) is obtained by the weighted sum of the distances of the neighboring points with an attenuation factor, as shown in equations (4) and (5):

[0079] Among them, P i Represents adjacent pixels, P i The corresponding pixel value represents the distance between the corresponding feature point and the laser radar, W i Indicates P iWeight information, d i Indicates P i The pixel distance between P and P′.

[0080] Optionally, in this embodiment, before the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated”, the point cloud data processing method may further include:

[0081] When the depth information corresponding to the mapped pixel does not meet the preset depth condition, the mapped pixel is determined as an abnormal pixel;

[0082] The step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated” may include:

[0083] Detect whether each adjacent pixel is an abnormal pixel and obtain the detection result;

[0084] Based on the detection results and weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0085] The preset depth condition can be set according to actual conditions. For example, the preset depth condition can be set to define a pixel as a non-abnormal pixel when the depth value corresponding to the depth information of the pixel is within a preset range. Depth values ​​that are too large or too small are usually outliers. This embodiment can reduce the negative impact of abnormal pixels (i.e., outliers) on the interpolation process by performing anomaly detection on adjacent pixels.

[0086] Optionally, in this embodiment, the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on the detection results and weight information to determine depth information of the pixel to be interpolated” may include:

[0087] According to the detection results, the weight adjustment coefficient of adjacent pixels is set;

[0088] Based on the weight adjustment coefficient, update the weight information of adjacent pixels;

[0089] Based on the updated weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0090] When an adjacent pixel is detected as an abnormal pixel, the weight adjustment coefficient of the adjacent pixel can be set to 0 to reduce the influence of the abnormal pixel. When an adjacent pixel is detected as not an abnormal pixel, the weight adjustment coefficient of the adjacent pixel can be set to 1.

[0091] In some embodiments, the weight adjustment coefficient may be fused with the weight information of adjacent pixels to update the weight information of adjacent pixels. The fusion method may be multiplication or the like.

[0092] Optionally, in this embodiment, the step of “determining weight information of adjacent pixels based on pixel distance” may include:

[0093] Determining first weight information of adjacent pixels according to pixel distance;

[0094] Compare the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point;

[0095] Determining second weight information of adjacent pixels based on the relative depth information;

[0096] The first weight information and the second weight information are fused to obtain weight information of adjacent pixels.

[0097] There are many ways to fuse the first weight information and the second weight information, which are not limited in this embodiment. For example, the fusion method may be multiplication.

[0098] The reference depth information can be set according to actual conditions.

[0099] Optionally, in this embodiment, the step of “comparing the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point” may include:

[0100] Selecting reference depth information from the depth information represented by the pixel values ​​of each adjacent pixel point according to the size of the depth information represented by the pixel values ​​of each adjacent pixel point;

[0101] The depth information represented by the pixel values ​​of each adjacent pixel point is compared with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point.

[0102] In some embodiments, depth information with the smallest depth value among the depth information of adjacent pixels of the pixel to be interpolated may be selected as reference depth information.

[0103] Among them, the depth information represented by the pixel values ​​of adjacent pixels and the reference depth information are compared and processed. Specifically, the depth information represented by the pixel values ​​of adjacent pixels and the reference depth information can be differenced, and the obtained difference is the relative depth information corresponding to the adjacent pixels.

[0104] In a specific example, the depth image is a range image, and the depth image can be upsampled by pixel distance and range weighted interpolation to obtain an interpolated extended depth image. Specifically, each weight W can be readjusted according to the relative depth range between adjacent pixels. i To enhance upsampling, the relative depth range between adjacent pixels specifically refers to the difference in distance between the two pixels and the lidar. The closer the depth of the pixel (referring to the distance from the lidar) and the reference depth (specifically, the reference depth information), the greater the weight given in the interpolation, which is similar to assuming that adjacent pixels with a close distance are likely to be projected onto the same object. In addition, in the weighted sum process, the outliers (abnormal pixels) among the adjacent pixels can be added using the coefficient s i Skip, the outlier can be a pixel whose depth value corresponding to the depth information is zero or far greater than the preset threshold. For the outlier, the coefficient s i The value of is 0; otherwise, the coefficient s i The value of is 1. The corrected interpolation equations are shown in equations (6) and (7):

[0105] Among them, s i That is, the weight adjustment coefficient in the above embodiment, R i is the distance between the i-th adjacent pixel and the laser radar, R min is the minimum distance value among the distance information of each adjacent pixel point of the pixel to be interpolated, W i is the weight information of adjacent pixels. Specifically, the first weight information in the above embodiment, Specifically, it is the second weight information in the above embodiment, R i -R min That is, the relative depth information in the above embodiment.

[0106] By upsampling the depth image through the above-mentioned pixel distance and range weighted interpolation method, the influence of outliers can be minimized and the robustness of the proposed interpolation method can be improved.

[0107] 204. Perform spatial position transformation on the interpolated pixel points according to the pixel positions of the interpolated pixel points in the extended depth image to obtain spatial position information of target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0108] The pixel position of the interpolated pixel point in the extended depth image may include a horizontal azimuth angle and a vertical angle. The spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene may be three-dimensional position information, which may be specific coordinate values ​​of the x-axis, y-axis, and z-axis in a three-dimensional coordinate system.

[0109] Optionally, in this embodiment, the step of “performing spatial position transformation on the interpolated pixel points according to the pixel positions of the interpolated pixel points in the extended depth image to obtain spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene” may include:

[0110] Determine the horizontal azimuth and vertical angle corresponding to the interpolated pixel point based on the pixel position of the interpolated pixel point in the extended depth image;

[0111] According to the depth information, horizontal azimuth and vertical angle of the interpolated pixel points, the spatial position of the interpolated pixel points is transformed to obtain the spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0112] The depth information of the interpolated pixel point can specifically represent the distance value between the corresponding feature point and the target measuring instrument (specifically, the laser radar).

[0113] Specifically, in this embodiment, after obtaining the interpolated extended depth image, the obtained interpolated pixel points can be spatially transformed, that is, converted from two-dimensional data to three-dimensional data. Specifically, the two-dimensional distance image can be converted into a high-dimensional distance image with 4-Ch Cartesian coordinates (x, y, z) and intensity values. The values ​​of x, y, and z can be obtained from the horizontal azimuth angle, elevation angle (specifically, vertical angle), and distance value r, as shown in equations (8), (9), and (10): x = r cosωcosα (8) y = r cosωsinα (9) z = r sinω (10)

[0114] Among them, the azimuth angle (α) and elevation angle (ω) represent the coordinates of the interpolated pixel point in the two-dimensional range image, and r represents the distance value between the corresponding pixel point and the lidar, that is, the depth information.

[0115] 205. Determine target point cloud data at a second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0116] The target point cloud data is the original point cloud data after upsampling. The target point cloud data may include the original point cloud data and the point cloud data obtained after interpolation. The point cloud data obtained after interpolation may include the spatial position information corresponding to the target feature points.

[0117] In some embodiments, the raw point cloud data may include only the spatial position information of each feature point. In other embodiments, the raw point cloud data may include information in at least four dimensions, for example, three-dimensional spatial position information and attribute information. Attribute information is information that characterizes the surface characteristics of the target, such as reflection intensity information or color information.

[0118] Among them, the reflection intensity information is specifically the laser beam point intensity value corresponding to the corresponding feature point, wherein the laser beam point intensity value is related to the material of the object hit by the laser and the distance between the laser and the object; for example, the intensity value obtained when the laser hits a tree and that when it hits glass are generally different.

[0119] Specifically, different point cloud data acquisition principles correspond to different attribute information. For example, if the point cloud data acquisition device is based on laser measurement principles, the attribute information may be reflection intensity information. Reflection intensity information refers to the intensity of the laser radar pulse echo reflection corresponding to the data collected by the acquisition device at a certain point. Different objects reflect lasers to varying degrees, and reflection intensity information can be used to distinguish different objects. If the point cloud data acquisition device is based on photogrammetry principles, the attribute information may be color information.

[0120] Optionally, in this embodiment, the original point cloud data also includes attribute information corresponding to each feature point, the image channel of the depth image includes a depth channel and an attribute channel, the pixel value of the mapped pixel point in the depth channel represents the depth information of the corresponding feature point, and the pixel value of the mapped pixel point in the attribute channel represents the attribute information of the corresponding feature point;

[0121] Before the step of “determining target point cloud data at a second resolution based on the spatial position information of the target feature points and the original point cloud data”, the method may further include:

[0122] Based on the attribute information corresponding to the mapped pixel points in the attribute channel, the pixel points in the depth image under the attribute channel are interpolated to obtain the pixel value of the interpolated pixel point under the attribute channel. The pixel value of the interpolated pixel point under the attribute channel represents the attribute information of the corresponding target feature point in the three-dimensional scene;

[0123] The step of “determining target point cloud data at a second resolution based on the spatial position information of the target feature points and the original point cloud data” may include:

[0124] Target point cloud data at a second resolution is determined based on the spatial position information, attribute information, and original point cloud data of the target feature point.

[0125] Among them, if the original point cloud data also includes attribute information of each feature point, this embodiment can also interpolate the attribute information of the point cloud data. The interpolation process of the attribute information can specifically refer to the interpolation process of the depth information in the above embodiment, and will not be repeated here.

[0126] Among them, through depth information interpolation and attribute information interpolation, the spatial position information and attribute information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene can be obtained.

[0127] Optionally, in this embodiment, the step of “interpolating pixels in the attribute channel of the depth image based on the attribute information corresponding to the mapped pixel points in the attribute channel to obtain pixel values ​​of the interpolated pixel points in the attribute channel” may include:

[0128] Determine a pixel position of a pixel to be interpolated in the depth image, and determine adjacent pixels of the pixel to be interpolated from the mapped pixel based on the mapped pixel and the pixel position of the pixel to be interpolated in the depth image;

[0129] The attribute information corresponding to adjacent pixels is fused to determine the attribute information of the pixel to be interpolated.

[0130] Optionally, in this embodiment, the step of “fusing attribute information corresponding to adjacent pixels to determine attribute information of the pixel to be interpolated” may include:

[0131] Determine the pixel distance between each adjacent pixel and the pixel to be interpolated according to the pixel positions of each adjacent pixel and the pixel to be interpolated in the depth image;

[0132] Determine the weight information of adjacent pixels based on pixel distance;

[0133] Based on the weight information, the attribute information corresponding to adjacent pixels is fused to determine the attribute information of the pixel to be interpolated.

[0134] Optionally, in this embodiment, before the step of “fusing attribute information corresponding to adjacent pixels based on weight information to determine attribute information of the pixel to be interpolated”, the following steps may also be included:

[0135] Detecting the attribute information corresponding to the mapped pixel points in the depth image;

[0136] When it is detected that the attribute information corresponding to the mapped pixel point does not meet the preset attribute conditions, the mapped pixel point is determined as an abnormal pixel point;

[0137] The step of “fusing attribute information corresponding to adjacent pixels based on weight information to determine attribute information of the pixel to be interpolated” may include:

[0138] Detect whether each adjacent pixel is an abnormal pixel and obtain the detection result;

[0139] Based on the detection results and weight information, the attribute information corresponding to adjacent pixels is fused to determine the attribute information of the pixel to be interpolated.

[0140] Optionally, in this embodiment, the step of “fusing attribute information corresponding to adjacent pixels based on the detection results and weight information to determine attribute information of the pixel to be interpolated” may include:

[0141] According to the detection results, the weight adjustment coefficient of adjacent pixels is set;

[0142] Based on the weight adjustment coefficient, update the weight information of adjacent pixels;

[0143] Based on the updated weight information, the attribute information corresponding to adjacent pixels is fused to determine the attribute information of the pixel to be interpolated.

[0144] Optionally, in this embodiment, the step of “determining weight information of adjacent pixels based on pixel distance” may include:

[0145] Determining first weight information of adjacent pixels according to pixel distance;

[0146] Compare the attribute information corresponding to each adjacent pixel point with the reference attribute information to determine the relative attribute information corresponding to each adjacent pixel point;

[0147] Determining second weight information of adjacent pixels based on the relative attribute information;

[0148] The first weight information and the second weight information are fused to obtain weight information of adjacent pixels.

[0149] Optionally, in this embodiment, the step of “comparing the attribute information corresponding to each adjacent pixel with the reference attribute information to determine the relative attribute information corresponding to each adjacent pixel” may include:

[0150] Selecting reference attribute information from the attribute information corresponding to each adjacent pixel point according to the size of the attribute information corresponding to each adjacent pixel point;

[0151] The attribute information corresponding to each adjacent pixel point is compared with the reference attribute information to determine the relative attribute information corresponding to each adjacent pixel point.

[0152] In a specific embodiment, the point cloud data modeling effect obtained based on the interpolation method provided by this application can be compared with the ground truth modeling effect. As shown in Figure 5 (a), it is the original point cloud data of 16-line LiDAR, Figure 5 (b) is the ground truth of 64-line LiDAR, and Figures 5 (c) and (d) are both based on the method proposed by this application.

[0153] Specifically, the distance image is upsampled by the pixel distance weighted interpolation method in the above embodiment to obtain an interpolated extended distance image. This embodiment can convert the interpolation output in the two-dimensional extended distance image into a three-dimensional point cloud and perform a visual analysis of the interpolation effect, as shown in (c) in Figure 5, where the distance weighted interpolation result depicts a denser 3D output that can fully describe the shape of the object, but the noise data is amplified.

[0154] The pixel distance and range weighted interpolation method increases the size of the 2D range image and constructs a 3D point cloud from the interpolated extended range image. This effectively removes noise and further improves interpolation performance. As shown in Figure 5(d), the interpolation result clearly removes the noisy interpolated data, upsamples and enhances the point cloud data, and fully represents the shape of the high-resolution LiDAR point. Compared with Figure 5(b), the result obtained by the pixel distance and range weighted interpolation method is close to the 64-Ch (64-line) true value. The simulated model obtained after interpolation is highly consistent with the shape obtained by the 64-line LiDAR sampling. The pixel distance and range weighted interpolation method can reduce the negative impact of abnormal pixels (i.e., outliers) on the interpolation process by removing outlier neighboring pixels. Neighboring points whose depth information corresponds to zero or significantly greater than a preset threshold compared to the surrounding data are considered outliers. Specifically, zero or maximum value pixels in the 2D range image refer to LiDAR beam points that are not projected onto the object within the measurable range.

[0155] This application can combine distance images and point cloud data, and proposes a new point cloud data upsampling method, which can effectively improve the accuracy and robustness of three-dimensional object detection. The point cloud data processing method provided by this application includes pixel distance weighted interpolation and pixel distance and range weighted interpolation. The pixel distance weighted interpolation method takes into account the distance between pixels and assigns higher weights to pixels with closer distances, thereby more accurately describing the contours and edges of the object; the pixel distance and range weighted interpolation method further takes into account the distance and range between pixels, and can better handle complex object structures and shapes.

[0156] In this example, the proposed lidar upsampling method was used as input to a selected 3D object detection model. The detection accuracy was evaluated, verifying the effectiveness of the method. The PointPillar model, an effective and widely used network model for 3D object detection using lidar data, was used as the 3D model. This model was trained using this method by upsampling a dataset from baseline point cloud data.

[0157] The baseline point cloud dataset includes low-resolution (32-line) LiDAR point cloud data, which is generated by downsampling the KITTI 3D dataset of the 64-line LiDAR. Specifically, the 64-line LiDAR point cloud data can be converted into a range image to obtain a high-resolution LiDAR range image. Some lines are extracted from the 64-line LiDAR range image to obtain a low-resolution range image. For example, the range image corresponding to the 32-line LiDAR can be extracted and then converted into the point cloud data corresponding to the 32-line LiDAR. The point cloud data corresponding to the 32-line LiDAR is then upsampled by various methods, and the obtained processing results are input into the 3D target detection model (i.e., the PointPillar model) to evaluate the effectiveness of target detection.

[0158] The number of data files in the PointPillar model's training, validation, and test datasets were 7481, 1856, and 3769, respectively. The overall performance of the 3D detection task was evaluated based on the mAP (mean Average Precision) of the most common obstacles in autonomous driving environments—cars, pedestrians, and cyclists—based on easy, moderate, and difficult levels.

[0159] Specifically, the upsampling method proposed in this application and the currently widely used upsampling method were used to reconstruct a 32-line LiDAR into a 64-line LiDAR, and the reconstruction results were compared. The same baseline dataset was used for the upsampling process, and the PointPillar model was used for detection.

[0160] Table 1 compares the object detection performance of a pretrained PointPillar model (upsampling from 32 lines to 64 channels). It presents the mAP values ​​for pedestrian, cyclist, and car detection, comparing the detection enhancement achieved by different upsampling methods. An Intersection over Union (IoU) of 0.7 was used for evaluation. E, M, and H represent the difficulty levels of easy, medium, and hard, respectively.

[0161] Table 1

[0162] Among them, IoU (Intersection over Union) refers to the ratio of intersection to union, which is often used to measure the degree of overlap between two sets, especially in target detection and image segmentation to evaluate the accuracy of model prediction results.

[0163] As shown in Table 1, the 3D detection performance of existing methods such as nearest neighbor interpolation, bilinear interpolation, CNN-based ESPCN upsampling method, and Shan upsampling is poor. This 3D detection performance reflects the upsampling performance of the corresponding methods in low-resolution (16-line) lidar point clouds to high-resolution point clouds (64 lines). These methods can reconstruct denser point clouds with lower overall loss scores. However, they cannot accurately reconstruct the object shape, resulting in poor detection performance.

[0164] Specifically, the nearest neighbor interpolation method cannot describe complex shapes such as curves. Bilinear interpolation generates noise points around the target. When upsampling using a deep learning-based model, the point distribution cannot fully describe the object's shape, and there will also be a lot of noise around the target.

[0165] ESPCN (efficient sub-pixel convolutional neural network) is an upsampling method that extracts multiple features from aggregated low-resolution feature maps to reconstruct a high-resolution image output. Shan upsampling is an upsampling method that uses a residual-connected encoder-decoder architecture. Both ESPCN and Shan upsampling are deep learning-based upsampling methods that require a large amount of training data to help deep learning models perform image reconstruction. They also have high computational complexity and difficult parameter selection.

[0166] Compared with the upsampling method based on deep learning, the method of the present application has obvious advantages: in addition to being able to better describe the shape of the object and greatly reduce the generation of noise, the pixel distance weighted interpolation and pixel distance and range weighted interpolation methods used in this application are not based on deep learning, do not require too many parameters, and have obvious advantages in computational efficiency; moreover, the method of the present application does not require a large amount of training data and has better versatility and practicality.

[0167] In addition, Table 1 also shows a comparison between mid-level class detections. It can be clearly seen that compared with existing upsampling methods, the proposed solution (Case C) shows the best performance in each class detection, improving the performance of the 3D object detection task. The mAP for the overall class (level M) is 45.4%, which is approximately 7% higher than the baseline dataset.

[0168] In summary, this application proposes an upsampling method for low-resolution lidar, which can convert low-resolution point cloud data with coarse details into corresponding high-resolution point cloud data with fine details, thereby enhancing the three-dimensional target detection capability.

[0169] This application provides an effective upsampling method for low-resolution LiDAR point clouds. This method improves the accuracy of 3D object detection by reconstructing objects from sparse point cloud data into data with higher vertical angular resolution. This method can be used in multiple scenarios for autonomous driving, including but not limited to 3D object detection, road surface detection, terrain modeling, and environmental perception.

[0170] For example, autonomous vehicles need to detect obstacles, pedestrians, vehicles, and other objects in their surroundings in real time to ensure driving safety. The 3D point cloud upsampling method proposed in this application can increase the density and resolution of the point cloud, thereby improving the accuracy and robustness of object detection.

[0171] For example, autonomous vehicles need to monitor road conditions in real time, including road surface smoothness, levelness, and potholes, to ensure smooth and safe driving. The 3D point cloud upsampling method proposed in this application can increase the density and resolution of the point cloud, thereby improving the accuracy and robustness of road surface detection.

[0172] For example, autonomous vehicles need to construct a 3D terrain model of their surroundings in real time for path planning and navigation. Using the 3D point cloud upsampling method proposed in this application can increase the density and resolution of the point cloud, thereby improving the accuracy and robustness of terrain modeling.

[0173] For example, autonomous vehicles need to perceive the state and changes of their surroundings in real time, including weather, lighting, road conditions, and traffic conditions, in order to make intelligent decisions and control the system. The 3D point cloud upsampling method proposed in this application can increase the density and resolution of the point cloud, thereby improving the accuracy and robustness of environmental perception.

[0174] In a specific scenario, as shown in Figure 6, an example of 3D object detection is shown. Figure 6 (a) shows an RGB image of the 3D scene, Figure 6 (b) shows the 3D detection result using 64-line ground truth, and Figure 6 (c) shows the 3D detection result after upsampling using the point cloud data processing method proposed in this application. By comparison, it can be seen that the point cloud data processing method provided in this application can well describe the shape of objects and improve the accuracy of pedestrian detection.

[0175] In a specific scenario, as shown in Figure 7, another example of 3D target detection is shown. Figure 7(a) shows an RGB image of the 3D scene, Figure 7(b) shows the 3D detection result using 64-line true values, and Figure 7(c) shows the 3D detection result after upsampling using the point cloud data processing method proposed in this application. By comparison, it can be seen that the point cloud data processing method provided in this application can well describe the shape of objects, improving the accuracy of vehicle detection.

[0176] It should be noted that the upsampling method proposed in this application can also be combined with a 3D target detection model based on 2D range images, such as a range-sparse network (RSN), to build an end-to-end network.

[0177] This application proposes a simple and effective low-resolution lidar sparse point cloud upsampling method, which projects disordered three-dimensional point cloud data onto an ordered multi-channel range map image, and then reconstructs the range image into a dense three-dimensional point cloud. Specifically, as shown in Figure 8, the low-resolution three-dimensional point cloud data can be first converted into a low-resolution two-dimensional range image, and then upsampled on the two-dimensional range image to obtain a high-resolution two-dimensional range image. Then, the high-resolution two-dimensional range image is converted into high-resolution three-dimensional point cloud data, so that target detection can be performed on the high-resolution three-dimensional point cloud data through a three-dimensional target detection model. This can improve the accuracy of target detection and enhance the three-dimensional target detection capability of the low-resolution lidar.

[0178] As can be seen from the above, this embodiment can obtain original point cloud data at a first resolution, and the original point cloud data includes multiple feature points sampled for the three-dimensional scene, and spatial position information of the feature points, the spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene; based on the spatial position information of each feature point in the original point cloud data, each feature point is mapped to a mapping pixel point in the depth image at the first resolution, and the pixel value of the mapping pixel point represents the depth information of the mapped feature point; based on the depth information represented by the pixel value of the mapping pixel point in the depth image, the depth image is interpolated to obtain an extended depth image at a second resolution; according to the pixel position of the interpolated pixel point in the extended depth image, the interpolated pixel point is spatially transformed to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; based on the spatial position information of the target feature point and the original point cloud data, the target point cloud data at the second resolution is determined.

[0179] This application can first convert the point cloud data into the corresponding depth image, then perform interpolation based on the depth image, and finally convert the processed extended depth image into high-resolution point cloud data. This can effectively improve the resolution of the point cloud data, and the distribution of the obtained points can fully describe the shape of the object, which is conducive to improving the accuracy of object detection; in addition, it does not require a large amount of training data, the computational complexity is low, and it has good versatility and practicality.

[0180] According to the method described in the previous embodiment, the point cloud data processing device will be further described in detail below by taking the example of being specifically integrated in a server.

[0181] The present application provides a point cloud data processing method. As shown in FIG9 , the specific process of the point cloud data processing method may be as follows:

[0182] 901. The server obtains original point cloud data at a first resolution, where the original point cloud data includes multiple feature points sampled from a three-dimensional scene and spatial position information of the feature points, where the spatial position information of the feature points represents the corresponding spatial positions of the feature points in the three-dimensional scene.

[0183] Raw point cloud data is a collection of point data in a certain coordinate system obtained by scanning a 3D scene using a measuring instrument, such as a LiDAR. Raw point cloud data consists of multiple sampling points, each of which is also a feature point.

[0184] Specifically, the feature points are obtained by rotating the lidar to simultaneously emit multiple groups of laser samples along different vertical angles to the three-dimensional scene according to the horizontal azimuth resolution. Different groups of lasers correspond to different vertical channels, that is, to different vertical angles.

[0185] 902. The server maps each feature point to a mapping pixel point in the depth image at a first resolution based on the spatial position information of each feature point in the original point cloud data, where the pixel value of the mapping pixel point represents the depth information of the mapped feature point.

[0186] The depth information of a feature point can be the distance from the feature point to a measuring instrument, which is specifically the distance from the feature point to a laser radar. The depth information of the feature point can be determined based on the spatial position information of the feature point.

[0187] The preset image may be a blank image, in which the pixel value of each pixel is 0. The rows of the preset image may represent horizontal azimuth angles (ie, horizontal angles), and the columns may represent vertical angles.

[0188] For each feature point in the original point cloud data, the spatial position information of that feature point can be transformed, converting the 3D coordinates into 2D coordinates, which can then be mapped onto the preset image. This position transformation process allows the unordered 3D point cloud data to be projected onto the ordered depth image. Each feature point in the original point cloud data corresponds one-to-one to each mapped pixel in the depth image.

[0189] Among them, a depth image is an image that maps point cloud data in a three-dimensional space to an image on a two-dimensional plane. It can also be called a distance image. The pixel value of a pixel in the image represents the depth information of the corresponding point. A pixel in a depth image can have one or more image channels (Ch, Channels). An image channel can correspond to one type of information. The number of image channels is related to the number of types of information. In a specific embodiment, a pixel in a depth image can have a depth channel and an attribute channel. The depth channel is also a distance measurement channel. The attribute channel can include an intensity channel or a color channel, etc. The depth value corresponding to the pixel in the distance measurement channel represents the distance from the corresponding feature point to the laser radar, the depth value corresponding to the pixel in the intensity channel represents the intensity value of the laser beam point corresponding to the corresponding feature point, and the pixel value corresponding to the pixel in the color channel represents the color information corresponding to the corresponding feature point.

[0190] 903. The server determines a pixel position of the pixel to be interpolated in the depth image, and determines adjacent pixels of the pixel to be interpolated from the mapped pixels based on the mapped pixels and the pixel positions of the pixel to be interpolated in the depth image.

[0191] In some embodiments, mapped pixels whose pixel distance to the pixel to be interpolated is less than a preset distance can be determined as neighboring pixels of the pixel to be interpolated, where the pixel distance is the distance between two pixels on the depth image. In other embodiments, the neighboring pixels of a pixel to be interpolated can be the k mapped pixels with existing depth information that are closest to the pixel to be interpolated, where k can be set based on actual conditions, such as 6.

[0192] 904. The server fuses depth information represented by pixel values ​​of adjacent pixels to determine depth information of the pixel to be interpolated, and obtains an extended depth image at the second resolution.

[0193] Among them, the second resolution is greater than the first resolution. The second resolution may include horizontal azimuth resolution and vertical angular resolution. Specifically, this embodiment may perform upsampling (i.e., interpolation) in the horizontal direction and the elevation direction, or may perform upsampling in only one of the directions. For example, upsampling may be performed only in the elevation direction to increase the density of the vertical angle of the lidar channel, so that compared to the first resolution, the second resolution only changes the vertical angular resolution, and the horizontal azimuth resolution remains unchanged. For another example, upsampling may be performed only in the horizontal direction, so that compared to the first resolution, the second resolution only changes the horizontal azimuth resolution, and the vertical angular resolution remains unchanged.

[0194] The extended depth image is the interpolated depth image.

[0195] Optionally, in this embodiment, the step of “fusing depth information represented by pixel values ​​of adjacent pixels to determine depth information of a pixel to be interpolated” may include:

[0196] Determine the pixel distance between each adjacent pixel and the pixel to be interpolated according to the pixel positions of each adjacent pixel and the pixel to be interpolated in the depth image;

[0197] Determine the weight information of adjacent pixels based on pixel distance;

[0198] Based on the weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0199] Among them, the greater the pixel distance, the smaller the weight information of adjacent pixel points; conversely, the smaller the pixel distance, the greater the weight information of adjacent pixel points.

[0200] Optionally, in this embodiment, before the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated”, the point cloud data processing method may further include:

[0201] When the depth information corresponding to the mapped pixel does not meet the preset depth condition, the mapped pixel is determined as an abnormal pixel;

[0202] The step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated” may include:

[0203] Detect whether each adjacent pixel is an abnormal pixel and obtain the detection result;

[0204] Based on the detection results and weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0205] The preset depth condition can be set according to actual conditions. For example, the preset depth condition can be set to define a pixel as a non-abnormal pixel when the depth value corresponding to the depth information of the pixel is within a preset range. Depth values ​​that are too large or too small are usually outliers. This embodiment can reduce the negative impact of abnormal pixels (i.e., outliers) on the interpolation process by performing anomaly detection on adjacent pixels.

[0206] Optionally, in this embodiment, the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on the detection results and weight information to determine depth information of the pixel to be interpolated” may include:

[0207] According to the detection results, the weight adjustment coefficient of adjacent pixels is set;

[0208] Based on the weight adjustment coefficient, update the weight information of adjacent pixels;

[0209] Based on the updated weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0210] When an adjacent pixel is detected as an abnormal pixel, the weight adjustment coefficient of the adjacent pixel can be set to 0 to reduce the influence of the abnormal pixel. When an adjacent pixel is detected as not an abnormal pixel, the weight adjustment coefficient of the adjacent pixel can be set to 1.

[0211] In some embodiments, the weight adjustment coefficient may be fused with the weight information of adjacent pixels to update the weight information of adjacent pixels. The fusion method may be multiplication or the like.

[0212] Optionally, in this embodiment, the step of “determining weight information of adjacent pixels based on pixel distance” may include:

[0213] Determining first weight information of adjacent pixels according to pixel distance;

[0214] Compare the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point;

[0215] Determining second weight information of adjacent pixels based on the relative depth information;

[0216] The first weight information and the second weight information are fused to obtain weight information of adjacent pixels.

[0217] There are many ways to fuse the first weight information and the second weight information, which are not limited in this embodiment. For example, the fusion method may be multiplication.

[0218] The reference depth information can be set according to actual conditions.

[0219] Optionally, in this embodiment, the step of “comparing the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point” may include:

[0220] Selecting reference depth information from the depth information represented by the pixel values ​​of each adjacent pixel point according to the size of the depth information represented by the pixel values ​​of each adjacent pixel point;

[0221] The depth information represented by the pixel values ​​of each adjacent pixel point is compared with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point.

[0222] In some embodiments, depth information with the smallest depth value among the depth information of adjacent pixels of the pixel to be interpolated may be selected as reference depth information.

[0223] Among them, the depth information represented by the pixel values ​​of adjacent pixels and the reference depth information are compared and processed. Specifically, the depth information represented by the pixel values ​​of adjacent pixels and the reference depth information can be differenced, and the obtained difference is the relative depth information corresponding to the adjacent pixels.

[0224] 905. The server performs spatial position transformation on the interpolated pixel points according to the pixel positions of the interpolated pixel points in the extended depth image to obtain spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0225] Optionally, in this embodiment, the step of “performing spatial position transformation on the interpolated pixel points according to the pixel positions of the interpolated pixel points in the extended depth image to obtain spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene” may include:

[0226] Determine the horizontal azimuth and vertical angle corresponding to the interpolated pixel point based on the pixel position of the interpolated pixel point in the extended depth image;

[0227] According to the depth information, horizontal azimuth and vertical angle of the interpolated pixel points, the spatial position of the interpolated pixel points is transformed to obtain the spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0228] The depth information of the interpolated pixel point can specifically represent the distance value between the corresponding feature point and the target measuring instrument (specifically, the laser radar).

[0229] 906. The server determines target point cloud data at a second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0230] This application provides an effective upsampling method for low-resolution LiDAR point clouds. This method improves the accuracy of 3D object detection by reconstructing objects from sparse point cloud data into data with higher vertical angular resolution. This method can be used in multiple scenarios for autonomous driving, including but not limited to 3D object detection, road surface detection, terrain modeling, and environmental perception.

[0231] As can be seen from the above, this embodiment can obtain original point cloud data at a first resolution through a server, where the original point cloud data includes multiple feature points sampled for a three-dimensional scene, and spatial position information of the feature points, where the spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene; based on the spatial position information of each feature point in the original point cloud data, each feature point is mapped to a mapping pixel point in a depth image at the first resolution, and the pixel value of the mapping pixel point represents the depth information of the mapped feature point; the pixel position of the pixel point to be interpolated in the depth image is determined, and based on the mapping pixel point and the pixel position of the pixel point to be interpolated in the depth image, the adjacent pixels of the pixel point to be interpolated are determined from the mapping pixel points; the depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel point to be interpolated, and an extended depth image at a second resolution is obtained; according to the pixel position of the interpolated pixel point in the extended depth image, the interpolated pixel point is spatially transformed to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; based on the spatial position information of the target feature point and the original point cloud data, the target point cloud data at the second resolution is determined.

[0232] This application can first convert the point cloud data into the corresponding depth image, then perform interpolation based on the depth image, and finally convert the processed extended depth image into high-resolution point cloud data. This can effectively improve the resolution of the point cloud data, and the distribution of the obtained points can fully describe the shape of the object, which is conducive to improving the accuracy of object detection; in addition, it does not require a large amount of training data, the computational complexity is low, and it has good versatility and practicality.

[0233] To better implement the above method, an embodiment of the present application further provides a point cloud data processing device, as shown in FIG10 . The point cloud data processing device may include an acquisition unit 1001, a mapping unit 1002, an interpolation unit 1003, a position transformation unit 1004, and a determination unit 1005, as follows:

[0234] (1) Acquisition unit 1001;

[0235] An acquisition unit is used to acquire original point cloud data at a first resolution, where the original point cloud data includes multiple feature points sampled from a three-dimensional scene and spatial position information of the feature points, where the spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene.

[0236] (2) mapping unit 1002;

[0237] A mapping unit is used to map each feature point into a mapping pixel point in the depth image at a first resolution based on the spatial position information of each feature point in the original point cloud data, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point.

[0238] (3) interpolation unit 1003;

[0239] The interpolation unit is configured to interpolate the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution.

[0240] Optionally, in some embodiments of the present application, the interpolation unit may include a determination subunit and a fusion subunit, as follows:

[0241] A determination subunit is configured to determine a pixel position of a pixel to be interpolated in the depth image, and determine adjacent pixels of the pixel to be interpolated from the mapped pixel based on the mapped pixel and the pixel position of the pixel to be interpolated in the depth image;

[0242] The fusion subunit is used to fuse the depth information represented by the pixel values ​​of adjacent pixels to determine the depth information of the pixel to be interpolated and obtain the extended depth image at the second resolution.

[0243] Optionally, in some embodiments of the present application, the fusion sub-unit can be specifically used to determine the pixel distance between each adjacent pixel point and the pixel point to be interpolated based on the pixel positions of each adjacent pixel point and the pixel point to be interpolated in the depth image; determine the weight information of the adjacent pixel points based on the pixel distance; based on the weight information, fuse the depth information represented by the pixel values ​​of the adjacent pixel points to determine the depth information of the pixel point to be interpolated.

[0244] Optionally, in some embodiments of the present application, before the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated”, the point cloud data processing method may include:

[0245] When the depth information corresponding to the mapped pixel does not meet the preset depth condition, the mapped pixel is determined as an abnormal pixel;

[0246] The step of “fusing depth information represented by pixel values ​​of adjacent pixels based on weight information to determine depth information of the pixel to be interpolated” may include:

[0247] Detect whether each adjacent pixel is an abnormal pixel and obtain the detection result;

[0248] Based on the detection results and weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0249] Optionally, in some embodiments of the present application, the step of “fusing depth information represented by pixel values ​​of adjacent pixels based on the detection results and weight information to determine depth information of the pixel to be interpolated” may include:

[0250] According to the detection results, the weight adjustment coefficient of adjacent pixels is set;

[0251] Based on the weight adjustment coefficient, update the weight information of adjacent pixels;

[0252] Based on the updated weight information, the depth information represented by the pixel values ​​of adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

[0253] Optionally, in some embodiments of the present application, the step of “determining weight information of adjacent pixels based on pixel distance” may include:

[0254] Determining first weight information of adjacent pixels according to pixel distance;

[0255] Compare the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point;

[0256] Determining second weight information of adjacent pixels based on the relative depth information;

[0257] The first weight information and the second weight information are fused to obtain weight information of adjacent pixels.

[0258] Optionally, in some embodiments of the present application, the step of “comparing the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point” may include:

[0259] Selecting reference depth information from the depth information represented by the pixel values ​​of each adjacent pixel point according to the size of the depth information represented by the pixel values ​​of each adjacent pixel point;

[0260] The depth information represented by the pixel values ​​of each adjacent pixel point is compared with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point.

[0261] (4) Position transformation unit 1004;

[0262] The position transformation unit is used to perform spatial position transformation on the interpolated pixel points according to the pixel positions of the interpolated pixel points in the extended depth image, so as to obtain the spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0263] Optionally, in some embodiments of the present application, the position transformation unit may include an angle determination subunit and a transformation subunit, as follows:

[0264] An angle determination subunit, configured to determine a horizontal azimuth angle and a vertical angle corresponding to an interpolated pixel point based on a pixel position of the interpolated pixel point in the extended depth image;

[0265] The transformation subunit is used to perform spatial position transformation on the interpolated pixel points according to the depth information, horizontal azimuth and vertical angle corresponding to the interpolated pixel points, so as to obtain the spatial position information of the target feature points corresponding to the interpolated pixel points in the three-dimensional scene.

[0266] (5) determining unit 1005;

[0267] The determining unit is configured to determine target point cloud data at a second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0268] Optionally, in some embodiments of the present application, the original point cloud data further includes attribute information corresponding to each feature point, the image channel of the depth image includes a depth channel and an attribute channel, the pixel value of the mapped pixel point in the depth channel represents the depth information of the corresponding feature point, and the pixel value of the mapped pixel point in the attribute channel represents the attribute information of the corresponding feature point;

[0269] The point cloud data processing device may further include an attribute interpolation unit as follows:

[0270] An attribute interpolation unit is used to interpolate the pixels of the depth image under the attribute channel based on the attribute information corresponding to the mapped pixel points in the attribute channel to obtain the pixel value of the interpolated pixel point under the attribute channel, and the pixel value of the interpolated pixel point under the attribute channel represents the attribute information of the corresponding target feature point in the three-dimensional scene;

[0271] The determining unit may be specifically configured to determine target point cloud data at the second resolution based on spatial position information, attribute information, and original point cloud data of the target feature point.

[0272] As can be seen from the above, in this embodiment, the acquisition unit 1001 can acquire original point cloud data at a first resolution, where the original point cloud data includes multiple feature points sampled for a three-dimensional scene, and spatial position information of the feature points, where the spatial position information of the feature points represents the corresponding spatial position of the feature points in the three-dimensional scene; the mapping unit 1002 maps each feature point to a mapping pixel point in a depth image at the first resolution based on the spatial position information of each feature point in the original point cloud data, where the pixel value of the mapping pixel point represents the depth information of the mapped feature point; the interpolation unit 1003 interpolates the depth image based on the depth information represented by the pixel value of the mapping pixel point in the depth image to obtain an extended depth image at a second resolution; the position transformation unit 1004 performs spatial position transformation on the interpolated pixel point according to the pixel position of the interpolated pixel point in the extended depth image to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; the determination unit 1005 determines the target point cloud data at the second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0273] This application can first convert the point cloud data into the corresponding depth image, then perform interpolation based on the depth image, and finally convert the processed extended depth image into high-resolution point cloud data. This can effectively improve the resolution of the point cloud data, and the distribution of the obtained points can fully describe the shape of the object, which is conducive to improving the accuracy of object detection; in addition, it does not require a large amount of training data, the computational complexity is low, and it has good versatility and practicality.

[0274] The present application also provides an electronic device, as shown in FIG11 , which shows a schematic diagram of the structure of the electronic device involved in the present application embodiment. The electronic device may be a terminal or a server, etc. Specifically:

[0275] The electronic device may include components such as a processor 1101 with one or more processing cores, a memory 1102 with one or more computer-readable storage media, a power supply 1103, and an input unit 1104. Those skilled in the art will appreciate that the electronic device structure shown in FIG11 does not limit the electronic device and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0276] Processor 1101 is the control center of the electronic device, connecting all parts of the electronic device using various interfaces and circuits. It executes the various functions of the electronic device and processes data by running or executing software programs and / or modules stored in memory 1102 and accessing data stored in memory 1102. Optionally, processor 1101 may include one or more processing cores; preferably, processor 1101 may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1101.

[0277] The memory 1102 can be used to store software programs and modules. The processor 1101 executes various functional applications and data processing by running the software programs and modules stored in the memory 1102. The memory 1102 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1102 may also include a memory controller to provide the processor 1101 with access to the memory 1102.

[0278] The electronic device also includes a power supply 1103 for supplying power to various components. Preferably, the power supply 1103 can be logically connected to the processor 1101 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 1103 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0279] The electronic device may further include an input unit 1104, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0280] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1101 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 1102 according to the following instructions, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions as follows:

[0281] Acquire original point cloud data at a first resolution, the original point cloud data including a plurality of feature points sampled for a three-dimensional scene, and spatial position information of the feature points, wherein the spatial position information of the feature points represents the spatial position corresponding to the feature points in the three-dimensional scene; based on the spatial position information of each feature point in the original point cloud data, map each feature point to a mapping pixel point in a depth image at the first resolution, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point; interpolate the depth image based on the depth information represented by the pixel value of the mapping pixel point in the depth image to obtain an extended depth image at a second resolution; perform spatial position transformation on the interpolated pixel point according to the pixel position of the interpolated pixel point in the extended depth image to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; determine the target point cloud data at the second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0282] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0283] As can be seen from the above, this embodiment can obtain original point cloud data at a first resolution, wherein the original point cloud data includes multiple feature points sampled for a three-dimensional scene, and spatial position information of the feature points, wherein the spatial position information of the feature points represents the spatial position corresponding to the feature points in the three-dimensional scene; based on the spatial position information of each feature point in the original point cloud data, each feature point is respectively mapped to a mapping pixel point in the depth image at the first resolution, and the pixel value of the mapping pixel point represents the depth information of the mapped feature point; based on the depth information represented by the pixel value of the mapping pixel point in the depth image, the depth image is interpolated to obtain an extended depth image at a second resolution; according to the pixel position of the interpolated pixel point in the extended depth image, the interpolated pixel point is spatially transformed to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; based on the spatial position information of the target feature point and the original point cloud data, the target point cloud data at the second resolution is determined.

[0284] This application can first convert the point cloud data into the corresponding depth image, then perform interpolation based on the depth image, and finally convert the processed extended depth image into high-resolution point cloud data. This can effectively improve the resolution of the point cloud data, and the distribution of the obtained points can fully describe the shape of the object, which is conducive to improving the accuracy of object detection; in addition, it does not require a large amount of training data, the computational complexity is low, and it has good versatility and practicality.

[0285] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0286] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the point cloud data processing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0287] Acquire original point cloud data at a first resolution, the original point cloud data including a plurality of feature points sampled for a three-dimensional scene, and spatial position information of the feature points, wherein the spatial position information of the feature points represents the spatial position corresponding to the feature points in the three-dimensional scene; based on the spatial position information of each feature point in the original point cloud data, map each feature point to a mapping pixel point in a depth image at the first resolution, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point; interpolate the depth image based on the depth information represented by the pixel value of the mapping pixel point in the depth image to obtain an extended depth image at a second resolution; perform spatial position transformation on the interpolated pixel point according to the pixel position of the interpolated pixel point in the extended depth image to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; determine the target point cloud data at the second resolution based on the spatial position information of the target feature point and the original point cloud data.

[0288] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0289] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0290] Since the instructions stored in the computer-readable storage medium can execute the steps in any point cloud data processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any point cloud data processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0291] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the aforementioned point cloud data processing.

[0292] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0293] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A point cloud data processing method, performed by an electronic device, comprising: Acquire original point cloud data at a first resolution, the original point cloud data including a plurality of feature points sampled from a three-dimensional scene and spatial position information of the feature points, wherein the spatial position information of the feature points represents a spatial position corresponding to the feature points in the three-dimensional scene; Based on the spatial position information of each feature point in the original point cloud data, each feature point is mapped to a mapping pixel point in the depth image at the first resolution, and the pixel value of the mapping pixel point represents the depth information of the mapped feature point; Interpolate the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution; According to the pixel position of the interpolated pixel point in the extended depth image, the interpolated pixel point is spatially transformed to obtain the spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene; and Based on the spatial position information of the target feature point and the original point cloud data, target point cloud data at the second resolution is determined.

2. The method according to claim 1, interpolating the depth image based on the depth information represented by the pixel values ​​of the mapped pixels in the depth image to obtain an extended depth image at a second resolution, comprising: Determine a pixel position of a pixel to be interpolated in the depth image, and determine adjacent pixels of the pixel to be interpolated from the mapped pixels based on the mapped pixels and the pixel position of the pixel to be interpolated in the depth image; The depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel to be interpolated, and to obtain an extended depth image at a second resolution.

3. The method according to claim 2, wherein the depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel to be interpolated, comprising: Determine the pixel distance between each adjacent pixel and the pixel to be interpolated according to the pixel positions of each adjacent pixel and the pixel to be interpolated in the depth image; Determining weight information of the adjacent pixel points according to the pixel distance; Based on the weight information, the depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

4. The method according to claim 3, further comprising: When the depth information corresponding to the mapped pixel does not meet the preset depth condition, determining the mapped pixel as an abnormal pixel; The step of fusing the depth information represented by the pixel values ​​of the adjacent pixels based on the weight information to determine the depth information of the pixel to be interpolated includes: Detect whether each adjacent pixel point is an abnormal pixel point and obtain the detection result; Based on the detection result and the weight information, the depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

5. The method according to claim 4, wherein the depth information represented by the pixel values ​​of the adjacent pixels is fused based on the detection result and the weight information to determine the depth information of the pixel to be interpolated, comprising: According to the detection result, setting the weight adjustment coefficient of the adjacent pixel points; Based on the weight adjustment coefficient, updating the weight information of the adjacent pixels; Based on the updated weight information, the depth information represented by the pixel values ​​of the adjacent pixels is fused to determine the depth information of the pixel to be interpolated.

6. The method according to any one of claims 3 to 5, wherein determining the weight information of the adjacent pixel points according to the pixel distance comprises: Determining first weight information of the adjacent pixel points according to the pixel distance; The depth information represented by the pixel values ​​of each adjacent pixel point is compared with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point; Determining second weight information of the adjacent pixels based on the relative depth information; The first weight information and the second weight information are merged to obtain weight information of the adjacent pixels.

7. The method according to claim 6, wherein the step of comparing the depth information represented by the pixel values ​​of each adjacent pixel point with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point comprises: Selecting reference depth information from the depth information represented by the pixel values ​​of each adjacent pixel point according to the size of the depth information represented by the pixel values ​​of each adjacent pixel point; The depth information represented by the pixel values ​​of each adjacent pixel point is compared with the reference depth information to determine the relative depth information corresponding to each adjacent pixel point.

8. According to the method according to any one of claims 1 to 7, the step of performing spatial position transformation on the interpolated pixel point according to the pixel position of the interpolated pixel point in the extended depth image to obtain spatial position information of the target feature point corresponding to the interpolated pixel point in the three-dimensional scene comprises: Determine, based on a pixel position of the interpolated pixel in the extended depth image, a horizontal azimuth and a vertical angle corresponding to the interpolated pixel; According to the depth information, horizontal azimuth and vertical angle corresponding to the interpolation pixel point, the interpolation pixel point is spatially transformed to obtain the spatial position information of the target feature point corresponding to the interpolation pixel point in the three-dimensional scene.

9. According to the method according to any one of claims 1 to 8, the original point cloud data also includes attribute information corresponding to each feature point, the image channel of the depth image includes a depth channel and an attribute channel, the pixel value of the mapped pixel point in the depth channel represents the depth information of the corresponding feature point, and the pixel value of the mapped pixel point in the attribute channel represents the attribute information of the corresponding feature point; The method further comprises: Based on the attribute information corresponding to the mapped pixel points in the attribute channel, interpolate the pixel points in the depth image under the attribute channel to obtain the pixel value of the interpolated pixel point under the attribute channel, wherein the pixel value of the interpolated pixel point under the attribute channel represents the attribute information of the target feature point corresponding to it in the three-dimensional scene; The determining, based on the spatial position information of the target feature point and the original point cloud data, the target point cloud data at the second resolution includes: Based on the spatial position information, attribute information, and the original point cloud data of the target feature point, target point cloud data at the second resolution is determined.

10. A point cloud data processing device, comprising: an acquisition unit, configured to acquire original point cloud data at a first resolution, wherein the original point cloud data includes a plurality of feature points sampled from a three-dimensional scene and spatial position information of the feature points, wherein the spatial position information of the feature points represents a spatial position corresponding to the feature points in the three-dimensional scene; A mapping unit, configured to map each of the feature points into a mapping pixel point in the depth image at the first resolution based on the spatial position information of each feature point in the original point cloud data, wherein the pixel value of the mapping pixel point represents the depth information of the mapped feature point; an interpolation unit, configured to interpolate the depth image based on depth information represented by pixel values ​​of mapped pixels in the depth image to obtain an extended depth image at a second resolution; A position transformation unit, configured to perform spatial position transformation on the interpolation pixel point according to the pixel position of the interpolation pixel point in the extended depth image, so as to obtain spatial position information of a target feature point corresponding to the interpolation pixel point in the three-dimensional scene; and A determination unit is used to determine the target point cloud data at the second resolution based on the spatial position information of the target feature point and the original point cloud data.

11. An electronic device comprising a memory and a processor; the memory stores an application, and the processor is used to run the application in the memory to perform the operations in the point cloud data processing method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor to execute the steps in the point cloud data processing method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the steps in the point cloud data processing method according to any one of claims 1 to 9.