Point cloud enhancement method, medium, program product, controller, system and vehicle

By generating and reconstructing the predicted depth map of millimeter-wave radar point cloud data, the problem of insufficient point cloud quality is solved and the accuracy of target detection and semantic segmentation is improved.

CN120725880APending Publication Date: 2025-09-30BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900928.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The poor quality of millimeter-wave radar point clouds limits the accuracy of downstream perception tasks such as object detection and semantic segmentation.

Method used

By performing prediction processing on the original point cloud data, an original predicted depth map is generated, and point cloud reconstruction is performed based on the depth map, including point cloud mapping and filtering processing, to remove tailing point cloud data and improve point cloud quality.

Benefits of technology

This improves the quality of point clouds and enhances the accuracy of downstream perception tasks such as object detection and semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725880A_ABST
    Figure CN120725880A_ABST
Patent Text Reader

Abstract

The invention relates to a point cloud enhancement method, a medium, a program product, a controller, a system and a vehicle. The method comprises the following steps: performing prediction processing on original point cloud data to obtain an original prediction depth map corresponding to the original point cloud data; and performing point cloud reconstruction according to the original prediction depth map to obtain target point cloud data corresponding to the original point cloud data. According to the invention, point cloud enhancement is realized, the point cloud quality is improved, and the precision of downstream perception tasks such as target detection and semantic segmentation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to a point cloud enhancement method, medium, program product, controller, system and vehicle. Background Art

[0002] As a core sensor for Advanced Driving Assistance Systems (ADAS), millimeter-wave radar has attracted widespread attention from industry, universities, and research institutions due to its robustness to adverse weather conditions (rain, snow, fog, etc.). However, the relatively poor quality of millimeter-wave radar point clouds limits the accuracy of downstream perception tasks such as object detection and semantic segmentation. Therefore, improving the quality of millimeter-wave radar point clouds is a technical issue currently under research in the industry. Summary of the Invention

[0003] The embodiments of the present application provide a point cloud enhancement method, medium, program product, controller, system and vehicle, which realize point cloud enhancement, improve the quality of point cloud, and help improve the accuracy of downstream perception tasks such as target detection and semantic segmentation, so as to at least partially solve the above-mentioned technical problems.

[0004] In order to achieve the above-mentioned purpose, according to the first aspect of the present application, a point cloud enhancement method is provided, which includes: performing prediction processing on the original point cloud data to obtain an original predicted depth map corresponding to the original point cloud data; reconstructing the point cloud according to the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data.

[0005] Optionally, the point cloud reconstruction is performed according to the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data, including: performing point cloud mapping according to the original predicted depth map to obtain intermediate point cloud data; and filtering the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data.

[0006] Optionally, performing point cloud mapping according to the original predicted depth map to obtain intermediate point cloud data includes: adjusting the size of the original predicted depth map to obtain a target predicted depth map; and performing mapping processing on the target predicted depth map to obtain intermediate point cloud data.

[0007] Optionally, filtering the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data includes: identifying trailing point cloud data in the intermediate point cloud data; and removing the trailing point cloud data from the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data.

[0008] Optionally, the predictive processing of the original point cloud data to obtain the original predicted depth map corresponding to the original point cloud data includes: inputting the original point cloud data into the generation network so that the generation network outputs the original predicted depth map corresponding to the original point cloud data; wherein the generation network is used to perform predictive processing on the point cloud data to output the predicted depth map.

[0009] Optionally, the generation network includes an encoding module and a decoding module; inputting the original point cloud data into the generation network so that the generation network outputs an original predicted depth map corresponding to the original point cloud data includes: inputting the original point cloud data into the encoding module so that the encoding module outputs point cloud feature data of the original point cloud data; inputting the point cloud feature data into the decoding module so that the decoding module outputs the original predicted depth map corresponding to the original point cloud data.

[0010] Optionally, the encoding module includes P downsampling units, where P is an integer greater than 1; inputting the original point cloud data into the encoding module so that the encoding module outputs the point cloud feature data of the original point cloud data includes: inputting the original point cloud data into the first downsampling unit so that the first downsampling unit outputs the spatial feature data of the original point cloud data at the first scale; inputting the spatial feature data at the Kth scale into the K+1th downsampling unit so that the K+1th downsampling unit outputs the spatial feature data of the original point cloud data at the K+1th scale; K is a positive integer less than P; wherein the point cloud feature data of the original point cloud data includes the spatial feature data of the original point cloud data at P scales.

[0011] Optionally, the decoding module includes P upsampling units and an output unit; the inputting the point cloud feature data into the decoding module so that the decoding module outputs the original predicted depth map corresponding to the original point cloud data includes: inputting the spatial feature data at the Pth scale into the first upsampling unit so that the first upsampling unit outputs the original predicted data of the original point cloud data at the Pth scale; inputting the original predicted data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original predicted data of the original point cloud data at the Kth scale; and inputting the original predicted data at the 1st scale into the output unit so that the output unit outputs the original predicted depth map corresponding to the original point cloud data.

[0012] Optionally, the P-K+1th upsampling unit is jump-connected to the K+1th downsampling unit; inputting the original prediction data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale includes: inputting the original prediction data at the K+1th scale and the spatial feature data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale.

[0013] Optionally, the method further includes: constructing a training data set; wherein each training sample in the training data set includes a training point cloud data and a labeled depth map; and training the generation network according to the training data set.

[0014] Optionally, constructing the training data set includes: thinning the first point cloud data to obtain the training point cloud data; performing projection processing on the first point cloud data to obtain the label depth map; and establishing the training samples based on the training point cloud data and the label depth map.

[0015] Optionally, the thinning of the first point cloud data to obtain the training point cloud data includes: performing occlusion processing on the first point cloud data to obtain second point cloud data; performing mirror reflection processing on the second point cloud data to obtain third point cloud data; and determining the training point cloud data based on the third point cloud data.

[0016] Optionally, the occlusion processing of the first point cloud data to obtain the second point cloud data includes: determining the position code of each point cloud in the first point cloud data in the voxel grid; performing point cloud deletion on each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data; and determining the second point cloud data based on the remaining point clouds in the voxel grid.

[0017] Optionally, determining the position coding of each point cloud in the first point cloud data in the voxel grid includes: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a spherical coordinate system; and determining the position coding of each point cloud in the first point cloud data in the voxel grid based on the coordinate data of each point cloud in the first point cloud data in the spherical coordinate system.

[0018] Optionally, point cloud deletion is performed on each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data, including: obtaining the minimum point cloud distance of each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data; and for each voxel in the voxel grid, eliminating the point cloud in the voxel whose point cloud distance is greater than the minimum point cloud distance.

[0019] Optionally, determining the second point cloud data based on the point cloud remaining in the voxel grid includes: converting the point cloud remaining in the voxel grid from a spherical coordinate system to a Cartesian coordinate system to obtain the second point cloud data.

[0020] Optionally, performing mirror reflection processing on the second point cloud data to obtain third point cloud data includes: determining bottom edge data based on the second point cloud data; wherein the bottom edge data includes multiple bottom edges in a bounding box corresponding to the second point cloud data; determining the mirror reflection angle of each point cloud in the second point cloud data based on the bottom edge data; and determining the third point cloud data based on the mirror reflection angle of each point cloud in the second point cloud data.

[0021] Optionally, determining the mirror reflection angle of each point cloud in the second point cloud data based on the bottom edge data includes: for each point cloud in the second point cloud data, determining the bottom edge to which the point cloud belongs based on the distance between the point cloud and each bottom edge in the bottom edge data; for each point cloud in the second point cloud data, determining the angle between the projection component of the point cloud and the bottom edge to which it belongs, as well as the pitch angle of the point cloud; for each point cloud in the second point cloud data, determining the mirror reflection angle of the point cloud based on the angle and the pitch angle of the point cloud.

[0022] Optionally, determining the third point cloud data based on the mirror reflection angle of each point cloud in the second point cloud data includes: classifying the second point cloud data according to the mirror reflection angle of each point cloud in the second point cloud data to obtain at least one scattered spot set; and obtaining the third point cloud data based on sampling processing of each scattered spot set.

[0023] Optionally, the third point cloud data is obtained based on the sampling processing of each of the scattered spot sets, including: extracting at least one point cloud from each of the scattered spot sets; and merging the point clouds extracted from at least one of the scattered spot sets to obtain the third point cloud data.

[0024] Optionally, determining the training point cloud data based on the third point cloud data includes: performing expansion processing on the third point cloud data to obtain fourth point cloud data; and determining the training point cloud data based on the fourth point cloud data.

[0025] Optionally, the expanding the third point cloud data to obtain fourth point cloud data includes: filling the point cloud that meets the first condition in the second point cloud data into the third point cloud data to obtain fourth point cloud data.

[0026] Optionally, determining the training point cloud data based on the fourth point cloud data includes: performing analog-to-digital conversion on the fourth point cloud data to obtain complex ADC data; and performing spectrum conversion on the complex ADC data to obtain the training point cloud data.

[0027] Optionally, the method further includes: generating the first point cloud data based on dense point cloud data of at least one object.

[0028] Optionally, generating the first point cloud data based on the dense point cloud data of at least one object includes: performing position transformation on the dense point cloud data of at least one object according to a point cloud generation condition to obtain the first point cloud data.

[0029] Optionally, the point cloud generation condition includes at least one of the following: at least one of the objects does not overlap in space, and at least one of the objects is within the radar perception range.

[0030] Optionally, the projecting processing of the first point cloud data to obtain the label depth map includes: determining the projection coordinates of each point cloud in the first point cloud data; and filling the depth value of each point cloud in the first point cloud data into the initial depth map according to the projection coordinates of each point cloud in the first point cloud data to obtain the label depth map.

[0031] Optionally, determining the projection coordinates of each point cloud in the first point cloud data includes: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a projection coordinate system; and determining the projection coordinates of each point cloud in the first point cloud data based on the coordinate data of each point cloud in the first point cloud data in the projection coordinate system.

[0032] Optionally, filling the depth value of each point cloud in the first point cloud data into the initial depth map according to the projection coordinates of each point cloud in the first point cloud data to obtain the label depth map includes: determining the corresponding depth position of each point cloud in the first point cloud data in the initial depth map according to the projection coordinates of each point cloud in the first point cloud data; for each depth position in the initial depth map, filling the depth value of the point cloud corresponding to the depth position into the depth position to obtain the label depth map.

[0033] Optionally, establishing the training sample based on the training point cloud data and the label depth map includes: performing a first data preprocessing on the training point cloud data; performing a second data preprocessing on the label depth map; and combining the training point cloud data after the first data preprocessing and the label depth map after the second data preprocessing to obtain the training sample.

[0034] Optionally, the first data preprocessing includes at least one of the following: noise addition processing and first normalization processing.

[0035] Optionally, the second data preprocessing includes a second normalization process.

[0036] Optionally, the training of the generative network according to the training data set includes: inputting the training point cloud data into the generative network so that the generative network outputs a training predicted depth map; constructing a target loss value according to the label depth map and the training predicted depth map; and training the generative network according to the target loss value.

[0037] Optionally, constructing a target loss value based on the label depth map and the training predicted depth map includes: inputting the label depth map and the training predicted depth map into a decision network, respectively, so that the decision network outputs a first decision result and a second decision result, respectively; wherein the first decision result includes the true or false state of the label depth map, and the second decision result includes the true or false state of the training predicted depth map; and determining the target loss value based on the first decision result and the second decision result.

[0038] Optionally, the decision network includes a first feature extraction module, a second feature extraction module and a fully connected module; the label depth map and the training prediction depth map are respectively input into the decision network so that the decision network outputs a first decision result and a second decision result respectively, including: inputting the label depth map into the first feature extraction module so that the first feature extraction module outputs first label feature data; inputting the training point cloud data into the second feature extraction module so that the second feature extraction module outputs second feature data; inputting the first label feature data and the second feature data into the fully connected module so that the fully connected module outputs the first decision result; inputting the training prediction depth map into the first feature extraction module so that the first feature extraction module outputs first predicted feature data; inputting the training point cloud data into the second feature extraction module so that the second feature extraction module outputs second feature data; inputting the first predicted feature data and the second feature data into the fully connected module so that the fully connected module outputs the second decision result.

[0039] Optionally, determining the target loss value based on the first judgment result and the second judgment result includes: determining the first loss value of the judgment network based on the first judgment result and the second judgment result; determining the second loss value of the generation network based on the second judgment result; wherein the target loss value includes the first loss value and the second loss value.

[0040] Optionally, determining the first loss value of the decision network based on the first decision result and the second decision result includes: determining first decision expected data based on the first decision result and the first data coefficient; determining second decision expected data based on the second decision result and the second data coefficient; wherein the second data coefficient is smaller than the first data coefficient; and determining the first loss value of the decision network based on the first decision expected data and the second decision expected data.

[0041] Optionally, determining the second loss value of the generating network based on the second judgment result includes: determining first generated expected data based on the second judgment result and the first data coefficient; determining second generated expected data based on the label depth map and the training prediction depth map; and determining the second loss value of the generating network based on the first generated expected data and the second generated expected data.

[0042] Optionally, the training of the generation network according to the target loss value includes: alternately adjusting the parameters of the decision network and the parameters of the generation network according to the first loss value and the second loss value so that the first loss value and the second loss value converge.

[0043] According to a second aspect of the present application, a computer-readable storage medium is provided, on which a computer program or instructions are stored. When the computer program or instructions are executed by a processor, the point cloud enhancement method as described above is implemented.

[0044] According to a third aspect of the present application, a computer program product is provided, comprising a computer program or instructions, which implement the point cloud enhancement method described above when executed by a processor.

[0045] According to a fourth aspect of the present application, a controller is provided, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the point cloud enhancement method as described above is implemented.

[0046] According to a fifth aspect of the present application, a point cloud enhancement system is provided, comprising the controller as described above and a radar connected to the controller; wherein the radar is used to: collect original point cloud data.

[0047] Optionally, the radar includes a millimeter wave radar.

[0048] According to a sixth aspect of the present application, a vehicle is provided, comprising the electronic device as described above, or comprising the point cloud enhancement system as described above.

[0049] The technical solution provided in the embodiment of the present application enhances point cloud data based on a cross-modal approach. First, the original predicted depth map is obtained based on the original point cloud data, and then the point cloud is reconstructed based on the original predicted depth map to obtain the target point cloud data corresponding to the original point cloud data. By obtaining the original predicted depth map, more effective and accurate spatial features can be extracted from the original point cloud data. Point cloud reconstruction based on these spatial features can not only effectively remove clutter and noise in the original point cloud data, but also enrich the point cloud data to improve the sparse state of the point cloud data. Therefore, the embodiment of the present application realizes point cloud enhancement, improves point cloud quality, and helps to improve the accuracy of downstream perception tasks such as target detection and semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0051] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.

[0052] Figure 1 This is a flow chart of a point cloud enhancement method provided in an embodiment of the present application;

[0053] Figure 2 is a flowchart of another point cloud enhancement method provided in an embodiment of the present application;

[0054] Figure 3 is a schematic diagram of a voxel grid provided in an embodiment of the present application;

[0055] Figure 4 This is a schematic diagram of generating point cloud data provided by an embodiment of the present application;

[0056] Figure 5 is a schematic diagram of a depth map provided in an embodiment of the present application;

[0057] Figure 6 is a schematic diagram of a generation network provided in an embodiment of the present application;

[0058] Figure 7 is a schematic diagram of a decision network provided in an embodiment of the present application;

[0059] Figure 8 is a schematic diagram comparing the depth map prediction performance provided by an embodiment of the present application;

[0060] Figure 9 1 is a comparative schematic diagram of point cloud enhancement performance provided by an embodiment of the present application;

[0061] Figure 10 It is a schematic diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0063] As a core sensor in Advanced Driving Assistance Systems (ADAS), millimeter-wave radar has attracted widespread attention from industry, universities, and research institutions due to its robustness in adverse weather conditions (rain, snow, fog, etc.). However, compared to lidar, millimeter-wave radar's point cloud quality is relatively poor, with issues such as sparse point clouds, irregular point cloud distribution, and high levels of clutter and noise. This severely limits the accuracy of downstream perception tasks such as object detection and semantic segmentation.

[0064] In order to improve the point cloud quality of millimeter-wave radar, point cloud generation methods have been successively applied to automotive millimeter-wave radar, including the constant false alarm rate (CFAR) detection method, the angle super-resolution method based on array signal processing, the imaging method based on synthetic aperture radar (SAR), and the point cloud enhancement method based on deep learning.

[0065] Traditional CFAR detection methods rely heavily on accurate estimation of background noise distribution, and their performance degrades dramatically, particularly in extended target detection scenarios accompanied by clutter. Array signal processing methods can overcome the Rayleigh limit and achieve angular super-resolution, but they place high demands on the signal-to-noise ratio, number of snapshots, and correlation of target echoes. Especially in complex and changing detection scenarios, they face problems such as model mismatch, poor real-time performance, and the disappearance of vehicle outlines caused by specular reflections. Vehicle-mounted SAR imaging methods can overcome specular reflections, but they are primarily mounted on the side of the vehicle, making it difficult to address obstacle avoidance issues in front of or directly ahead of the vehicle while driving. Deep learning-based point cloud enhancement methods can significantly improve the density and quality of point clouds, promising angular super-resolution, mitigating specular reflections, and eliminating clutter. However, they rely heavily on the quality and scale of training data, and therefore face challenges such as high raw data production costs, poor labeling accuracy, and poor generalization capabilities.

[0066] In view of this, the embodiments of the present application provide a point cloud enhancement method, medium, program product, controller, system and vehicle. The point cloud enhancement method in the embodiments of the present application is a cross-modal radar point cloud imaging enhancement method. In order to solve technical problems such as the difficulty of collecting training data, low efficiency, and difficulty in obtaining high-quality learning labels, the embodiments of the present application combine simulation modeling to achieve the effect of data enhancement by generating large amounts of training data, thereby improving the generalization ability of the network model; in addition, in order to improve the prediction accuracy of the network model, the embodiments of the present application provide a radar point cloud enhancement model based on a super-resolution generative adversarial network. For the specific point cloud enhancement method, the structure and training of the network model, the construction of the training data set and its beneficial effects, please refer to the following embodiments, which will not be elaborated here.

[0067] According to the first aspect of the present application, an embodiment of the present application provides a point cloud enhancement method.

[0068] See also Figure 1 , Figure 1 This is a flow chart of a point cloud enhancement method provided by an embodiment of the present application. Figure 1 As shown, the point cloud enhancement method may include the following steps:

[0069] Step S100: performing prediction processing on the original point cloud data to obtain an original predicted depth map corresponding to the original point cloud data;

[0070] Step S200: reconstructing the point cloud according to the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data.

[0071] Raw point cloud data refers to point cloud data that requires point cloud enhancement. The raw point cloud data can be unprocessed point cloud data after collection, or it can be point cloud data that has been pre-processed by filtering after collection. The raw point cloud data can be collected by radar, which can be a millimeter-wave radar, a lidar, etc., and this embodiment of the application is not limited to this. The raw point cloud data may have problems such as sparse point clouds, irregular point cloud arrangement, and a lot of clutter and noise. In order to improve the quality of the raw point cloud data and further improve the accuracy of downstream perception tasks, the embodiment of the application enhances the raw point cloud data.

[0072] After acquiring the raw point cloud data, embodiments of the present application perform prediction processing on the raw point cloud data to obtain a predicted depth map corresponding to the raw point cloud data, namely, an original predicted depth map. The original predicted depth map may include depth values ​​of multiple point clouds. Based on the original predicted depth map, point cloud reconstruction may be further performed to obtain target point cloud data. The target point cloud data is the point cloud data obtained by performing point cloud enhancement on the raw point cloud data.

[0073] The technical solution provided in the embodiment of the present application enhances point cloud data based on a cross-modal approach. First, the original predicted depth map is obtained based on the original point cloud data, and then the point cloud is reconstructed based on the original predicted depth map to obtain the target point cloud data corresponding to the original point cloud data. By obtaining the original predicted depth map, more effective and accurate spatial features can be extracted from the original point cloud data. Point cloud reconstruction based on these spatial features can not only effectively remove clutter and noise in the original point cloud data, but also enrich the point cloud data to improve the sparse state of the point cloud data. Therefore, the embodiment of the present application realizes point cloud enhancement, improves point cloud quality, and helps to improve the accuracy of downstream perception tasks such as target detection and semantic segmentation.

[0074] In some embodiments, the above step S200 may include the following steps:

[0075] Step S210: performing point cloud mapping based on the original predicted depth map to obtain intermediate point cloud data;

[0076] Step S220: Filter the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data.

[0077] Since the depth map is two-dimensional data and the point cloud data is three-dimensional data, it is necessary to perform point cloud mapping based on the original point cloud data to obtain intermediate point cloud data. To further improve the quality of the point cloud data, the intermediate point cloud data is filtered to obtain the target point cloud data.

[0078] To facilitate the execution of point cloud mapping, in some embodiments, the above step S210 may include: adjusting the size of the original predicted depth map to obtain a target predicted depth map; and performing mapping processing on the target predicted depth map to obtain intermediate point cloud data.

[0079] Adjusting the size of the original predicted depth map includes enlarging the original predicted depth map. When obtaining the original predicted depth map, a smaller depth map size may be set for considerations such as computational complexity and processing accuracy. However, a smaller depth map size is not convenient for point cloud mapping. Therefore, it is necessary to adjust the size of the original predicted depth map, such as enlarging the original predicted depth map, to obtain the target predicted depth map.

[0080] The target predicted depth map is then mapped, such as mapping the target predicted depth map from two dimensions to three dimensions to obtain intermediate point cloud data. To improve the accuracy of the intermediate point cloud data, in some embodiments, the above-mentioned mapping of the target predicted depth map to obtain intermediate point cloud data includes: mapping the target predicted depth map to three-dimensional space according to the optical parameters of the depth camera to obtain intermediate point cloud data. The optical parameters of the depth camera can be pre-set. In some embodiments, the optical parameters of the depth camera include the internal parameters of the depth camera, which include but are not limited to at least one of the following: depth map size, focal length, principal point coordinates, measurement range, imaging plane, etc., which are not limited in the embodiments of the present application.

[0081] To further improve the quality of the point cloud data, in some embodiments, the above step S220 may include: identifying the trailing point cloud data in the intermediate point cloud data; removing the trailing point cloud data from the intermediate point cloud data to obtain the target point cloud data corresponding to the original point cloud data.

[0082] Tailing point cloud data refers to unrealistic point cloud data reconstructed in three-dimensional space due to noise or errors in the depth map. Usually, tailing point cloud data appears at the edge of the object's contour or in areas where the depth changes dramatically, forming a tailing phenomenon along the edge of the object or the depth gradient direction. Therefore, it is necessary to identify and filter out the tailing point cloud data in the intermediate point cloud data to obtain higher quality target point cloud data. The embodiment of the present application does not limit the identification method of the tailing point cloud data. In actual applications, it can be flexibly set according to the needs. For example, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm or its variant algorithms, such as the GDBSCAN (Generalized Density-Based Spatial Clustering of Applications with Noise) algorithm, the FDBSCAN (Fuzzy Density-Based Spatial Clustering of Applications with Noise) algorithm, etc. can be used.

[0083] In some embodiments, step S100 may include inputting raw point cloud data into a generative network, causing the generative network to output a raw predicted depth map corresponding to the raw point cloud data. The generative network is configured to perform predictive processing on the point cloud data to output a predicted depth map. The generative network is a deep learning-based network model. By performing predictive processing on the generative network to output a predicted depth map, the predicted depth map can be automatically generated, improving the efficiency and accuracy of generating the predicted depth map.

[0084] In some embodiments, the generation network includes an encoding module and a decoding module. Therefore, inputting the raw point cloud data into the generation network so that the generation network outputs a raw predicted depth map corresponding to the raw point cloud data includes: inputting the raw point cloud data into the encoding module so that the encoding module outputs point cloud feature data of the raw point cloud data; and inputting the point cloud feature data into the decoding module so that the decoding module outputs a raw predicted depth map corresponding to the raw point cloud data.

[0085] The encoding module (Encoder) is used to extract the spatial features of the point cloud data. The original point cloud data is input into the encoding module so that the encoding module can extract the spatial features of the original point cloud data, that is, the point cloud feature data of the original point cloud data. The decoding module (Decoder) is used to predict the depth map based on the spatial features extracted by the encoding module. The point cloud feature data of the original point cloud data is input into the decoding module so that the decoding module performs depth map prediction and outputs the original predicted depth map. The embodiment of the present application constructs a generation network through the encoding module and the decoding module. The network model structure is simple and easy to implement.

[0086] To extract more effective and accurate spatial features, in some embodiments, the encoding module includes P downsampling units. Based on this, the above-mentioned inputting the raw point cloud data into the encoding module so that the encoding module outputs point cloud feature data of the raw point cloud data includes: inputting the raw point cloud data into the first downsampling unit so that the first downsampling unit outputs spatial feature data of the raw point cloud data at the first scale; inputting the spatial feature data at the Kth scale into the K+1th downsampling unit so that the K+1th downsampling unit outputs spatial feature data of the raw point cloud data at the K+1th scale. The point cloud feature data of the raw point cloud data includes the spatial feature data of the raw point cloud data at P scales.

[0087] Wherein, P is an integer greater than 1; K is a positive integer less than P. In the embodiment of the present application, multiple downsampling units are provided in the encoding module of the generation network, and each downsampling unit is used to extract the spatial features of the point cloud data at a scale. Thus, the multi-scale spatial features of the point cloud data can be extracted through the encoding module. In the embodiment of the present application, the specific value of P is not limited. In actual application, it can be flexibly set according to the needs. For example, P can be set to an integer greater than or equal to 3, such as 3, 5, 6, or 7.

[0088] Taking P as an integer greater than or equal to 3 as an example, the original point cloud data is processed by the first downsampling unit in the encoding module to obtain the spatial feature data of the original point cloud data at the first scale; then the spatial feature data at the first scale is processed by the second downsampling unit in the encoding module to obtain the spatial feature data of the original point cloud data at the second scale... and so on, until the spatial feature data at the P-1th scale is processed by the Pth downsampling unit in the encoding module to obtain the spatial feature data of the original point cloud data at the Pth scale.

[0089] The embodiment of the present application does not limit the specific structure of the downsampling unit. In actual application, it can be flexibly configured according to the needs. For example, each downsampling unit includes but is not limited to at least one of the following: a two-dimensional convolution layer, a two-dimensional activation layer, a two-dimensional batch normalization layer, a two-dimensional maximum pooling layer, etc. In addition, the scales corresponding to the multiple downsampling units in the encoding module can be reduced in sequence. For example, from the first downsampling unit to the Pth downsampling unit, their corresponding scales are reduced in sequence. At a larger scale, global features corresponding to the overall structure of the point cloud data can be extracted; at a smaller scale, local features corresponding to the texture details of the point cloud data can be extracted.

[0090] Based on the network structure of the encoding module, in some embodiments, the decoding module includes P upsampling units and an output unit. Based on this, the above-mentioned inputting of point cloud feature data into the decoding module so that the decoding module outputs the original predicted depth map corresponding to the original point cloud data includes: inputting the spatial feature data at the Pth scale into the first upsampling unit so that the first upsampling unit outputs the original predicted data of the original point cloud data at the Pth scale; inputting the original predicted data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original predicted data of the original point cloud data at the Kth scale; and inputting the original predicted data at the 1st scale into the output unit so that the output unit outputs the original predicted depth map corresponding to the original point cloud data.

[0091] In an embodiment of the present application, multiple upsampling units and an output unit are set in the decoding module of the generation network, each upsampling unit is used to predict the predicted data of the point cloud data at a scale, and the output unit is used to perform prediction based on the predicted data output by the last upsampling unit to output a predicted depth map.

[0092] Taking P as an integer greater than or equal to 3 as an example, the spatial feature data at the P-th scale is processed by the first upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the P-th scale; the original prediction data at the P-th scale is processed by the second upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the P-1-th scale... and so on, until the original prediction data at the second scale is processed by the P-th upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the first scale.

[0093] The embodiments of the present application do not limit the specific structures of the upsampling unit and the output unit. In actual applications, they can be flexibly configured based on the needs. For example, each upsampling unit can be implemented based on bilinear interpolation; the output unit can include a convolution layer and an activation layer. The convolution layer is used to achieve feature dimension compression and angular super-resolution, and the activation layer is used to map the depth value to a specific interval, such as between 0 and 1. In addition, the scales corresponding to the multiple upsampling units in the decoding module can be increased in sequence. For example, from the first upsampling unit to the Pth upsampling unit, the scales corresponding to them are increased in sequence.

[0094] To further improve the prediction accuracy of the upsampling unit, in some embodiments, the P-K+1th upsampling unit is jump-connected to the K+1th downsampling unit. Based on this, the above-mentioned inputting the original prediction data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale includes: inputting the original prediction data at the K+1th scale and the spatial feature data at the K+1th scale into the P-K+1th upsampling unit so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale.

[0095] The P-K+1th upsampling unit is jump-connected to the K+1th downsampling unit, and the P-K+1th upsampling unit is used to process the original prediction data at the K+1th scale, and the K+1th downsampling unit is used to output the spatial feature data at the K+1th scale. Therefore, through the jump connection, the original prediction data at the K+1th scale and the spatial feature data at the K+1th scale can be input into the P-K+1th upsampling unit, and the original prediction data at the Kth scale can be output through the processing of the P-K+1th upsampling unit.

[0096] Taking P as an integer greater than or equal to 3 as an example, the spatial feature data at the P-th scale is processed by the first upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the P-th scale; the original prediction data at the P-th scale and the spatial feature data at the P-th scale are processed by the second upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the P-1-th scale... and so on, until the original prediction data at the second scale and the spatial feature data at the second scale are processed by the P-th upsampling unit in the decoding module to obtain the original prediction data of the original point cloud data at the first scale.

[0097] In some embodiments, the point cloud enhancement method further includes the following steps:

[0098] Step S010: constructing a training data set;

[0099] Step S020: training the generated network according to the training data set.

[0100] In an embodiment of the present application, a training data set is constructed, in which each training sample in the training data set includes a training point cloud data and a labeled depth map. There is a corresponding relationship between the training point cloud data and the labeled depth map in a training sample, that is, in a training sample, the labeled depth map is a depth map constructed for the training point cloud data. The generative network is trained based on the training data set to improve the prediction accuracy of the generative network for the depth map. In order to further improve and verify the prediction accuracy of the generative network, in some embodiments, the training data set can be divided into a learning data set and a verification data set. The learning data set is used to adjust the parameters of the generative network during the training process, and the verification data set is used to verify the prediction effect of the generative network during the training process.

[0101] To improve the accuracy of the generated network in predicting the depth map of the sparse point cloud data, in some embodiments, step S010 may include the following steps:

[0102] Step S011: Sparse the first point cloud data to obtain training point cloud data;

[0103] Step S012: Projecting the first point cloud data to obtain a label depth map;

[0104] Step S013: Create training samples based on the training point cloud data and the label depth map.

[0105] The first point cloud data is obtained from a database, can be based on radar acquisition, or can be generated by simulation, which is not limited in the present embodiment. To facilitate the generation of large quantities of first point cloud data, simulation generation can be used. Therefore, in some embodiments, the above point cloud enhancement method can also include the following steps:

[0106] Step S0110: Generate first point cloud data based on dense point cloud data of at least one object.

[0107] Among them, dense point cloud data of at least one object can be obtained from a point cloud database (Dense Point Cloud Library, DPCL). The point cloud database stores dense point cloud data of multiple objects. The embodiment of the present application does not limit the number of objects and object types corresponding to the dense point cloud data obtained in step S0110, and can be flexibly determined in combination with needs in actual applications. For example, the dense point cloud data of the objects to be obtained can be determined in combination with the application scenario of the generation network. When the generation network is applied to a vehicle-mounted scenario, dense point cloud data of objects such as vehicles, pedestrians, obstacles and / or road facilities can be obtained to generate the first point cloud data. In some embodiments, in order to improve the generalization and robustness of the generation network, multiple dense point cloud data of multiple objects can be obtained to generate the first point cloud data.

[0108] After obtaining dense point cloud data of at least one object, it is necessary to combine it into first point cloud data. Therefore, in some embodiments, the above step S0110 may include: performing position transformation on the dense point cloud data of at least one object according to the point cloud generation conditions to obtain the first point cloud data.

[0109] The present application does not limit the specific content of the point cloud generation conditions. In actual applications, they can be flexibly set based on actual needs. In some embodiments, the point cloud generation conditions include at least one of the following: at least one object does not overlap in space, and at least one object is within the radar perception range. The radar perception range can be determined based on the radar's physical and / or optical parameters, and the radar's physical and / or optical parameters can be pre-set when simulating and generating the first point cloud data.

[0110] According to the point cloud generation conditions, the position of the dense point cloud data of at least one object in space can be adjusted, such as by performing a position transformation such as rotation and / or translation, to obtain first point cloud data when the point cloud generation conditions are met. Based on this, in some embodiments, the above-mentioned position transformation of the dense point cloud data of at least one object according to the point cloud generation conditions to obtain the first point cloud data includes: determining the position transformation data of at least one object; performing a position transformation on the dense point cloud data of at least one object based on the position transformation data of at least one object; if the dense point cloud data of at least one object after the position transformation meets the point cloud generation conditions, using the dense point cloud data of at least one object after the position transformation as the first point cloud data; if the dense point cloud data of at least one object after the position transformation does not meet the point cloud generation conditions, re-execution from the step of determining the position transformation data of at least one object until the dense point cloud data of at least one object after the position transformation meets the point cloud generation conditions. The position transformation data of at least one object refers to the position data of the dense point cloud data in space, and the position transformation data includes but is not limited to rotation angle data and / or position offset data. In addition, the embodiment of the present application can obtain position transformation data based on random number generation. The position transformation data of at least one object relative to the radar observation center can be initialized based on random number generation, and then position transformation and condition judgment can be performed. If the point cloud generation conditions are not met, the position transformation data can be obtained again based on random number generation until the point cloud generation conditions are met.

[0111] In step S011, the first point cloud data is thinned to obtain training point cloud data. The present embodiment does not limit the specific method of thinning. To achieve a realistic and natural thinning effect, the thinning can be performed by simulating adverse effects of radar perception, such as occlusion effects and / or mirror reflection effects. Based on this, in some embodiments, the above step S011 may include the following steps:

[0112] Step S0111: performing occlusion processing on the first point cloud data to obtain second point cloud data;

[0113] Step S0112: performing mirror reflection processing on the second point cloud data to obtain third point cloud data;

[0114] Step S0113: Determine training point cloud data based on the third point cloud data.

[0115] Occlusion processing is used to simulate the occlusion effect in radar reflection characteristics. To facilitate occlusion processing, in some embodiments, step S0111 may include: determining a position code for each point cloud in the first point cloud data in a voxel grid; performing point cloud pruning on each voxel in the voxel grid based on the position code for each point cloud in the first point cloud data; and determining second point cloud data based on the remaining point clouds in the voxel grid.

[0116] A voxel grid is a data structure for representing a three-dimensional space, which divides the three-dimensional space into multiple cubes, i.e., multiple voxels. Each point cloud in the first point cloud data can be mapped to a voxel grid, so as to grasp the distribution state of the point cloud in the first point cloud data, thereby facilitating occlusion processing. In an embodiment of the present application, a voxel grid can be constructed based on simulation parameters. In some embodiments, the above-mentioned point cloud enhancement method further includes: determining radar resolution data based on radar simulation parameters; and constructing a voxel grid based on radar resolution data. The embodiment of the present application does not limit the specific contents of radar simulation parameters and radar resolution data. In some embodiments, radar simulation parameters include but are not limited to at least one of the following: operating frequency band, frequency modulation slope, sampling rate, number of sampling points, arrangement data of array antenna (such as the number of horizontal antennas and the number of vertical antennas of the array antenna), etc.; radar resolution data include but are not limited to at least one of the following: distance resolution, horizontal angle resolution, vertical angle resolution, etc.

[0117] When mapping the point cloud to the voxel grid, the position code of the point cloud in the voxel grid can be obtained. The position code indicates the voxel where the point cloud is located, so that the point cloud in each voxel can be obtained for point cloud deletion to simulate the radar occlusion effect.

[0118] The first point cloud data may be point cloud data based on a Cartesian coordinate system, while the voxel grid is generally based on a spherical coordinate system. Therefore, in some embodiments, the above-mentioned determination of the position coding of each point cloud in the first point cloud data in the voxel grid includes: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a spherical coordinate system; and determining the position coding of each point cloud in the first point cloud data in the voxel grid according to the coordinate data of each point cloud in the first point cloud data in the spherical coordinate system.

[0119] The embodiment of the present application performs point cloud deletion for each voxel in the voxel grid, so that the radar obstruction effect can be simulated more finely and accurately. Based on this, in some embodiments, the above-mentioned point cloud deletion for each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data includes: obtaining the minimum point cloud distance of each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data; for each voxel in the voxel grid, eliminating the point cloud in the voxel whose point cloud distance is greater than the minimum point cloud distance. The minimum point cloud distance refers to the minimum distance of the point cloud in the voxel grid. The embodiment of the present application eliminates the point cloud in the voxel whose point cloud distance is greater than the minimum point cloud distance for each voxel, that is, only retains the point cloud corresponding to the minimum point cloud distance. Then, the second point cloud data is determined based on the remaining point clouds in the voxel grid. Since the point cloud is converted from the Cartesian coordinate system to the spherical coordinate system when point cloud deletion is performed, in order to restore the coordinate system type of the second point cloud data, in some embodiments, the above-mentioned determination of the second point cloud data based on the remaining point cloud in the voxel grid includes: converting the remaining point cloud in the voxel grid from the spherical coordinate system to the Cartesian coordinate system to obtain the second point cloud data.

[0120] Specular reflection processing is used to simulate the specular reflection effect of radar. To facilitate specular reflection processing, the specular reflection angle of each point cloud can be calculated. Based on this, in some embodiments, step S0112 may include: determining base edge data based on the second point cloud data; determining the specular reflection angle of each point cloud in the second point cloud data based on the base edge data; and determining the third point cloud data based on the specular reflection angle of each point cloud in the second point cloud data.

[0121] The bottom edge data is used to indicate the bottom edge of the second point cloud data, and the bottom edge data includes multiple bottom edges in the bounding box corresponding to the second point cloud data. The bounding box of the second point cloud data may include multiple edges, and the bottom edge is an edge located in the bottom edge plane of the second point cloud data. Depending on the shape or posture of the three-dimensional object, the number of bottom edges in the bottom edge data may vary. In practical applications, for ease of processing, the three-dimensional object can be divided into a cuboid, which at least includes the three-dimensional object, so that the bottom edge data can uniformly include four bottom edges.

[0122] In some embodiments, the above-mentioned determination of the mirror reflection angle of each point cloud in the second point cloud data based on the bottom edge data includes: for each point cloud in the second point cloud data, determining the bottom edge to which the point cloud belongs based on the distance between the point cloud and each bottom edge in the bottom edge data; for each point cloud in the second point cloud data, determining the angle between the projection component of the point cloud and the bottom edge to which it belongs, as well as the pitch angle of the point cloud; for each point cloud in the second point cloud data, determining the mirror reflection angle of the point cloud based on the angle and pitch angle of the point cloud. Among them, the projection component of the point cloud refers to the component of the point cloud in the bottom edge plane of the second point cloud data. The embodiment of the present application calculates the angle between the projection component of each point cloud and the bottom edge to which it belongs, and calculates the pitch angle of the point cloud. For the specific calculation method of the angle and the pitch angle, please refer to the following embodiment, which will not be elaborated here. Afterwards, based on the included angle and the pitch angle of the point cloud, the specular reflection angle of the point cloud may be further determined. For example, the specular reflection angle of each point cloud may be the maximum value of the included angle and the pitch angle of the point cloud.

[0123] Scatter speckles are formed when radar waves are scattered in non-specular reflection directions due to irregularities or edges on the target surface. The intensity and distribution of scatter speckles are affected by the specular reflection effect. Therefore, in some embodiments, determining the third point cloud data based on the specular reflection angle of each point cloud in the second point cloud data includes: classifying the second point cloud data based on the specular reflection angle of each point cloud in the second point cloud data to obtain at least one scatter speckle set; and sampling each scatter speckle set to obtain the third point cloud data. By classifying the second point cloud data based on the specular reflection angle to obtain at least one scatter speckle set, the intensity and distribution of scatter speckles in different scatter speckle sets vary, reflecting the variations in the radar's specular reflection effect. Sampling each scatter speckle set individually allows for more precise and accurate simulation of the radar's specular reflection effect.

[0124] Exemplarily, the above-described classification of the second point cloud data based on the specular reflection angle of each point cloud in the second point cloud data to obtain at least one scattered spot set includes: for each point cloud in the second point cloud data, if the specular reflection angle of the point cloud is less than or equal to a first angle threshold, adding the point cloud to the first scatter spot set; for each point cloud in the second point cloud data, if the specular reflection angle of the point cloud is greater than the first angle threshold and less than or equal to a second angle threshold, adding the point cloud to the second scatter spot set; for each point cloud in the second point cloud data, if the specular reflection angle of the point cloud is greater than the second angle threshold, adding the point cloud to the third scatter spot set. The second angle threshold is greater than the first angle threshold, and the first and second angle thresholds may be preset. The embodiments of the present application do not limit the specific values ​​of the first and second angle thresholds. In actual applications, they can be flexibly set based on needs. For example, the first angle threshold can be 0° or 1°, and the second angle threshold can be 14°, 15°, or 16°. Furthermore, in practical applications, other classification angle thresholds may be set. For example, a third angle threshold greater than the second angle threshold may be set, such as 24°, 25°, or 26°. If the specular reflection angle of a point cloud is greater than the second angle threshold and less than or equal to the third angle threshold, the point cloud is added to the third scattering spot set. Of course, in practical applications, more scattering spot sets may be used, such as four, five, or six, and this is not limited in this embodiment of the present application.

[0125] For each scattered spot set obtained by classification, the scattered spot set is sampled and processed to construct third point cloud data. Based on this, in some embodiments, the third point cloud data is obtained according to the sampling process of each scattered spot set, including: extracting at least one point cloud from each scattered spot set; merging the point clouds extracted from at least one scattered spot set to obtain third point cloud data. The number of point clouds extracted for different scattered spot sets can be the same or different, and this embodiment of the present application is not limited to this. For example, for the three scattered spot sets obtained by the above division, a first number of point clouds can be extracted from the first scattered spot set, a second number of point clouds can be extracted from the second scattered spot set, and a third number of point clouds can be extracted from the third scattered spot set, where the second number is greater than the first number, and the first number is greater than the third number. The third point cloud data can be obtained by merging the point clouds extracted from at least one scattered spot set.

[0126] In some embodiments, the above step S0113 may include the following steps: performing expansion processing on the third point cloud data to obtain fourth point cloud data; and determining training point cloud data based on the fourth point cloud data.

[0127] By expanding the data, the point cloud data is enriched to avoid the point cloud data being too sparse, which affects the training effect of the generated network. In some embodiments, the above expansion processing of the third point cloud data to obtain the fourth point cloud data includes: filling the point cloud that meets the first condition in the second point cloud data into the third point cloud data to obtain the fourth point cloud data. The embodiment of the present application extracts the point cloud that meets the first condition from the second point cloud data, and fills the point cloud that meets the first condition into the third point cloud data. Among them, in order to simulate the real sparse effect, in some embodiments, the first condition includes that the distance between the point cloud data and the third point cloud data is less than a first distance threshold. The embodiment of the present application does not limit the specific value of the first distance threshold. In actual application, it can be flexibly set according to the needs. For example, the first distance threshold can be set to 0.2m, 0.3m or 0.4m, etc. Exemplarily, the point cloud in the second point cloud data whose distance to the third point cloud data is less than the first distance threshold can be selected based on the ellipsoid rule.

[0128] In some embodiments, determining the training point cloud data based on the fourth point cloud data includes: performing analog-to-digital conversion on the fourth point cloud data to obtain complex ADC data; and performing spectrum conversion on the complex ADC data to obtain training point cloud data.

[0129] The analog-to-digital converter (ADC) digitizes the point cloud data based on the radar imaging system. In some embodiments, the analog-to-digital conversion of the fourth point cloud data to obtain complex ADC data includes: determining a snapshot signal of each point cloud in the fourth point cloud data; and superimposing the snapshot signals of multiple point clouds in the fourth point cloud data to obtain complex ADC data. The snapshot signal refers to the original echo signal captured by the radar in a short period of time. The embodiment of the present application obtains the snapshot signal of each point cloud in the fourth point cloud data, and then superimposes the snapshot signals of multiple point clouds in the fourth point cloud data to obtain complex ADC data. For the calculation and superposition method of the snapshot signal, please refer to the following embodiment, which will not be described here.

[0130] The training point cloud data is a radar data cube (RDC). In order to obtain the training point cloud data, the complex ADC data needs to be spectrally converted. The embodiment of the present application does not limit the specific method of spectrum conversion, and it can be flexibly set in combination with the needs in actual applications. In some embodiments, the above-mentioned spectrum conversion of the complex ADC data to obtain the training point cloud data includes: performing a fast Fourier transform (FFT) on the complex ADC data to obtain the training point cloud data. Among them, in order to integrate spatial features of multiple dimensions, in some embodiments, the fast Fourier transform includes at least one of the following: distance fast Fourier transform, Doppler fast Fourier transform, horizontal angle fast Fourier transform, and vertical angle fast Fourier transform. For other introductions to the training point cloud data, please refer to the following embodiments, which will not be elaborated here.

[0131] In step S012, the first point cloud data is projected to obtain a label depth map. The projection process in the embodiment of the present application can project the three-dimensional first point cloud data into two dimensions, and then obtain the label depth map through depth filling. Based on this, in some embodiments, the above step S012 may include the following steps:

[0132] Step S0121: determining the projection coordinates of each point cloud in the first point cloud data;

[0133] Step S0122: filling the depth value of each point cloud in the first point cloud data into the initial depth map according to the projection coordinates of each point cloud in the first point cloud data to obtain a label depth map.

[0134] Projection coordinates refer to the two-dimensional coordinates of a point cloud. Each point cloud in the first point cloud data can be projected onto a preset projection plane to obtain the projection coordinates of each point cloud. Based on this, in some embodiments, step S0121 may include: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a projection coordinate system; and determining the projection coordinates of each point cloud in the first point cloud data based on the coordinate data of each point cloud in the projection coordinate system.

[0135] In practical applications, the imaging plane of the depth camera can be selected as the projection plane, and each point cloud in the first point cloud data can be converted from the Cartesian coordinate system to the projection coordinate system according to the physical parameters of the depth camera (such as installation position, posture, etc.) and the selection of the projection plane. The coordinate system conversion process may involve geometric transformations such as rotation and / or translation, and the coordinate system conversion can be achieved by matrix multiplication. Then, based on the coordinate data of each point cloud in the first point cloud data in the projection coordinate system, the projection coordinates of each point cloud in the first point cloud data in the projection plane are determined. Among them, when obtaining the projection coordinates of each point cloud, the optical parameters of the depth camera, such as internal parameters such as focal length, can be combined. For other introductions and explanations on the calculation of projection coordinates, please refer to the following embodiments, which will not be elaborated here.

[0136] Based on the projection coordinates, the position of each point cloud in the first point cloud data in the depth map can be clarified, thereby facilitating the filling of depth values ​​in the depth map. Based on this, in some embodiments, the above step S0122 may include: determining the corresponding depth position of each point cloud in the first point cloud data in the initial depth map based on the projection coordinates of each point cloud in the first point cloud data; for each depth position in the initial depth map, filling the depth value of the point cloud corresponding to the depth position into the depth position to obtain a label depth map.

[0137] The initial depth map can be generated by initializing the depth map based on the optical parameters of the depth camera, such as image size, measurement range and other internal parameters. The depth map can be a two-dimensional array or matrix. Based on the projection coordinates of each point cloud in the first point cloud data, the depth position corresponding to each point cloud in the initial depth map can be determined, and then for each depth position in the initial depth map, the depth value of the point cloud corresponding to the depth position is filled into the depth position, and a labeled depth map is obtained after filling. Among them, the depth map of the point cloud can be the value of the point cloud on a coordinate axis in the Cartesian coordinate system. The coordinate axis is related to the selection of the projection plane. Please refer to the following embodiments and will not be elaborated here.

[0138] In the case where multiple point clouds correspond to a single depth position, statistical processing can be performed on the depth values ​​of the multiple point clouds projected to the same depth position. Based on this, in some embodiments, the depth values ​​of the point cloud corresponding to the depth position are filled into the depth position to obtain a labeled depth map, including: when the depth position corresponds to multiple point clouds, statistical processing is performed on the depth values ​​of the multiple point clouds to obtain statistical values ​​of the multiple point clouds; and the statistical values ​​of the multiple point clouds are filled into the depth position to obtain a labeled depth map. The statistical processing includes, but is not limited to, near-point processing or averaging processing.

[0139] In addition, in order to further improve the quality of the label depth map, the label depth map can also be filtered and / or interpolated. For example, median filtering, mean filtering and other methods can be used to remove noise or outliers in the label depth map; linear interpolation, bilinear interpolation and the like can also be used to fill in the blank areas of the label depth map. In addition, in order to intuitively present the label depth map and verify the quality and accuracy of the label depth map, the label depth map can be visualized. In actual applications, visualization tools or drawing libraries in programming languages ​​can be used to display the label depth map. According to the visualization results and actual application requirements, the projection parameters, filtering methods, interpolation strategies, etc. can also be adjusted to obtain better depth map effects.

[0140] In step S013, a training sample is created based on the training point cloud data obtained in step S011 and the labeled depth map obtained in step S012 to construct a training dataset. For example, an association relationship is established between the training point cloud data and the labeled depth map corresponding to the same first point cloud data to create a training sample. In some embodiments, step S013 may include the following steps:

[0141] Step S0131: performing first data preprocessing on the training point cloud data;

[0142] Step S0132: performing second data preprocessing on the tag depth map;

[0143] Step S0133: Combine the training point cloud data after the first data preprocessing and the label depth map after the second data preprocessing to obtain a training sample.

[0144] The embodiments of the present application do not limit the specific methods of the first data preprocessing and the second data preprocessing, and they can be flexibly set according to the needs in actual applications. In some embodiments, the first data preprocessing includes at least one of the following: noise addition processing and first normalization processing. Among them, the noise addition processing is used to add noise to the training point cloud data, such as Gaussian white noise. The training point cloud data can be expanded through the noise addition processing to facilitate the generation of a large number of training samples; the first normalization processing is used to limit the values ​​of the training point cloud data to a certain range, so as to facilitate subsequent prediction and reduce the amount of calculation. The first normalization processing can include modulo squaring and logarithm-minimum-maximum-normalization of the training point cloud data. In some embodiments, the second data preprocessing includes a second normalization processing. Among them, the second normalization processing is used to limit the values ​​of the label depth map to a certain range, so as to facilitate subsequent prediction and reduce the amount of calculation. For other introductions and descriptions of the noise addition processing, the first normalization processing and the second normalization processing, please refer to the following embodiments, which will not be elaborated here.

[0145] During the model training process, the generative network is trained based on the training data set constructed in the above embodiment. To facilitate parameter adjustment of the generative network, a loss value can be calculated and then the parameters of the generative network can be adjusted to converge the loss value. Based on this, in some embodiments, the above step S020 may include the following steps:

[0146] Step S021: inputting the training point cloud data into the generation network so that the generation network outputs the training predicted depth map;

[0147] Step S022: constructing a target loss value based on the label depth map and the training prediction depth map;

[0148] Step S023: Train the generated network according to the target loss value.

[0149] Based on the comparison between the label depth map and the training prediction depth map, a target loss value can be constructed, and then the parameters of the generated network can be adjusted according to the target loss value to make the target loss value converge. The embodiment of the present application does not limit the calculation method of the target loss value, and it can be flexibly set according to the needs in actual application. For example, the target loss value can be calculated based on the error between the label depth map and the training prediction depth map, such as mean square error, mean absolute error, cross entropy loss, etc.

[0150] To improve the training accuracy of the network model, in some embodiments, step S022 may include the following steps:

[0151] Step S0221: inputting the label depth map and the training prediction depth map into the decision network respectively, so that the decision network outputs a first decision result and a second decision result respectively;

[0152] Step S0222: Determine the target loss value based on the first judgment result and the second judgment result.

[0153] During the model training process, a decision network is introduced to judge the true or false status of the label depth map and the training prediction depth map. The first decision result output by the decision network for the input label depth map includes the true or false status of the label depth map, and the second decision result output by the decision network for the input training prediction depth map includes the true or false status of the training prediction depth map. The first decision result and the second decision result can be presented in the form of a scalar, such as a value range of 0 to 1, the closer to 0, the more the decision network tends to think it is fake (Fake), and the closer to 1, the more the decision network tends to think it is real (Real, authentic).

[0154] In some embodiments, the decision network includes a first feature extraction module, a second feature extraction module, and a fully connected module. In the case of inputting a labeled depth map, the above step S0221 may include: inputting the labeled depth map into the first feature extraction module so that the first feature extraction module outputs first labeled feature data; inputting the training point cloud data into the second feature extraction module so that the second feature extraction module outputs second feature data; inputting the first labeled feature data and the second feature data into the fully connected module so that the fully connected module outputs a first decision result. In the case of inputting a training predicted depth map, the above step S0221 may include: inputting the training predicted depth map into the first feature extraction module so that the first feature extraction module outputs first predicted feature data; inputting the training point cloud data into the second feature extraction module so that the second feature extraction module outputs second feature data; inputting the first predicted feature data and the second feature data into the fully connected module so that the fully connected module outputs a second decision result.

[0155] The decision network can be implemented as a dual-stream heterogeneous structure. The first feature extraction module and the second feature extraction module in the decision network can be connected to the output and input ends of the generation network respectively to perform feature extraction on the two types of inputs; the fully connected module in the decision network connects (concat) and makes a decision on the output of the first feature extraction module and the output of the second feature extraction module.

[0156] The first feature extraction module is a two-dimensional feature extraction module, which is used to extract features from the depth map. It can perform feature extraction on the label depth map to obtain first label feature data, and perform feature extraction on the training prediction depth map to obtain first prediction feature data. The embodiment of the present application does not limit the specific structure of the first feature extraction module. In actual applications, it can be flexibly set according to needs. For example, the first feature extraction module may include a two-dimensional convolution layer, a two-dimensional normalization layer, a two-dimensional LeakyReLU (Leaky Rectified Linear Unit) layer, a two-dimensional Dropout layer, etc. Among them, the first feature extraction module may include multiple layers, such as 6 layers and above, and finally output the first label feature data or the first prediction feature data through the view function.

[0157] The second feature extraction module is a three-dimensional feature extraction module, which is used to extract features from point cloud data. It can extract features from training point cloud data to output second feature data. The embodiment of the present application does not limit the specific structure of the second feature extraction module. In actual applications, it can be flexibly configured according to needs. For example, the second feature extraction module may include a three-dimensional convolution layer, a three-dimensional normalization layer, a three-dimensional LeakyReLU layer, etc. Among them, the second feature extraction module may include multiple layers, such as 6 layers or more, and finally output the second feature data through the view function.

[0158] Based on the structure of the decision network described above, the first and second decision results output by the decision network can be conditional. For example, the first decision result can be the true or false status of the labeled depth map under the condition of the training point cloud data, and the second decision result can be the true or false status of the training predicted depth map under the condition of the training point cloud data. For further information on the first and second decision results, as well as the structure of the decision network, please refer to the following embodiments and will not be elaborated here.

[0159] Since the decision network is introduced during the model training process, the generator network and the decision network can be jointly optimized to achieve training of the generator network. Based on this, the target loss value calculated in step S0222 includes the first loss value of the decision network and the second loss value of the generator network. In some embodiments, step S0222 may include: determining the first loss value of the decision network based on the first and second decision results; and determining the second loss value of the generator network based on the second decision result.

[0160] Because the output of the decision network includes a first decision result and a second decision result, the first loss value of the decision network can be calculated by combining the first decision result and the second decision result. In some embodiments, determining the first loss value of the decision network based on the first decision result and the second decision result includes: determining first decision expected data based on the first decision result and the first data coefficient; determining second decision expected data based on the second decision result and the second data coefficient; and determining the first loss value of the decision network based on the first decision expected data and the second decision expected data.

[0161] Among them, the second data coefficient is smaller than the first data coefficient. Since the first data coefficient corresponds to the first judgment result for the label depth map, the first data coefficient can be set to be relatively large, for example, it can be set to between 0.9 and 1.1; since the second data coefficient corresponds to the second judgment result for the training prediction depth map, the second data coefficient can be set to be relatively small, such as 0 or 0.1. In addition, the above-mentioned determination of the first loss value of the decision network based on the first decision expected data and the second decision expected data includes: summing the square of the first decision expected data and the square of the second decision expected data to obtain the first loss value of the decision network. For other introductions and explanations on the calculation of the first decision expected data, the second decision expected data, and the first loss value, please refer to the following embodiments, which will not be elaborated here.

[0162] Because the output of the generative network includes the training predicted depth map, the second loss value of the generative network can be calculated by combining the second judgment result and the training predicted depth map. In some embodiments, determining the second loss value of the generative network based on the second judgment result includes: determining first generated expected data based on the second judgment result and the first data coefficient; determining second generated expected data based on the labeled depth map and the training predicted depth map; and determining the second loss value of the generative network based on the first generated expected data and the second generated expected data.

[0163] Determining the second loss value of the generative network based on the first and second expected generation data includes summing the product of the first expected generation data and the hyperparameter with the second expected generation data to obtain the second loss value of the generative network. For further information on the calculation of the first and second expected generation data, please refer to the following examples and will not be elaborated upon here.

[0164] Based on the first loss value of the decision network and the second loss value of the generative network, the decision network and the generative network can be jointly optimized. The joint optimization includes alternating training of the decision network and the generative network. Based on this, in some embodiments, step S023 may include: alternatingly adjusting parameters of the decision network and the generative network based on the first loss value and the second loss value, so that both the first loss value and the second loss value converge. Alternatingly adjusting parameters of the decision network and the generative network based on the first loss value and the second loss value, so that both the first loss value and the second loss value converge, may include: adjusting parameters of the generative network based on the second loss value so that the second loss value converges; determining the first loss value based on the adjusted generative network; adjusting parameters of the decision network based on the first loss value so that the first loss value converges; determining the second loss value based on the adjusted decision network; and, if the second loss value does not converge, re-performing the step of adjusting parameters of the generative network based on the second loss value so that the second loss value converges, until both the first loss value and the second loss value converge.

[0165] It should be understood that in the embodiments of the present application, for the sake of convenience in distinction, the depth map output by the generated network during the application process is referred to as the original predicted depth map, the depth map output by the generated network during the training process is referred to as the trained predicted depth map, and the depth map used as a label is referred to as the labeled depth map. These names do not constitute a limitation on the embodiments of the present application.

[0166] Below, an example is used to introduce the point cloud enhancement method provided in the embodiment of the present application.

[0167] See also Figure 2 , Figure 2This is a flow chart of another point cloud enhancement method provided by an embodiment of the present application. Figure 2 As shown, the point cloud enhancement method may include the following steps S2100 to S2500.

[0168] Step S2100: Generate training point cloud data.

[0169] The training point cloud data can be a radar data cube (RDC), which can take the dense point cloud data of at least one object as input and obtain sparse point cloud data through rotation, translation, occlusion processing, and mirror reflection processing. Then, according to the setting of radar simulation parameters, the sparse point cloud data is converted into a three-dimensional RDC, whose dimensions can be distance dimension, horizontal angle dimension, and vertical angle dimension respectively. Among them, at least one object can be a three-dimensional object. For example, in a vehicle-mounted scene, at least one object can include a vehicle, a pedestrian, a traffic sign, an obstacle, etc. Among them, step S2100 can include the following steps S2110 to S2180.

[0170] Step S2110: Acquire dense point cloud data of at least one object.

[0171] The dense point cloud data uses the Cartesian coordinate system as a reference.

[0172] Step S2120: Setting radar simulation parameters.

[0173] Radar simulation parameters include but are not limited to at least one of the following: operating frequency band f c , frequency modulation slope S, ADC sampling rate f s , the number of ADC sampling points N ADC , the number of horizontal antennas of the array antenna N a 、N number of vertical antennas in the array antenna e , the array arrangement of the array antenna. In addition, it can be assumed that the radar works in FMCW (Frequency Modulated Continuous Wave). According to the radar simulation parameters set above, the radar's range resolution ΔR can be expressed as follows: max It can be shown as the following formula 2.

[0174] Formula 1:

[0175] Formula 2:

[0176] Where c is the speed of light.

[0177] In addition, according to the parameters of the array antenna, the horizontal angle resolution Δθ and vertical angle resolution of the radar can be determined. like Figure 3 As shown, according to the radar antenna's directional pattern, the radar's horizontal field of view, vertical field of view, and the transmitting antenna gain in all directions can be determined. and receiving antenna gain

[0178] Step S2130: performing position transformation on the dense point cloud data of at least one object according to the point cloud generation conditions to obtain first point cloud data P1.

[0179] Point cloud generation conditions include: at least one object does not overlap in space, and at least one object is within the radar sensing range. Figure 4 As shown in (a), the position transformation data of at least one object is obtained based on the random number generation method, and the position transformation data includes the rotation angle data ω and the position offset data O of the object relative to the radar observation center O. x / O y The dense point cloud data of at least one object is positionally transformed based on the positional transformation data of the at least one object, and it is determined whether the dense point cloud data after the positional transformation satisfies the point cloud generation conditions. If the point cloud generation conditions are not satisfied, the positional transformation data of the at least one object is re-obtained based on a random number generation method, and this process is iterated until the point cloud generation conditions are satisfied, thereby obtaining first point cloud data P1.

[0180] Step S2140: performing occlusion processing on the first point cloud data P1 to obtain second point cloud data P2.

[0181] like Figure 4 As shown in (b), based on the first point cloud data P1, the occlusion effect in the radar reflection characteristics is simulated to obtain the second point cloud data P2 after occlusion processing. Among them, step S2140 can include the following steps S2141 to S2144.

[0182] Step S2141: convert each point cloud in the first point cloud data P1 from a Cartesian coordinate system to a spherical coordinate system.

[0183] Assuming that the radar observation center O is at the origin, the coordinate data (x i ,y i , z i ) is converted into coordinate data in the spherical coordinate system (R i ,θ i , );

[0184] Step S2142: Determine radar resolution data according to radar simulation parameters, and construct a voxel grid according to the radar resolution data; determine the position code of each point cloud in the voxel grid.

[0185] According to the radar simulation parameters, the range resolution ΔR, horizontal angle resolution Δθ and vertical angle resolution can be obtained. To form a voxel grid (Voxel Grid). Then we can calculate the position code of each point cloud falling in the voxel grid

[0186] Step S2143: based on the minimum point cloud distance of each voxel in the voxel grid, eliminate the point clouds whose point cloud distance in the voxel is greater than the minimum point cloud distance.

[0187] For each voxel in the voxel grid, calculate the voxel (θ′, )’s minimum point cloud distance R min = min(R′), remove the voxel (θ′, )The point cloud distance is greater than the minimum point cloud distance R′>R min point cloud.

[0188] Step S2144: convert the remaining point cloud in the voxel grid from the spherical coordinate system to the Cartesian coordinate system to obtain the second point cloud data P2.

[0189] The remaining point cloud is converted from the spherical coordinate system to the Cartesian coordinate system to obtain the second point cloud data P2 after occlusion processing.

[0190] Step S2150: performing mirror reflection processing on the second point cloud data P2 to obtain third point cloud data P3.

[0191] like Figure 4 As shown in (c), based on the second point cloud data P2, the mirror reflection effect of the radar is simulated to obtain the third point cloud data P3 processed by the mirror reflection. Among them, step S2150 can include the following steps S2151 to S2154.

[0192] Step S2151: determining the base edge of each point cloud in the second point cloud data according to the base edge data of the second point cloud data.

[0193] Calculate the direction vectors of the four bottom edges in the 3D bounding box of the second point cloud data using them as a reference. and Calculate the distance from each point cloud in the second point cloud data to the four bottom edges, and take the bottom edge corresponding to the minimum distance as the bottom edge of the point cloud because The vertical component z of is 0 and can therefore be ignored.

[0194] Step S2152: Determine the angle between the projection component of the point cloud and its corresponding bottom edge, as well as the pitch angle of the point cloud.

[0195] The projection component of the direction vector from each point cloud in the second point cloud data P2 to the radar observation center O on the xoy plane is calculated as calculate The bottom edge of the point cloud The included angle α between them is shown in the following formula 3.

[0196] Formula 3:

[0197] in,

[0198] In addition, the pitch angle β of the point cloud can also be calculated as shown in the following formula 4.

[0199] Formula 4:

[0200] Step S2153: Calculate the mirror reflection angle of the point cloud according to the included angle and the pitch angle of the point cloud, so as to classify and sample the second point cloud data P2 according to the mirror reflection angle of the point cloud.

[0201] The specular reflection angle γ can be calculated based on the following formula 5.

[0202] Formula 5: γ = max(α, β)

[0203] When γ=0°, the point cloud can be added to the first scattered spot set. The first scattered spot set is randomly sampled, and the number of samples can be 10, so as to obtain the first set P3-1.

[0204] When 0°<γ and γ≤15°, the point cloud can be added to the second scattering spot set. The second scattering spot set is randomly sampled, and the number of samples can be 20, so that the second set P3-2 can be obtained;

[0205] When 15°<γ and γ≤25°, the point cloud can be added to the third scattering spot set. The third scattering spot set is randomly sampled, and the number of samples can be 5, so as to obtain a third set P3-3.

[0206] Step S2154: Merge the first set P3-1, the second set P3-2, and the third set P3-3 to obtain third point cloud data P3 processed by mirror reflection.

[0207] Step S2160: performing expansion processing on the third point cloud data P3 to obtain fourth point cloud data P4.

[0208] like Figure 4As shown in (d), the radar reflection points are filled, and the point clouds with a spatial distance of no more than 0.3 m in the second point cloud data P2 and the third point cloud data P3 are selected according to the ellipsoid law. The expanded point cloud data is obtained by merging, which is the fourth point cloud data P4.

[0209] Step S2170: performing analog-to-digital conversion on the fourth point cloud data to obtain complex ADC data.

[0210] According to the TDMA-FMCW radar imaging system, the snapshot signal S of each point cloud i in the fourth point cloud data P4 is calculated. i (r, θ, ), as shown in Formula 6 below.

[0211] Formula 6:

[0212] in, σ b is the radar cross-section of the object; R is the distance from the object to the observation center of the radar antenna array; τ is the time delay of the reflected wave compared to the transmitted wave, as shown in the following formula 7.

[0213] Formula 7:

[0214] Where dist represents the Euclidean distance; Tx represents the spatial position of the transmitting antenna element; Rx represents the spatial position of the receiving antenna element; P4 i Indicates the spatial position of the i-th point in the fourth point cloud data P4.

[0215] The snapshot signal S of all point clouds i (r, θ, ) are superimposed, as shown in Formula 8 below.

[0216] Formula 8:

[0217] Where M is the number of elements in the fourth point cloud data P4. According to formula 8, the complex ADC data received by the radar in the current scene can be obtained:

[0218] Step S2180: Perform fast Fourier transform on the complex ADC data to obtain training point cloud data.

[0219] According to the conventional signal processing of the complex ADC data, the range-FFT, Doppler-FFT, horizontal angle-FFT and vertical angle-FFT can be performed in sequence to obtain the radar data cube RDC, that is, in, Represents a complex tensor, N′ r , N′ v , N′a , N′ e are the number of samples of range-FFT, Doppler-FFT, horizontal angle-FFT and vertical angle-FFT respectively. Considering that it can mainly target stationary targets or moving targets that have been compensated by Doppler, RDC can be degenerated into

[0220] Step S2200: Generate a label depth map.

[0221] The label depth map is an image label based on the principle of visual imaging. Using the first point cloud data as input, a label depth map corresponding to the training point cloud data (RDC) is generated for network model learning. Step S2200 may include the following steps S2210 to S2280.

[0222] Step S2210: Acquire first point cloud data.

[0223] Among them, the first point cloud data (x i ,y i , z i ) can be the three-dimensional point cloud data output by the above step S2130.

[0224] Step S2220: Select a projection plane.

[0225] The imaging plane of the depth camera can be selected as the projection plane, and a two-dimensional coordinate system can be further constructed based on the selected projection plane. The two-dimensional coordinate system can be based on the upper left corner of the image as the origin, the horizontal right direction as the positive x-axis, and the vertical downward direction as the positive y-axis.

[0226] Step S2230: Setting the optical parameters of the depth camera.

[0227] The optical parameters of the depth camera include the internal parameters of the depth camera, which include but are not limited to: image size W×H pixels, focal length fx / fy, principal point coordinates cx / cy, measurement range I max wait;

[0228] Step S2240: Calculate the three-dimensional projection coordinates of each point cloud in the first point cloud data.

[0229] For each point cloud (x i ,y i , z i ), according to the physical parameters such as the installation position and posture of the depth camera and the selection of the projection plane, it is converted into the projection coordinate system to obtain the three-dimensional projection coordinates This usually involves geometric transformations such as rotations and translations, which can be achieved by matrix multiplication to achieve coordinate system transformations.

[0230] Step S2250: Calculate the two-dimensional projection coordinates of each point cloud in the first point cloud data.

[0231] According to the optical parameters of the depth camera, such as the intrinsic parameters, the three-dimensional projection coordinates of the point cloud in the projection coordinate system are projected onto the imaging plane of the depth camera to obtain the two-dimensional projection coordinates I(u, v), which can be shown in the following formula 9.

[0232]

[0233] Step S2260: Fill the depth value of each point cloud into the initial depth map according to the two-dimensional projection coordinates of each point cloud to obtain a label depth map.

[0234] like Figure 5 As shown, according to the internal parameters such as image size (width W and height H), the initialization depth map can be obtained, such as creating a two-dimensional array or matrix with the initial value set to I max Then, the depth value of each point cloud (usually the z coordinate value of the point cloud in the Cartesian coordinate system) is filled into the corresponding position in the initial depth map according to the 2D projection coordinates I(u, v) of each point cloud. If multiple point clouds are projected to the same position, they can be processed according to the depth value of the nearest point or the average depth value.

[0235] Step S2270: Filter and / or interpolate the tag depth map.

[0236] Median filtering, mean filtering and other methods can be used to remove noise or outliers in the label depth map; linear interpolation, bilinear interpolation and other methods can also be used to fill the blank areas of the label depth map.

[0237] Step S2280: Verify and adjust the label depth map.

[0238] The generated labeled depth map can be visualized to visually check its quality and accuracy. This can be displayed using visualization tools or drawing libraries in programming languages. Based on the visualization results and actual application requirements, adjustments can be made to projection parameters, filtering methods, interpolation strategies, and other factors to achieve a better depth map.

[0239] Step S2300: performing data preprocessing on the training point cloud data and the label depth map.

[0240] The training point cloud data (RDC) generated by simulation can be sorted, denoised, normalized and other data preprocessing. In addition, the label depth map I GT Perform data preprocessing such as normalization to obtain a large amount of training data sets. Step S2300 may include the following steps S2310 to S2330.

[0241] Step S2310: sort the training samples.

[0242] According to the training point cloud data and the label depth map, training samples can be constructed to construct a training data set. Among them, the training samples can be sorted to construct independent learning data sets D train And the test dataset D test .

[0243] Step S2320: performing first data preprocessing on the training point cloud data in the training sample.

[0244] The first data preprocessing may include noise addition processing and first normalization processing.

[0245] Noise processing can be done by adding complex Gaussian white noise to the training point cloud data. The amplitude of the Gaussian white noise can be gradually added at intervals of -5, 0, 5, and 10 dB according to the signal-to-noise ratio, thereby obtaining amplified training point cloud data. train And the test dataset D test , we can get the expanded learning data set and test dataset

[0246] Then the first normalization process is performed on the training point cloud data after noise processing, such as and test dataset The training point cloud data in is first normalized. The training point cloud data can be first modulo-squared, as shown in Formula 10 below; then the first normalization process is performed, such as log-min-max normalization, as shown in Formula 11 below.

[0247] Formula 10: R = |RDC| 2

[0248] Here, R refers to the point cloud data tensor heat map.

[0249] Formula 11:

[0250] Among them, R norm is the training point cloud data after the first normalization process; max[A] and min[A] are the maximum and minimum values ​​of vector A respectively; vec(B) means vectorizing tensor B; R log It represents the result of taking the logarithm of R, that is, R log =log(R).

[0251] Step S2330: performing a second normalization process on the label depth map.

[0252] It is also necessary to label the depth map I GT (u, v) is normalized as shown in the following formula 12.

[0253] Formula 12:

[0254] in, Represents a real number.

[0255] Step S2400: training the generation network and the decision network according to the training data set.

[0256] A point cloud imaging enhancement network model based on a super-resolution generative adversarial network (SRGAN) can be built, comprising a generator network G and a decision network D. Training point cloud data and a labeled depth map are input into the point cloud imaging enhancement network model to train the generator network G and the decision network D, and a loss value is defined to adjust the parameters of the network model until the network converges to an optimal state. Step S2400 may include steps S2410 to S2440.

[0257] Step S2410: Process the training point cloud data by generating the encoding module in the network G to obtain point cloud feature data.

[0258] like Figure 6 As shown, the encoding module (Encoder) is based on R norm The distance information is used as the learning feature and R is extracted by the downsampling unit norm The encoding module includes at least three downsampling units, each of which can be composed of a two-dimensional convolution block (Conv2D), a two-dimensional activation layer (ReLU-2D), a two-dimensional batch normalization (BatchNormalization-2D), a two-dimensional maximum pooling (Maxpooling-2D) and other structures.

[0259] Step S2420: Process the point cloud feature data by generating a decoding module in the network G to obtain a training predicted depth map.

[0260] like Figure 6As shown in the figure, the decoding module (Decoder) in the generated network G is symmetrically connected to the encoding module and contains multiple upsampling units. The upsampling units can be implemented through bilinear interpolation. Each upsampling unit can receive the spatial feature data output by the corresponding downsampling unit in the encoding module through a 2D-2D jump connection. In addition, there is no need for jump connections in the output unit of the decoding module. Two or more 1D convolutional layers are used to achieve feature dimension compression and angle super-resolution. Finally, the features are mapped to between 0 and 1 through the Sigmoid function to obtain the network output G (R) consistent with the dimension of the label depth map. norm ).

[0261] Step S2430: Through the decision network D, feature extraction is performed on the input and output of the generation network G to obtain a first decision result and a second decision result.

[0262] like Figure 7 As shown, the decision network D is respectively connected to the input R of the encoding module in the generation network G norm and the output of the decoding module G(R norm ) are connected, and features of the two types of data are extracted through a dual-stream heterogeneous feature network. Step S2430 may include the following steps S2431 to S2433.

[0263] Step S2431: The label depth map and the training prediction depth map are processed respectively by the first feature extraction module in the decision network D to obtain first label feature data and first prediction feature data.

[0264] like Figure 7 As shown, the first feature extraction module in the decision network D is a two-dimensional feature extraction module, mainly composed of a two-dimensional convolutional layer, a two-dimensional normalization layer, a two-dimensional LeakyReLU layer, a two-dimensional Dropout layer, etc. This module contains 6 or more layers, and finally obtains an F1×1 feature vector through the view function. When the first feature extraction module inputs the label depth map, the output F1×1 feature vector is the first label feature data; when the first feature extraction module inputs the training prediction depth map, the output F1×1 feature vector is the first prediction feature data.

[0265] Step S2432: Process the training point cloud data through the second feature extraction module in the decision network D to obtain second feature data.

[0266] like Figure 7 As shown, the second feature extraction module in the decision network D is a three-dimensional feature extraction module, which can be used to extract R norm Reshape into The second feature extraction module is mainly composed of a three-dimensional convolution layer, a three-dimensional normalization layer, and a three-dimensional LeakyReLU layer. The module has a total of 6 layers or more. Finally, the F2×1 feature vector, i.e., the second feature data, is obtained through the view function.

[0267] Step S2433: The fully connected module in the decision network D outputs a first decision result and a second decision result according to the first label feature data and the second feature data, and the first predicted feature data and the second feature data, respectively.

[0268] like Figure 7 As shown, the fully connected module in the decision network D concats the first label feature data and the second feature data, or concats the first predicted feature data and the second feature data to obtain a feature vector of (F1+F2)×1, which is then input into the fully connected layer to ultimately obtain the first and second decision results. When the first feature extraction module in the decision network D inputs the label depth map, the fully connected module in the decision network D processes the first label feature data and the second feature data to obtain the first decision result; when the first feature extraction module in the decision network D inputs the training predicted depth map, the fully connected module in the decision network D processes the first predicted feature data and the second feature data to obtain the second decision result.

[0269] Step S2440: Determine a target loss value based on the first judgment result and the second judgment result, and train the generation network and the judgment network based on the target loss value.

[0270] The target loss value includes the first loss value of the decision network D and the second loss value of the generation network G.

[0271] Determine the first loss value of network D It can be shown as the following formula 13.

[0272] Formula 13:

[0273] in, is the first data coefficient, representing the label smoothing random number, and its value can be between 0.9 and 1.1; GT is the label depth map, Indicates that the decision network D is Under the condition of GT The first judgment result; In addition, Indicates that the decision network D is Under the condition of input G(R norm )’s judgment.

[0274] Generate the second loss value of network G It can be shown as the following formula 14.

[0275] Formula 14:

[0276] Among them, λ1 is a hyperparameter with a value of 10 -3 .

[0277] In formula 14 It can be shown as the following formula 15.

[0278] Formula 15:

[0279] GAN loss value in formula 14 It can be shown as the following formula 16.

[0280] Formula 16:

[0281] Afterwards, the parameters of the decision network D and the generation network G can be updated using an alternating training strategy until both the first loss value and the second loss value converge.

[0282] Step S2500: performing point cloud enhancement processing according to the generated network.

[0283] After the training of the generative network G is completed, the original point cloud data collected by the radar can be input into the trained generative network G, so that the generative network G outputs the original predicted depth map corresponding to the original point cloud data, and the point cloud is reconstructed according to the original predicted depth map to obtain the target point cloud data corresponding to the original point cloud data, thereby realizing point cloud enhancement.

[0284] Among them, point cloud reconstruction is point cloud post-processing, which can amplify the original predicted depth map output by the generation network G to obtain the target predicted depth map I pred , as shown in Formula 17 below.

[0285] Formula 17: I pred =G(R norm )×I max

[0286] Then, the target predicted depth map can be mapped into three-dimensional point cloud data according to the optical parameters of the depth camera, such as internal parameters, and the tailing point cloud data can be filtered through the DBSCAN algorithm or its variant algorithm to obtain the target point cloud data after point cloud enhancement.

[0287] The embodiment of the present application provides a cross-modal radar point cloud imaging enhancement method, including radar data cube (RDC) generation, vision-based label depth map generation, data preprocessing, point cloud imaging enhancement network based on super-resolution generative adversarial network, point cloud post-processing, etc. Among them, the generation network G takes RDC as input and its distance dimension information as learning features, obtains the predicted depth map through asymmetric U-net, and uses the corresponding visual depth label to supervise the generation of the predicted depth map; it also provides a conditional decision network D, which takes RDC and predicted depth map as input and obtains the decision result through a dual-stream heterogeneous feature extraction network; finally, the optical parameters of the depth camera are used to post-process and map the predicted depth map into a three-dimensional point cloud. The embodiment of the present application can solve the technical problems of low point cloud imaging quality, sparse point cloud and lack of large-scale high-quality training samples in the related technology, and can improve the imaging performance of radar for close-range extended targets.

[0288] Below, a simulation experiment is used to introduce and illustrate the effectiveness of the point cloud enhancement method provided in the embodiment of the present application.

[0289] The radar simulation parameters are set as follows: working frequency band f c =77GHz, frequency modulation slope S=2×10 13 Hz / s, ADC sampling rate f s =2MHz, ADC sampling points N ADC =120; the radar working system can be set to TDMA-FMCW, the radar antenna array adopts a "mouth"-shaped antenna array with 40 transmit and 40 receive (equivalent to a two-dimensional dense planar array), and the radar signal processing parameter N′ r =96, N′ a =64, N′ e =32. In addition, the parameters of the depth camera are set as follows: image size W×H=256×128; focal length f x =255.270, f y =220.957; principal point coordinate c x =125.031, c y =70.308;I max =20m.

[0290] like Figure 8 As shown, the depth map prediction methods provided in the embodiment of the present application and the related art are compared. It can be seen that: first, compared with the traditional nearest neighbor interpolation method, the embodiment of the present application, Adjusted RadarHD and Hawkeye can effectively fill in the missing vehicle echo scattering depth pixels, effectively overcoming the radar mirror reflection effect; second, compared Figure 8As can be seen from the third and fifth columns in the figure, the present embodiment has improved image detail texture and depth map estimation compared to Hawkeye. This is due to the improvement of the generative network G, which replaces the 3D-2D jump connection in Hawkeye with a 2D-2D jump connection. Figure 8 As can be seen from the fourth and fifth columns, the embodiment of the present application has improved image detail texture compared to AdjustedRadarHD. This is due to the introduction of the decision network D, which uses the confrontation between the generation network G and the decision network D to enhance the rationality and detail of the depth map.

[0291] like Figure 9 As shown, Figure 9 The real point cloud data, Hawkeye, the adjusted RadarHD and the point cloud data generated by the embodiment of the present application were compared. Figure 9 As shown, compared with Haweye and RadarHD, the point cloud data output by the embodiment of the present application has clearer vehicle outlines and better point cloud enhancement effects.

[0292] According to the second aspect of the present application, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the above-mentioned point cloud enhancement method is implemented and has all the beneficial effects of the above-mentioned point cloud enhancement method. This application will not go into details here.

[0293] According to the third aspect of the present application, an embodiment of the present application also provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the above-mentioned point cloud enhancement method is implemented and has all the beneficial effects of the above-mentioned point cloud enhancement method. This application will not go into details here.

[0294] According to the fourth aspect of the present application, an embodiment of the present application also provides a controller on which a computer program or instruction is stored. When the computer program or instruction is executed by the processor, the above-mentioned point cloud enhancement method is implemented and has all the beneficial effects of the above-mentioned point cloud enhancement method. This application will not go into details here.

[0295] According to a fifth aspect of the present application, embodiments of the present application further provide a point cloud enhancement system, comprising a controller and a radar connected to the controller. The controller may be the controller described in the fourth aspect above. The radar may be used to collect raw point cloud data. In some embodiments, the radar comprises a millimeter-wave radar. Of course, the radar may also be implemented as a laser radar, etc., which is not limited in the embodiments of the present application. The point cloud enhancement system can be used to implement the above-mentioned point cloud enhancement method and has all the beneficial effects of the above-mentioned point cloud enhancement method, which will not be further elaborated in this application.

[0296] The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof, and this application does not specifically limit this. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0297] In some embodiments of the present application, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0298] The computer-readable storage medium may be included in the controller or may exist independently without being incorporated into the controller. The computer-readable storage medium carries one or more programs, which, when executed by the controller, cause the controller to:

[0299] Performing prediction processing on the original point cloud data to obtain an original predicted depth map corresponding to the original point cloud data;

[0300] Point cloud reconstruction is performed based on the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data.

[0301] Computer program code for performing the operations of some embodiments of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, using an Internet service provider to connect via the Internet).

[0302] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function.

[0303] It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures.

[0304] For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow charts, and combinations of blocks in the block diagrams and / or flow charts, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions.

[0305] The units described in some embodiments of the present application may be implemented in software or hardware, and may also be provided in a processor.

[0306] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.

[0307] According to the sixth aspect of this application, Figure 10 As shown, the embodiment of the present application further provides a vehicle 10, which includes the above-mentioned controller, or includes the above-mentioned point cloud enhancement system. The vehicle has all the beneficial effects of the above-mentioned controller and point cloud enhancement system, etc., and this application will not repeat them here.

[0308] The vehicle may be a fuel vehicle, a plug-in hybrid vehicle or a new energy vehicle, etc., and this application does not make any specific restrictions on this.

[0309] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0310] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0311] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0312] The above are merely preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of each embodiment in the embodiments of the present application have different focuses, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A point cloud enhancement method, characterized in that: The method comprises: Performing prediction processing on the original point cloud data to obtain an original predicted depth map corresponding to the original point cloud data; Point cloud reconstruction is performed based on the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data.

2. The method according to claim 1, characterized in that The step of reconstructing the point cloud according to the original predicted depth map to obtain target point cloud data corresponding to the original point cloud data includes: Performing point cloud mapping according to the original predicted depth map to obtain intermediate point cloud data; The intermediate point cloud data is filtered to obtain target point cloud data corresponding to the original point cloud data.

3. The method according to claim 2, characterized in that The performing point cloud mapping according to the original predicted depth map to obtain intermediate point cloud data includes: Adjusting the size of the original predicted depth map to obtain a target predicted depth map; Mapping processing is performed on the target predicted depth map to obtain intermediate point cloud data.

4. The method according to claim 2, characterized in that The filtering process on the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data includes: identifying tailing point cloud data in the intermediate point cloud data; The trailing point cloud data is removed from the intermediate point cloud data to obtain target point cloud data corresponding to the original point cloud data.

5. The method according to claim 1, wherein The predicting process of the original point cloud data to obtain an original predicted depth map corresponding to the original point cloud data includes: Inputting the original point cloud data into the generation network so that the generation network outputs the original predicted depth map corresponding to the original point cloud data; The generation network is used to perform prediction processing on the point cloud data to output a predicted depth map.

6. The method according to claim 5, characterized in that The generation network includes an encoding module and a decoding module; Inputting the original point cloud data into the generation network so that the generation network outputs the original predicted depth map corresponding to the original point cloud data includes: Inputting the original point cloud data into the encoding module so that the encoding module outputs point cloud feature data of the original point cloud data; The point cloud feature data is input into the decoding module so that the decoding module outputs an original predicted depth map corresponding to the original point cloud data.

7. The method according to claim 6, characterized in that The encoding module includes P downsampling units, where P is an integer greater than 1; Inputting the original point cloud data into the encoding module so that the encoding module outputs point cloud feature data of the original point cloud data includes: Inputting the original point cloud data into the first downsampling unit, so that the first downsampling unit outputs spatial feature data of the original point cloud data at a first scale; Inputting the spatial feature data at the Kth scale into the K+1th downsampling unit, so that the K+1th downsampling unit outputs the spatial feature data of the original point cloud data at the K+1th scale; wherein K is a positive integer less than P; The point cloud feature data of the original point cloud data includes the spatial feature data of the original point cloud data at P scales.

8. The method according to claim 7, characterized in that The decoding module includes P upsampling units and an output unit; Inputting the point cloud feature data into the decoding module so that the decoding module outputs an original predicted depth map corresponding to the original point cloud data includes: Inputting the spatial feature data at the P-th scale into the first upsampling unit, so that the first upsampling unit outputs original prediction data of the original point cloud data at the P-th scale; Inputting the original prediction data at the K+1th scale into the P-K+1th upsampling unit, so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale; The original predicted data at the first scale is input into the output unit, so that the output unit outputs an original predicted depth map corresponding to the original point cloud data.

9. The method according to claim 8, characterized in that The P-K+1th upsampling unit is jump-connected to the K+1th downsampling unit; Inputting the original prediction data at the K+1th scale into the P-K+1th upsampling unit, so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale, includes: The original prediction data at the K+1th scale and the spatial feature data at the K+1th scale are input into the P-K+1th upsampling unit, so that the P-K+1th upsampling unit outputs the original prediction data of the original point cloud data at the Kth scale.

10. The method according to claim 5, characterized in that The method further comprises: Constructing a training data set; wherein each training sample in the training data set includes a training point cloud data and a labeled depth map; The generation network is trained according to the training data set.

11. The method according to claim 10, characterized in that The constructing of the training data set includes: Sparsely processing the first point cloud data to obtain the training point cloud data; Performing projection processing on the first point cloud data to obtain the label depth map; The training samples are established according to the training point cloud data and the label depth map.

12. The method according to claim 11, characterized in that The thinning the first point cloud data to obtain the training point cloud data includes: Performing occlusion processing on the first point cloud data to obtain second point cloud data; performing mirror reflection processing on the second point cloud data to obtain third point cloud data; The training point cloud data is determined according to the third point cloud data.

13. The method according to claim 12, characterized in that The performing occlusion processing on the first point cloud data to obtain the second point cloud data includes: Determining a position code of each point cloud in the first point cloud data in a voxel grid; performing point cloud pruning on each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data; Second point cloud data is determined based on the remaining point clouds in the voxel grid.

14. The method according to claim 13, characterized in that The determining of a position code of each point cloud in the first point cloud data in a voxel grid includes: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a spherical coordinate system; According to the coordinate data of each point cloud in the first point cloud data in the spherical coordinate system, a position code of each point cloud in the first point cloud data in a voxel grid is determined.

15. The method according to claim 13, characterized in that The step of performing point cloud pruning on each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data comprises: Obtaining a minimum point cloud distance of each voxel in the voxel grid according to the position code of each point cloud in the first point cloud data; For each voxel in the voxel grid, the point cloud having a point cloud distance in the voxel greater than the minimum point cloud distance is eliminated.

16. The method according to claim 13, characterized in that Determining second point cloud data based on the remaining point cloud in the voxel grid includes: The remaining point cloud in the voxel grid is converted from a spherical coordinate system to a Cartesian coordinate system to obtain the second point cloud data.

17. The method according to claim 12, wherein: The performing mirror reflection processing on the second point cloud data to obtain third point cloud data includes: Determine bottom edge data based on the second point cloud data; wherein the bottom edge data includes multiple bottom edges in a bounding box corresponding to the second point cloud data; determining a mirror reflection angle of each point cloud in the second point cloud data according to the bottom edge data; The third point cloud data is determined according to the specular reflection angle of each point cloud in the second point cloud data.

18. The method according to claim 17, characterized in that The determining, based on the bottom edge data, a mirror reflection angle of each point cloud in the second point cloud data comprises: For each point cloud in the second point cloud data, determining the base edge to which the point cloud belongs based on the distance between the point cloud and each base edge in the base edge data; For each point cloud in the second point cloud data, determining an angle between a projection component of the point cloud and the corresponding bottom edge, and a pitch angle of the point cloud; For each point cloud in the second point cloud data, a mirror reflection angle of the point cloud is determined according to the included angle and the pitch angle of the point cloud.

19. The method according to claim 17, wherein The determining the third point cloud data according to the mirror reflection angle of each point cloud in the second point cloud data comprises: classifying the second point cloud data according to the specular reflection angle of each point cloud in the second point cloud data to obtain at least one scattered spot set; The third point cloud data is obtained by sampling each of the scattering spot sets.

20. The method according to claim 19, wherein The obtaining of the third point cloud data according to the sampling process of each of the scattered spot sets includes: Extracting at least one point cloud from each of the scattered spot sets; The point clouds extracted from at least one of the scattered spot sets are merged to obtain the third point cloud data.

21. The method according to claim 12, wherein The determining the training point cloud data according to the third point cloud data includes: performing expansion processing on the third point cloud data to obtain fourth point cloud data; The training point cloud data is determined according to the fourth point cloud data.

22. The method according to claim 21, characterized in that The expanding the third point cloud data to obtain fourth point cloud data includes: The point cloud that meets the first condition in the second point cloud data is filled into the third point cloud data to obtain fourth point cloud data.

23. The method according to claim 21, characterized in that The determining the training point cloud data according to the fourth point cloud data includes: Performing analog-to-digital conversion on the fourth point cloud data to obtain complex ADC data; Perform spectrum conversion on the complex ADC data to obtain the training point cloud data.

24. The method according to claim 12, wherein The method further comprises: The first point cloud data is generated according to dense point cloud data of at least one object.

25. The method according to claim 24, characterized in that Generating the first point cloud data according to the dense point cloud data of at least one object includes: According to the point cloud generation condition, position transformation is performed on the dense point cloud data of at least one object to obtain the first point cloud data.

26. The method according to claim 25, characterized in that The point cloud generation condition includes at least one of the following: at least one of the objects does not overlap in space, and at least one of the objects is within the radar perception range.

27. The method according to claim 11, wherein The projecting the first point cloud data to obtain the label depth map includes: Determining the projection coordinates of each point cloud in the first point cloud data; According to the projection coordinates of each point cloud in the first point cloud data, the depth value of each point cloud in the first point cloud data is filled into the initial depth map to obtain the label depth map.

28. The method according to claim 27, characterized in that The determining of the projection coordinates of each point cloud in the first point cloud data includes: converting each point cloud in the first point cloud data from a Cartesian coordinate system to a projected coordinate system; The projection coordinates of each point cloud in the first point cloud data are determined according to the coordinate data of each point cloud in the projection coordinate system.

29. The method according to claim 27, characterized in that The step of filling the initial depth map with the depth value of each point cloud in the first point cloud data according to the projection coordinates of each point cloud in the first point cloud data to obtain the label depth map includes: Determining, according to the projection coordinates of each point cloud in the first point cloud data, a depth position corresponding to each point cloud in the first point cloud data in the initial depth map; For each depth position in the initial depth map, the depth value of the point cloud corresponding to the depth position is filled into the depth position to obtain the label depth map.

30. The method according to claim 11, wherein The step of establishing the training sample according to the training point cloud data and the label depth map includes: Performing a first data preprocessing on the training point cloud data; Performing second data preprocessing on the label depth map; The training point cloud data after the first data preprocessing and the label depth map after the second data preprocessing are combined to obtain the training sample.

31. The method according to claim 30, wherein The first data preprocessing includes at least one of the following: noise addition processing and first normalization processing.

32. The method according to claim 30, wherein The second data preprocessing includes a second normalization process.

33. The method according to claim 10, wherein The step of training the generating network according to the training data set includes: Inputting the training point cloud data into the generative network so that the generative network outputs a training predicted depth map; Constructing a target loss value according to the label depth map and the training prediction depth map; The generating network is trained according to the target loss value.

34. The method according to claim 33, wherein The constructing a target loss value according to the label depth map and the training prediction depth map includes: Inputting the label depth map and the training prediction depth map into a decision network respectively, so that the decision network outputs a first decision result and a second decision result respectively; wherein the first decision result includes the true or false state of the label depth map, and the second decision result includes the true or false state of the training prediction depth map; A target loss value is determined according to the first judgment result and the second judgment result.

35. The method according to claim 34, wherein The decision network includes a first feature extraction module, a second feature extraction module and a fully connected module; Inputting the label depth map and the training prediction depth map into a decision network respectively so that the decision network outputs a first decision result and a second decision result respectively includes: Inputting the label depth map into the first feature extraction module so that the first feature extraction module outputs first label feature data; inputting the training point cloud data into the second feature extraction module so that the second feature extraction module outputs second feature data; inputting the first label feature data and the second feature data into the fully connected module so that the fully connected module outputs the first judgment result; The training predicted depth map is input into the first feature extraction module so that the first feature extraction module outputs first predicted feature data; the training point cloud data is input into the second feature extraction module so that the second feature extraction module outputs second feature data; the first predicted feature data and the second feature data are input into the fully connected module so that the fully connected module outputs the second judgment result.

36. The method according to claim 34, wherein The determining of the target loss value according to the first judgment result and the second judgment result includes: Determining a first loss value of the judgment network according to the first judgment result and the second judgment result; Determining a second loss value of the generating network according to the second judgment result; The target loss value includes the first loss value and the second loss value.

37. The method according to claim 36, wherein Determining a first loss value of the judgment network according to the first judgment result and the second judgment result includes: Determining first judgment expected data according to the first judgment result and the first data coefficient; Determining second decision expected data according to the second decision result and the second data coefficient; wherein the second data coefficient is smaller than the first data coefficient; A first loss value of the decision network is determined according to the first decision expectation data and the second decision expectation data.

38. The method according to claim 36, characterized in that Determining a second loss value of the generating network according to the second judgment result includes: determining first generated expected data according to the second judgment result and the first data coefficient; Determining second generated expected data according to the label depth map and the training predicted depth map; A second loss value of the generation network is determined based on the first generation expected data and the second generation expected data.

39. The method according to claim 36, wherein The step of training the generating network according to the target loss value includes: According to the first loss value and the second loss value, parameters of the decision network and parameters of the generation network are alternately adjusted so that both the first loss value and the second loss value converge.

40. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the point cloud enhancement method according to any one of claims 1 to 39 is implemented.

41. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instructions are executed by a processor, the point cloud enhancement method according to any one of claims 1 to 39 is implemented.

42. A controller having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the point cloud enhancement method according to any one of claims 1 to 39 is implemented.

43. A point cloud enhancement system, characterized in that: The point cloud enhancement system comprises a controller as claimed in claim 42 and a radar connected to the controller; wherein, The radar is used to collect original point cloud data.

44. The point cloud enhancement system according to claim 43, characterized in that The radar includes a millimeter wave radar.

45. A vehicle, characterized in that: The vehicle comprises a controller as claimed in claim 42, or comprises a point cloud enhancement system as claimed in claim 43 or 44.