Sparse point cloud data completion method for automatic driving and related equipment

By building a point cloud completion model, using the technical means of coding modules and generation modules, the problems of insufficient global structure information and limited local geometric feature expression in sparse point cloud data completion are solved, high-quality point cloud completion is achieved, and the environmental perception ability of the autonomous driving system is improved.

CN120219692AActive Publication Date: 2025-06-27湖南工商大学
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510696332.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient global structural information, limited local geometric feature expression and blurred obstacle boundaries in sparse point cloud data completion, which affects the environmental perception and object recognition capabilities of the autonomous driving system.

Method used

A sparse point cloud data completion method for autonomous driving is proposed. By building a point cloud completion model including coding module and generation module, using training sparse point cloud for model training, and optimizing model parameters to improve the quality of completion. The encoding module extracts global features through the category feature attention unit, and the generation module gradually optimizes point cloud completion through upsampling and category feature attention unit.

Benefits of technology

The consistency of the complete point cloud is improved, the environment perception capability of the autonomous driving system is enhanced, and the generalization capability and completion accuracy are ensured in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219692A_ABST
    Figure CN120219692A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving-oriented sparse point cloud data completion method and related equipment, a first category feature attention unit is introduced into a coding module of a point cloud completion model, and through joint modeling of category information and point cloud geometric features, a completion network is enabled to complete point cloud data in a point cloud completion process. Structural features of different types of objects can be recognized more accurately, features of sparse point clouds are dynamically adjusted through a second type feature attention unit in the generation module, it is ensured that the model can adapt to different environments, the generalization ability under multiple scenes is improved, and the robustness of the model is improved. And through an up-sampling unit, a third category feature attention unit and a sixth multi-layer perceptron in the generation module, the whole complementing process of the sparse point cloud is gradually optimized from coarse to fine, so that the consistency of complementing the point cloud is improved, and the environmental perception capability of the automatic driving system is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud completion, and particularly to a sparse point cloud data completion method and related devices for autonomous driving. Background Art

[0002] With the rapid development of autonomous driving technology, the demand for three-dimensional environment perception is increasing day by day. As an important data form for three-dimensional environment perception, point cloud data can provide rich spatial and geometric information, which makes three-dimensional point cloud data widely used in autonomous driving. However, due to the limitations of sensor acquisition such as lidar or environmental occlusion, point cloud data often has sparsity and incompleteness, which directly affects the environment perception and object recognition capabilities of autonomous driving vehicles. Therefore, many existing technologies have explored the research on efficient sparse point cloud data completion, and these existing technologies mainly include completion methods based on geometric interpolation and completion methods based on rules.

[0003] In recent years, deep learning methods have shown great potential in the task of point cloud data completion. In particular, methods based on models such as generative adversarial networks and variational autoencoders can learn the global and local features of point clouds to a certain extent and achieve high-quality completion. However, there are still the following problems when relying solely on sparse point clouds for completion: 1. Insufficient global structure information, making it difficult to restore the complete environment: Due to the limited local information of sparse point clouds, when relying solely on the existing point clouds for completion without sufficient context information, the generated point cloud structure may be incomplete or distorted in shape. In the autonomous driving scenario, there may be only a small amount of point data for distant vehicles or obstacles, and direct completion may lead to inaccurate shape reconstruction, affecting the accuracy of target detection and path planning; 2. Limited expression of local geometric features, affecting the completion accuracy: Point cloud data is essentially unstructured, and traditional completion methods (such as interpolation) are difficult to accurately capture the geometric features of complex shapes. In the case of sparse point clouds, local geometric features may not be sufficient to accurately infer the details of the object surface, resulting in deformation or even pseudo-features in the completed point cloud, thereby affecting the stability of target detection and tracking tasks; 3. Blurred obstacle boundaries: The autonomous driving system requires high-precision environment information for path planning and obstacle avoidance. The defects of sparse point clouds will lead to blurred obstacle boundaries, making it difficult for vehicles to accurately judge the drivable area. Summary of the Invention

[0004] The present invention provides a sparse point cloud data completion method and related devices for autonomous driving, and its purpose is to improve the consistency of the completed point cloud to enhance the environment perception ability of the autonomous driving system.

[0005] To achieve the above object, the present invention provides a sparse point cloud data completion method for autonomous driving, including: Step 1, obtaining sparse point clouds for training in an autonomous driving scenario; Step 2, constructing a point cloud completion model, which includes an encoding module and a generation module; Step 3, training the point cloud completion model using the sparse point clouds for training, and optimizing the parameters of the trained point cloud completion model using a point cloud completion loss and a class prediction loss to obtain a trained point cloud completion model; Step 4, inputting the sparse point cloud data of the target autonomous driving vehicle into the trained point cloud completion model for completion to obtain the dense point cloud of the target autonomous driving vehicle; The encoding module includes a splicing unit, a first multi-layer perceptron, a first class feature attention unit for fusing class attention and feature attention, and a max pooling unit connected in sequence; The generation module includes a second multi-layer perceptron, a second class feature attention unit for fusing class attention and feature attention, a third multi-layer perceptron, a farthest point sampling unit, a fourth multi-layer perceptron, a fifth multi-layer perceptron, a splicing unit, an upsampling unit, a third class feature attention unit for fusing class attention and feature attention, and a sixth multi-layer perceptron.

[0006] Furthermore, before training the point cloud completion model using the sparse point clouds for training, it further includes: Calculating the centroid of the sparse point clouds for training; Translating all the sparse point clouds for training to a new coordinate system with the centroid of the point cloud as the origin; Calculating the maximum distance from each point cloud in the sparse point clouds for training to the centroid of the point cloud, and normalizing all the sparse point clouds for training based on the maximum distance to obtain the normalized sparse point clouds for training; Using one-hot encoding to convert the class label of each point cloud in the sparse point clouds for training into a class embedding vector to obtain a class embedding vector matrix; Performing denoising processing on the normalized sparse point clouds for training through a local neighborhood analysis method to obtain denoised sparse point clouds for training.

[0007] Furthermore, the generation module includes: Processing the global features sequentially through a second multi-layer perceptron, a second class feature attention unit, a third multi-layer perceptron, and a farthest point sampling unit to generate a seed point cloud; Processing the seed point cloud through a fourth multi-layer perceptron to obtain a processed seed point cloud; Processing the global features through a fifth multi-layer perceptron to obtain processed global features; The processed seed point cloud and the processed global features are concatenated by a concatenation unit to obtain a seed point cloud feature matrix; The seed point cloud feature matrix is cyclically optimized by an upsampling unit, a third-category feature attention unit, and a sixth multi-layer perceptron to generate a dense point cloud.

[0008] Furthermore, both the first-category feature attention unit and the second-category feature attention unit include: A category attention layer, a feature attention layer, a fusion layer, a residual connection layer, a normalization layer, and a feed-forward network layer; The input ends of the category attention layer and the feature attention layer are both the input ends of the first-category feature attention unit and the second-category feature attention unit; The output ends of the category attention layer and the feature attention layer are both connected to the input end of the fusion layer; The output end of the fusion layer and the input end of the feature attention layer are both connected to the input end of the residual connection layer. The output end of the residual connection layer is connected to the input end of the normalization layer, and the output end of the normalization layer is connected to the input end of the feed-forward network layer; The output end of the feed-forward network layer is the output end of the first-category feature attention unit and the second-category feature attention unit.

[0009] Furthermore, the category attention layer is used to: Perform a Softmax operation after taking the inner product of the input category embedding vector matrix to obtain a category attention weight matrix. The calculation expression is: where represents the category attention weight matrix, represents the input category embedding vector matrix, represents the transposed category embedding vector matrix.

[0010] Furthermore, the feature attention layer is used to: Use a linear transformation to transform the input point feature matrix to obtain a query matrix and a key matrix; Use the query matrix and the key matrix to transform and obtain a feature attention weight matrix. The expression is: where represents the feature attention weight matrix, represents the query matrix, represents a key matrix, represents a value matrix, represents the dimension of the point feature matrix, represents a query weight matrix, represents a key weight matrix, represents a value weight matrix, represents the input point feature matrix.

[0011] Furthermore, the point cloud completion loss is used to measure the geometric difference between the generated point cloud and the real point cloud, and its calculation expression is: The class prediction loss is used to measure the accuracy of the model's point cloud class prediction, and its calculation expression is: where, represents the value of the point cloud completion loss, represents the number of point clouds, represents the th point in the real point cloud, represents the th point in the real point cloud, represents the value of the class prediction loss, represents the probability that the th point predicted by the model belongs to its real class .

[0012] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a sparse point cloud data completion method for autonomous driving.

[0013] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a sparse point cloud data completion method for autonomous driving.

[0014] The above solution of the present invention has the following beneficial effects: The present invention obtains sparse point clouds for training in an autonomous driving scenario; constructs a point cloud completion model including an encoding module and a generation module; trains the point cloud completion model using the sparse point clouds for training, and optimizes the parameters of the trained point cloud completion model using a point cloud completion loss and a class prediction loss to obtain a trained point cloud completion model; inputs the sparse point cloud data of a target autonomous vehicle into the trained point cloud completion model for completion to obtain the dense point cloud of the target autonomous vehicle; the encoding module includes a splicing unit, a first multi-layer perceptron, a first class feature attention unit, and a max pooling unit connected in sequence; the generation module includes a second multi-layer perceptron, a second class feature attention unit for fusing class attention and feature attention, a third multi-layer perceptron, a farthest point sampling unit, a fourth multi-layer perceptron, a fifth multi-layer perceptron, a splicing unit, an upsampling unit, a third class feature attention unit for fusing class attention and feature attention, and a sixth multi-layer perceptron; compared with the prior art, the present invention introduces a first class feature attention unit in the encoding module of the point cloud completion model, and through the joint modeling of class information and point cloud geometric features, enables the completion network to more accurately identify the structural features of different class objects during the process of completing the point cloud. The second class feature attention unit in the generation module dynamically adjusts the features of the sparse point cloud to ensure that the model can adapt to different environments and improve the generalization ability in multiple scenarios. The upsampling unit, the third class feature attention unit, and the sixth multi-layer perceptron in the generation module gradually optimize the process of completing the sparse point cloud from coarse to fine, thereby improving the consistency of the completed point cloud and further enhancing the environmental perception ability of the autonomous driving system.

[0015] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic flowchart of an embodiment of the present invention; Figure 2 is a schematic structural diagram of the point cloud completion model in an embodiment of the present invention; Figure 3 is a schematic structural diagram of the class feature attention unit in an embodiment of the present invention; Figure 4 is a schematic structural diagram of the terminal device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will describe in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0019] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0020] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0021] The present invention provides a sparse point cloud data completion method and related devices for autonomous driving in view of existing problems.

[0022] As Figure 1 、 Figure 2 shown, an embodiment of the present invention provides a sparse point cloud data completion method for autonomous driving, including: Step 1, obtaining training sparse point clouds in an autonomous driving scenario; Step 2, constructing a point cloud completion model, where the point cloud completion model includes an encoding module and a generation module; Step 3, training the point cloud completion model using the training sparse point clouds, and optimizing the parameters of the trained point cloud completion model using the point cloud completion loss and the class prediction loss to obtain the trained point cloud completion model; Step 4, inputting the sparse point cloud data of the target autonomous driving vehicle into the trained point cloud completion model for completion to obtain the dense point cloud of the target autonomous driving vehicle.

[0023] To improve the consistency and quality of the model input, the embodiments of the present invention need to preprocess the obtained training sparse point clouds to ensure the normalization of the input data and the effective utilization of class information; therefore, before training the point cloud completion model using the training sparse point clouds, it further includes: Calculate the centroid of the sparse point cloud for training. The calculation formula is: where, represents the centroid of the sparse point cloud for training , represents the number of the sparse point clouds for training, represents the th sparse point cloud in the sparse point clouds for training; Translate all the sparse point clouds for training to a new coordinate system with the centroid of the point cloud as the origin. The translation formula is: where, represents the th sparse point cloud in the sparse point clouds for training in the new coordinate system; Calculate the maximum distance from each point cloud in the sparse point clouds for training to the centroid of the point cloud. The calculation expression is: where, represents the maximum distance, represents the distance from the , th sparse point cloud in the sparse point clouds for training in the new coordinate system to the centroid of the point cloud, represents the coordinates of the Based on the maximum distance, normalize all the sparse point clouds for training, scale the unit range of all the sparse point clouds for training, and obtain the normalized sparse point clouds for training. The normalization expression is: where, represents the th sparse point cloud in the normalized sparse point clouds for training, Use one-hot encoding to convert the class label of each point cloud in the sparse point clouds for training into a class embedding vector, and obtain the class embedding vector matrix ; Denoise the normalized sparse point clouds for training through the local neighborhood analysis method to obtain the denoised sparse point clouds for training.

[0024] Since removing outliers and noise points is an important step in data preprocessing, especially when processing point cloud data, in the embodiments of the present invention, the normalized sparse point clouds for training are denoised through the local neighborhood analysis method. The specific process is as follows: For each point cloud in the training sparse point cloud after standardization, calculate its set of neighboring points. The calculation expression is as follows: Wherein, represents the training sparse point cloud after standardization in the th point cloud and the th point cloud represents the training sparse point cloud after standardization in the th point cloud Select the k point clouds closest to the th point cloud from the training sparse point cloud after standardization to form a k-nearest neighbor point set ; Use principal component analysis to calculate the local normal vector of the th point cloud in the training sparse point cloud after standardization , and obtain: Calculate the angular deviation between the th point cloud in the training sparse point cloud after standardization and its neighboring points: Wherein, represents the included angle between two points, represents the norm of the local normal vector.

[0025] If the angular deviation of the th point cloud in the training sparse point cloud after standardization is significantly greater than the average deviation of its neighboring points, it is considered an outlier and removed.

[0026] In the embodiment of the present invention, the main task of the encoding module in the point cloud completion model is to extract global features from the input sparse point cloud and enhance the understanding of the point cloud context by combining category information.

[0027] Specifically, the encoding module includes a splicing unit, a first multi-layer perceptron, a first category feature attention unit for fusing category attention and feature attention, and a max pooling unit connected in sequence; Perform preliminary feature extraction on the input sparse point cloud through the first multi-layer perceptron to obtain a point feature matrix. The expression is as follows: Among them, represents the point feature matrix, represents multi-layer perception processing, represents the real number field, represents the feature dimension. The input layer of the first multi-layer perceptron receives the preprocessed sparse point cloud , and the hidden layer of the first multi-layer perceptron uses multiple fully connected layers for feature extraction, and then outputs the point feature matrix through the output layer ; Input the point feature matrix and the category embedding vector matrix into the first category feature attention unit for category attention and feature attention fusion to obtain the updated point feature matrix. The expression is: Among them, represents the updated point feature matrix, represents the fusion operation, represents the category embedding vector matrix; Input the updated point feature matrix into the max pooling unit for max pooling operation to generate the global feature. The global feature contains the local geometric structure and category distribution information of the point cloud, and provides a context-aware shape encoding for subsequent generation tasks. The expression is: Among them, represents the global feature, represents the max pooling operation.

[0028] In the embodiment of the present invention, in the point cloud completion model, the generation module gradually generates a high-quality dense point cloud based on the global feature extracted by the encoding module.

[0029] Specifically, the generation module includes a second multi-layer perceptron, a second category feature attention unit for fusing category attention and feature attention, a third multi-layer perceptron, a farthest point sampling unit, a fourth multi-layer perceptron, a fifth multi-layer perceptron, a splicing unit, an upsampling unit, a third category feature attention unit for fusing category attention and feature attention, and a sixth multi-layer perceptron; Process the global feature through the second multi-layer perceptron, the second category feature attention unit, the third multi-layer perceptron, and the farthest point sampling unit in sequence to generate the seed point cloud; Process the seed point cloud through the fourth multi-layer perceptron to obtain the processed seed point cloud; Process the global feature through the fifth multi-layer perceptron to obtain the processed global feature; Splice the processed seed point cloud and the processed global feature through the splicing unit to obtain the seed point cloud feature matrix; The seed point cloud feature matrix is cyclically optimized by using an upsampling unit, a third-class feature attention unit, and a sixth multi-layer perceptron to generate a dense point cloud.

[0030] In the embodiment of the present invention, the generation module first maps the global feature to the initial point feature through the second multi-layer perceptron, and the expression is: where represents the initial point feature, represents the mapping operation of the second multi-layer perceptron; Subsequently, the category attention and feature attention are fused through the second-class feature attention unit: where is the feature matrix enhanced by the second-class feature attention unit; Then, the feature is mapped to the three-dimensional space through the third multi-layer perceptron to generate the initial seed point cloud: where represents the initial seed point cloud, represents the mapping operation of the third multi-layer perceptron; The spatial distribution of the initial seed point cloud is optimized through the farthest point sampling unit to obtain the seed point cloud: where represents the seed point cloud, represents the farthest point sampling process; The seed point cloud and the global feature are processed by the fourth multi-layer perceptron and the fifth multi-layer perceptron and then concatenated to generate the seed point cloud feature matrix: where and are respectively used to convert the seed point cloud and the global feature into matching feature matrices, represents the feature concatenation operation for merging and ; The feature matrix after concatenation is subjected to feature extraction through the upsampling unit to obtain the feature of the seed point cloud feature matrix: where is a linear transformation matrix used to expand the feature dimension, is the number of points after upsampling; Subsequently, the upsampled features are dynamically adjusted by the third category feature attention unit: The adjusted features are inversely mapped to the three-dimensional space through the sixth multi-layer perceptron to generate the initial dense point cloud :

[0031] The above upsampling, dynamic adjustment, and inverse mapping steps are repeated L times for the initial dense point cloud to gradually refine the point cloud resolution, and finally a complete and high-density dense point cloud is output : Among them, represents the number of upsampling layers. The progressive generation method can, to a certain extent, avoid the problems of blurring or overfitting caused by single upsampling.

[0032] According to the above process, the output of the generator is a complete dense point cloud, which not only retains the global shape information but also has rich local details and category perception ability.

[0033] Most preferably, the first category feature attention unit and the second category feature attention unit are designed to enhance the model's understanding ability of point cloud data by dynamically weighing the category information and local geometric structure information, as Figure 3 shown, both include: category attention layer, feature attention layer, fusion layer, residual connection layer, normalization layer, feed-forward network layer; The input ends of the category attention layer and the feature attention layer are both the input ends of the first category feature attention unit and the second category feature attention unit; The output ends of the category attention layer and the feature attention layer are both connected to the input end of the fusion layer; The output end of the fusion layer and the input end of the feature attention layer are both connected to the input end of the residual connection layer. The output end of the residual connection layer is connected to the input end of the normalization layer, and the output end of the normalization layer is connected to the input end of the feed-forward network layer; The output end of the feed-forward network layer is the output end of the first category feature attention unit and the second category feature attention unit.

[0034] In the embodiment of the present invention, the category feature attention unit can receive two inputs, namely the point feature matrix and the category embedding vector matrix. Among them, the point feature matrix is obtained by feature extraction from the preprocessed point cloud data, and the category embedding vector matrix It is obtained after one-hot encoding.

[0035] Specifically, the class attention layer is used for: Performing a Softmax operation after taking the inner product of the input class embedding vector matrix to obtain the class attention weight matrix. The calculation expression is: Among them, represents the class attention weight matrix, represents the input class embedding vector matrix, represents the transposed class embedding vector matrix.

[0036] Specifically, the feature attention layer is used for: Using a linear transformation to transform the input point feature matrix to obtain a query matrix and a key matrix; Using the query matrix and the key matrix to transform and obtain the feature attention weight matrix. The expression is: Among them, represents the feature attention weight matrix, represents the query matrix, represents the key matrix, represents the value matrix, represents the dimension of the point feature matrix, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the input point feature matrix, The purpose of the operation is to prevent numerical overflow.

[0037] Specifically, the fusion layer is used to weight-fuse the class attention weight matrix and the feature attention weight matrix to generate a comprehensive attention weight matrix. The expression is: Among them, represents the comprehensive attention weight matrix, represents a hyperparameter used to balance the contributions of classes and features; The fusion layer is also used to fuse the value matrix using the comprehensive attention weight matrix to generate an updated point feature matrix. The expression is: Among them, represents the updated point feature matrix, Represents a value matrix.

[0038] Specifically, the residual connection layer is used to perform a residual connection between the updated point feature matrix and the point feature matrix, and the normalization layer is used to normalize the point feature matrix after the residual connection to enhance the stability of the model. The expression is: Where, Represents the normalized point feature matrix, Represents the normalization operation.

[0039] Specifically, the feed-forward network layer is used to enhance the feature representation ability. The expression is: Where, Represents the output result of the category feature attention unit, Represents the enhancement operation, Represents the activation function, 、 Represents the weight matrix of the feed-forward network, 、 Represents the bias vector.

[0040] The objective of optimizing the parameters of the trained point cloud completion model using the point cloud completion loss and the category prediction loss in the embodiments of the present invention is to minimize the difference between the generated point cloud and the real point cloud, and at the same time, combine the category prediction information to improve the completion quality and classification accuracy of the model.

[0041] Specifically, the point cloud completion loss is used to measure the geometric difference between the generated point cloud and the real point cloud. The calculation expression is: The category prediction loss is used to measure the accuracy of the model's prediction of the point cloud category. The calculation expression is: Where, Represents the point cloud completion loss value, Represents the number of point clouds, Represents the th point in the real point cloud, Represents the th point in the real point cloud, Represents the category prediction loss value, Represents the probability that the th point predicted by the model belongs to its real category .

[0042] The comprehensive loss function combines the above two parts of losses with weights to form the final optimization objective, ensuring that the model maintains the accuracy of class prediction while completing the point cloud. The expression is as follows: Wherein, represents the comprehensive loss function, represents an adjustable hyperparameter, and .

[0043] In the embodiment of the present invention, sparse point clouds for training in the autonomous driving scenario are obtained; a point cloud completion model including an encoding module and a generation module is constructed; the point cloud completion model is trained using the training sparse point clouds, and the parameters of the trained point cloud completion model are optimized using the point cloud completion loss and the class prediction loss to obtain the trained point cloud completion model; the sparse point cloud data of the target autonomous driving vehicle is input into the trained point cloud completion model for completion to obtain the dense point cloud of the target autonomous driving vehicle; the encoding module includes a splicing unit, a first multi-layer perceptron, a first class feature attention unit, and a max pooling unit connected in sequence; the generation module includes a second multi-layer perceptron, a second class feature attention unit for fusing class attention and feature attention, a third multi-layer perceptron, a farthest point sampling unit, a fourth multi-layer perceptron, a fifth multi-layer perceptron, a splicing unit, an upsampling unit, a third class feature attention unit for fusing class attention and feature attention, and a sixth multi-layer perceptron; compared with the prior art, in the encoding module of the point cloud completion model in the embodiment of the present invention, a first class feature attention unit is introduced. Through the joint modeling of class information and point cloud geometric features, the completion network can more accurately identify the structural features of different class objects during the process of completing the point cloud. The second class feature attention unit in the generation module dynamically adjusts the features of the sparse point cloud to ensure that the model can adapt to different environments and improve the generalization ability in multiple scenarios. The upsampling unit, the third class feature attention unit, and the sixth multi-layer perceptron in the generation module enable the completion process of the sparse point cloud to be gradually optimized from coarse to fine, thereby improving the consistency of the completed point cloud and further enhancing the environmental perception ability of the autonomous driving system.

[0044] The embodiment of the present invention also provides a terminal device, as Figure 4 shown. The terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 4 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the above-mentioned method for completing sparse point cloud data for autonomous driving is implemented.

[0045] The terminal device D10 may be a computing device such as a desktop computer, a notebook, a palm computer, a server, a server cluster, and a cloud server. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art can understand that Figure 4 merely examples of the terminal device D10, which do not constitute a limitation on the terminal device D10, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0046] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and the processor D100 may also be other general-purpose processors, digital signal processors (DSPs, Digital Signal Processors), application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), off-the-shelf programmable gate arrays (FPGAs, Field-Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0047] The memory D101 may be an internal storage unit of the terminal device D10 in some embodiments, such as the hard disk or memory of the terminal device D10. The memory D101 may also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0048] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.

[0049] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0050] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements a sparse point cloud data completion method for autonomous driving.

[0051] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0052] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A sparse point cloud data completion method for autonomous driving, characterized in that, Including: Step 1: Obtain the sparse point cloud for training in the autonomous driving scenario; Step 2: Construct a point cloud completion model, where the point cloud completion model includes an encoding module and a generation module; Step 3: Use the sparse point cloud for training to train the point cloud completion model, and optimize the parameters of the trained point cloud completion model using the point cloud completion loss and the class prediction loss to obtain the trained point cloud completion model; Step 4: Input the sparse point cloud data of the target autonomous driving vehicle into the trained point cloud completion model for completion to obtain the dense point cloud of the target autonomous driving vehicle; The encoding module includes a splicing unit, a first multi-layer perceptron, a first class feature attention unit for fusing class attention and feature attention, and a max pooling unit connected in sequence; The generation module includes a second multi-layer perceptron, a second class feature attention unit for fusing class attention and feature attention, a third multi-layer perceptron, a farthest point sampling unit, a fourth multi-layer perceptron, a fifth multi-layer perceptron, a splicing unit, an upsampling unit, a third class feature attention unit for fusing class attention and feature attention, and a sixth multi-layer perceptron.

2. The method for completing sparse point cloud data for autonomous driving according to claim 1, wherein Before using the sparse point cloud for training to train the point cloud completion model, it further includes: Calculate the centroid of the sparse point cloud for training; Translate all the sparse point clouds for training to a new coordinate system with the centroid of the point cloud as the origin; Calculate the maximum distance from each point cloud in the sparse point cloud for training to the centroid of the point cloud, and normalize all the sparse point clouds for training based on the maximum distance to obtain the normalized sparse point cloud for training; Use one-hot encoding to convert the class label of each point cloud in the sparse point cloud for training into a class embedding vector to obtain a class embedding vector matrix; Perform denoising processing on the normalized sparse point cloud for training through a local neighborhood analysis method to obtain the denoised sparse point cloud for training.

3. The method for completing sparse point cloud data for autonomous driving according to claim 2, characterized in that The encoding module includes: Perform preliminary feature extraction on the input sparse point cloud through the first multi-layer perceptron to obtain a point feature matrix; Input the point feature matrix and the class embedding vector matrix into the first class feature attention unit for fusing class attention and feature attention to obtain an updated point feature matrix; Input the updated point feature matrix into the max pooling unit for max pooling operation to generate a global feature.

4. The method for completing sparse point cloud data for autonomous driving according to claim 3, characterized in that, The generation module includes: Process the global feature through the second multi-layer perceptron, the second class feature attention unit, the third multi-layer perceptron, and the farthest point sampling unit in sequence to generate a seed point cloud; Process the seed point cloud through the fourth multi-layer perceptron to obtain a processed seed point cloud; Process the global feature through the fifth multi-layer perceptron to obtain a processed global feature; Splice the processed seed point cloud and the processed global feature through the splicing unit to obtain a seed point cloud feature matrix; Using the upsampling unit, the third category feature attention unit, and the sixth multi-layer perceptron to perform cyclic optimization processing on the seed point cloud feature matrix to generate a dense point cloud.

5. The method for completing sparse point cloud data for autonomous driving according to claim 4, wherein Both the first category feature attention unit and the second category feature attention unit include: A category attention layer, a feature attention layer, a fusion layer, a residual connection layer, a normalization layer, and a feed-forward network layer; The input ends of the category attention layer and the feature attention layer are both the input ends of the first category feature attention unit and the second category feature attention unit; The output ends of the category attention layer and the feature attention layer are both connected to the input end of the fusion layer; The output end of the fusion layer and the input end of the feature attention layer are both connected to the input end of the residual connection layer. The output end of the residual connection layer is connected to the input end of the normalization layer, and the output end of the normalization layer is connected to the input end of the feed-forward network layer; The output end of the feed-forward network layer is the output end of the first category feature attention unit and the second category feature attention unit.

6. The method for completing sparse point cloud data for autonomous driving according to claim 5, wherein The category attention layer is used to: Perform a Softmax operation after taking the inner product of the input category embedding vector matrix to obtain a category attention weight matrix. The calculation expression is: Among them, represents the class attention weight matrix, represents the input class embedding vector matrix, represents the transposed class embedding vector matrix.

7. The method for completing sparse point cloud data for autonomous driving according to claim 5, characterized in that, The feature attention layer is used to: Use a linear transformation to transform the input point feature matrix to obtain a query matrix and a key matrix; Use the query matrix and the key matrix to transform and obtain a feature attention weight matrix. The expression is: Among them, represents the feature attention weight matrix, represents the query matrix, represents the key matrix, represents the value matrix, represents the dimension of the point feature matrix, represents the query weight matrix, represents the key weight matrix, represents the value weight matrix, represents the input point feature matrix.

8. The method for sparse point cloud data completion for autonomous driving according to claim 1, wherein The point cloud completion loss is used to measure the geometric difference between the generated point cloud and the real point cloud. The calculation expression is: The category prediction loss is used to measure the accuracy of the model's prediction of the point cloud category. The calculation expression is: Among them, represents the point cloud completion loss value, represents the number of point clouds, represents the th point in the real point cloud, represents the th point in the real point cloud, represents the class prediction loss value, represents the probability that the th point predicted by the model belongs to its real class .

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for sparse point cloud data completion for autonomous driving according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for sparse point cloud data completion for autonomous driving according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Three-dimensional point cloud semantic segmentation method based on local and global context awareness

    CN117218351A

  • Airborne point cloud classification method and system considering global-local self-attention mechanism

    CN119445208A

  • Weak supervision point cloud semantic segmentation method based on correlation enhancement

    CN119723095A

  • Data active selection and annotation method and apparatus for point cloud

    US20250077622A1

  • Method, apparatus and device for processing three-dimensional point cloud, and storage medium

    WO2022166400A1