Image-assisted point cloud generation method, system, device and storage medium

Through the image-assisted point cloud generation method, the attention mechanism and normalized flow model are used to solve the problems of point cloud data uneven resolution and edge information loss in the existing technology, and high-precision three-dimensional point cloud map construction is achieved.

CN118505855BActive Publication Date: 2025-05-09NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410603220.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-05-09
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high precision when generating three-dimensional point cloud maps of unknown scenes, and the point cloud data acquired by lidar is uneven in resolution and edge area information is lost, which affects the efficiency of scene construction.

Method used

The image-assisted point cloud generation method is adopted to generate high-precision point cloud data through an attention mechanism based on point cloud density and a normalized flow model, combining VAE and GAN, combining the mapping relationship between environmental image data and point cloud data.

Benefits of technology

The representation ability of point cloud data is improved, especially in the sparse and edge areas of point cloud data, the representation ability of long-distance objects is enhanced, and the accuracy and efficiency of three-dimensional point cloud map construction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118505855B_ABST
    Figure CN118505855B_ABST
Patent Text Reader

Abstract

The present invention discloses an image-assisted point cloud generation method, system, device and storage medium, which relates to the field of point cloud generation technology, and includes the steps of: collecting environmental image data and point cloud data, and using an attention mechanism based on point cloud density to obtain the attention weight of the point cloud data; weighted fusion of the attention weights of the environmental image data, simulating the distribution space of the features of the environmental image data and the features of the point cloud data after weighted fusion, learning the reversible linear transformation between the feature distributions, and obtaining the mapping relationship between the data through the reversible linear transformation of the feature distribution space after fusion weights; according to the mapping relationship, the environmental image data is used as prior knowledge to obtain the corresponding point cloud data. The present invention highlights the image features corresponding to the sparse areas of the point cloud with the help of the attention mechanism, obtains the conversion relationship between different modal data, generates point cloud edge area data on the known environmental image data, and improves the point cloud edge perception range and the accuracy of the point cloud edge area data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud generation, and in particular to an image-assisted point cloud generation method, system, device and storage medium. Background Art

[0002] Point cloud generation technology has been widely used in many fields, including robotics, autonomous driving, and the military. Point cloud models can provide detailed three-dimensional environmental information, which is essential for environmental perception and scene understanding. In the field of robotics, when a robot moves in an unknown environment, it needs to obtain spatial information about the surrounding area. By carrying sensors such as lidar, the robot obtains a point cloud map of the environment. In the field of autonomous driving, autonomous vehicles need to accurately perceive the surrounding environment, including roads, traffic signs, vehicles, and pedestrians. The point cloud data collected by sensors such as lidar provides vehicles with high-resolution three-dimensional environmental information to help vehicles make corresponding decisions. In the military field, troops need to use robots equipped with sensors such as lidar to accurately map the battlefield environment in order to perform related tasks.

[0003] However, due to the limitation of hardware resources, the resolution of point cloud data acquired by LiDAR is uneven, and the point cloud data has sparse edges, which leads to information loss in the edge areas of the point cloud. In the process of using point cloud data for scene construction, sparse edges means that more agents are needed to perform multiple perception tasks. When facing extreme terrains such as gullies and mountains that ground agents cannot reach, relying solely on LiDAR perception data can no longer meet the task requirements.

[0004] In addition, when building a scene, it is necessary to select a larger point cloud overlapping area for retrieval and calibration. When building a three-dimensional scene in some low-overlapping areas such as the edge of the point cloud, the error caused by sparse data will greatly affect the accuracy of point cloud registration, thereby affecting the efficiency of three-dimensional point cloud scene construction. To solve the problems of insufficient edge representation capabilities and data sparsity of point cloud data perceived by sensors such as lidar, using environmental image data to generate point cloud data is undoubtedly a good way, because environmental image data has uniform data distribution and long perception distance, and its representation capabilities for distant objects are usually stronger than point cloud data. When point cloud generation technology is used to build a map, the point cloud data generated in real time is fused with the previous map data to improve the accuracy and range of the point cloud data obtained by the intelligent agent. It has great practical value.

[0005] Existing point cloud generation technologies usually learn the process from single noise to original point cloud, and the process from original distribution diffusion to noise distribution. By using GAN, VAE and other architectures to reduce the distance between random noise and real point cloud, point cloud data that tends to be real point cloud is generated. However, these methods consume a lot of memory space and are directly generated from scratch without any prior knowledge. The model converges slowly, the generation efficiency is low, and it can only generate point cloud data with dense center and sparse edges. This is similar to the point cloud data representation effect directly perceived by LiDAR, and cannot achieve the purpose of enhancing the point cloud data representation ability. In addition, it is difficult to achieve high precision in the construction of 3D point cloud maps of unknown scenes. Summary of the invention

[0006] The purpose of the present invention is to provide an image-assisted point cloud generation method, system, device and storage medium to address the deficiencies of the above-mentioned prior art, so as to solve the problem that the prior art cannot be applied in real scenes and it is difficult to achieve high precision in constructing three-dimensional point cloud maps of unknown scenes.

[0007] The present invention specifically provides the following technical solution: an image-assisted point cloud generation method, comprising the following steps:

[0008] Collect environmental image data and point cloud data, and use the attention mechanism based on point cloud density to obtain the attention weight of point cloud data;

[0009] Performing weighted fusion of attention weights on the environment image data, and using VAE to simulate the distribution space of features of the environment image data and features of the point cloud data after weighted fusion;

[0010] The normalized flow model NFs is used to learn the reversible linear transformation from the feature distribution of the weighted fused environment image data to the feature distribution space of the real point cloud data, and the mapping relationship between the feature distribution of the weighted fused environment image data and the feature distribution space of the real point cloud data is obtained;

[0011] According to the mapping relationship, the environmental image data is used as prior knowledge to obtain corresponding point cloud data.

[0012] Preferably, the step of obtaining the attention weight of the point cloud data using the attention mechanism based on the point cloud density comprises the following steps:

[0013] Use the nearest neighbor search algorithm to calculate the number of points in the neighborhood near the point cloud data and obtain the density value of the point cloud;

[0014] The calculated density value is normalized by the softmax function, and the normalized density value is used as the attention score of each point.

[0015] Preferably, the normalized density value is used as the attention score of each point, and the specific expression is:

[0016]

[0017] Among them, d i represents the normalized density value of point i, a is a parameter used to adjust the relationship between the density value and the attention weight, and N is the total number of points in the point cloud.

[0018] Preferably, the environment image data is subjected to weighted fusion of attention weights, and the specific expression is:

[0019]

[0020] Among them, Attention(d i ) represents the attention score of the i-th point, Image Feature is the image feature, and N is the total number of points in the point cloud.

[0021] Preferably, the mapping relationship between the feature distribution of the weighted fused environment image data and the feature distribution space of the real point cloud data is specifically expressed as follows:

[0022]

[0023] Among them, z is a sample from the feature distribution of the weighted fusion environment image data, x' is a sample from the feature distribution of the point cloud, A is a learnable linear transformation matrix, b is a learnable bias vector, ° represents the composite of the function, f i is the mapping coefficient, and affine is the mapping function.

[0024] Preferably, after obtaining the corresponding point cloud data by taking the environmental image data as prior knowledge according to the mapping relationship, the Pointnet++ model is used as a discriminator to extract the features between the generated point cloud data obtained by generator reasoning and the real point cloud data, and the realism of the generated point cloud data is judged by calculating the cosine distance between the features of the generated point cloud data and the features of the real point cloud data.

[0025] Preferably, when calculating the cosine distance between the generated point cloud data features and the real point cloud data features, the cosine distance is optimized and minimized, and the parameters of the discriminator are updated so that the discriminator can reduce the cosine distance between the two point clouds to a minimum. The specific expression is:

[0026]

[0027] Among them, A and B are the feature vectors of two point clouds, · represents the dot product of the vectors, ||·|| represents the norm of the vector, and Cosine Similarity is the cosine function.

[0028] The present invention provides an image-assisted point cloud generation method system, comprising:

[0029] The acquisition module is used to collect environmental image data and point cloud data, and use the attention mechanism based on point cloud density to obtain the attention weight of the point cloud data;

[0030] A fusion module, used for performing weighted fusion of attention weights on the environment image data, and using VAE to simulate the distribution space of features of the environment image data and features of the point cloud data after weighted fusion;

[0031] A conversion module is used to use the normalized flow model NFs to learn the reversible linear transformation of the feature distribution of the weighted fused environment image data to the feature distribution space of the real point cloud data, and obtain the mapping relationship between the feature distribution of the weighted fused environment image data and the feature distribution space of the real point cloud data;

[0032] A generation module is used to obtain corresponding point cloud data using the environmental image data as prior knowledge according to the mapping relationship.

[0033] The present invention provides a computer device, comprising a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the image-assisted point cloud generation method.

[0034] The present invention provides a storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the image-assisted point cloud generation method are implemented.

[0035] Compared with the prior art, the present invention has the following significant advantages:

[0036] The present invention uses the attention mechanism to focus on the features corresponding to the sparse areas of the point cloud data according to the density value of the point cloud data, and then uses NFs to achieve reversible linear transformation between sample data that conforms to the image feature distribution after weighted fusion and sample data that conforms to the point cloud feature distribution. The mapping relationship between cross-modal data features is directly learned in the generation module, and the mapping parameters between different modal data features are obtained to achieve the generation of data in the edge area of ​​the point cloud based on the known environmental image data. The prior knowledge of the environmental image data is used to improve the convergence speed of the entire model during the training process. At the same time, with the help of the GAN unsupervised learning method, the distance between the generated point cloud data and the real point cloud data is optimized without labels. Through the adversarial learning process, the model is made more robust, and has a stronger ability to resist noise and changes, greatly reducing the overall training cost of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is the overall flow chart of the present invention;

[0038] Figure 2 A pseudo code representation diagram of the method for generating point clouds from images using normalized flow based on the attention mechanism of the present invention;

[0039] Figure 3 This is a pseudocode representation of the reversible linear transformation of the flow model of the present invention for learning weighted image feature distribution and point cloud feature distribution. DETAILED DESCRIPTION

[0040] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0041] Existing point cloud generation algorithms can be divided into flow-based, variational encoder (VAE)-based, and GAN-based algorithms. Although the VAE-based point cloud generation model can process continuous and discrete data and generate new samples, its generation process is random and cannot guarantee that each sample is of high quality, so training VAE requires a lot of computing resources and time. The flow-based point cloud generation model is usually used to model and learn data distribution to generate realistic point cloud data. However, although the flow-based generation model solves the problem of single generated samples in the variational encoder-based generation model, the model itself is very complex and difficult to generalize to other tasks.

[0042] At present, only a small number of works are based on GAN algorithms, using environmental image data to generate point cloud data. Although point cloud data can be generated based on a single environmental image data, the shape of the generated point cloud and the density of the generated point cloud must also be considered. The shape of point cloud data in a real environment is very complex, and generating point cloud data of a specific shape cannot meet the needs in a real environment. In addition, the point cloud data generated by the existing methods is also accompanied by problems such as uneven point cloud data acquired by sensors.

[0043] Based on the above related work, it is concluded that the existing point cloud generation technology has problems such as requiring a lot of computing resources and time overhead, high model training cost, fixed shape of point cloud data generated by the model, and low resolution.

[0044] In view of the deficiencies of the prior art, the present invention proposes an image-assisted point cloud generation method, system, device and storage medium to solve the problems of the prior art.

[0045] The present invention mainly solves the problems of sparse and uneven point cloud data and insufficient representation capability for distant objects in point cloud generation tasks. This method simulates the distribution space of environmental image data and point cloud data features with the help of flow models, and learns the conversion parameters between environmental image data and point cloud data by subjecting the sample data sampled in the distribution to multiple linear transformations; uses GAN to design a method for comparing the cosine similarity between the generated point cloud data features and the real point cloud data features, and efficiently generates point cloud data that tends to be realistic through the prior knowledge provided by the environmental image data. With the help of the attention mechanism, the point cloud density in the sparse area of ​​the point cloud data is directionally enhanced, thereby improving the representation capability of the point cloud data for distant objects.

[0046] This method mainly includes two stages: model training stage and joint optimization stage.

[0047] like Figure 1 As shown, an embodiment of the present application provides an image-assisted point cloud generation method, comprising the following steps:

[0048] Step S1: In the model training stage, the environment image data and point cloud data are first collected, and the attention weight of the point cloud data is obtained using the attention mechanism based on the point cloud density.

[0049] In this step, the attention weight of the point cloud data is obtained using the attention mechanism based on point cloud density, including the following steps:

[0050] The nearest neighbor search algorithm is used to calculate the number of points in the neighborhood near the point cloud data to obtain the density value of the point cloud.

[0051] The calculated density value is normalized by the softmax function, and the normalized density value is used as the attention score of each point.

[0052] Among them, the normalized density value is used as the attention score of each point. The specific expression is:

[0053]

[0054] Among them, d i represents the normalized density value of point i, a is a parameter used to adjust the relationship between the density value and the attention weight, and N is the total number of points in the point cloud.

[0055] A variational encoder (VAE) is used to learn the mean and variance of the environment image data and point cloud data.

[0056] Step S2: simulate the feature distribution space of the environmental image data and the feature distribution space of the point cloud data, perform weighted fusion of the environmental image data with attention weights, and use VAE to simulate the distribution space of the features of the environmental image data and the features of the point cloud data after weighted fusion.

[0057] In this step, the specific expression is:

[0058]

[0059] Among them, Attention(d i ) represents the attention score of the i-th point, Image Feature is the image feature, and N is the total number of points in the point cloud.

[0060] Step S3: Use the normalized flow model NFs to learn the reversible linear transformation of the feature distribution of the weighted fused environmental image data to the feature distribution space of the real point cloud data, and obtain the mapping relationship between the feature distribution space of the weighted fused environmental image data and the feature distribution space of the real point cloud data through the reversible linear transformation of the feature distribution space of the environmental image data and the feature distribution space of the point cloud data.

[0061] In this step, the specific expression of the mapping relationship is:

[0062]

[0063] Among them, z is a sample from the feature distribution of the weighted fusion environment image data, x' is a sample from the feature distribution of the point cloud, A is a learnable linear transformation matrix, b is a learnable bias vector, represents the composition of functions, f i is the mapping coefficient, and affine is the mapping function.

[0064] Step S4: According to the conversion parameters, the environmental image data is used as prior knowledge to obtain the corresponding point cloud data.

[0065] After obtaining the corresponding point cloud data using the environmental image data as prior knowledge according to the mapping relationship, the Pointnet++ model is used as a discriminator to extract the features between the generated point cloud data obtained by generator reasoning and the real point cloud data. The realism of the generated point cloud data is determined by calculating the cosine distance between the features of the generated point cloud data and the features of the real point cloud data.

[0066] In this step, the cosine similarity between the generated point cloud data features and the real point cloud data features is calculated to determine the realism of the point cloud data. That is, the cosine distance is minimized using gradient descent, and the parameters of the discriminator are updated so that the discriminator can reduce the cosine distance between the two point clouds to the minimum. The specific expression is:

[0067]

[0068] Among them, A and B are the feature vectors of two point clouds, · represents the dot product of the vectors, ||·|| represents the norm of the vector, and Cosine Similarity is the cosine function.

[0069] In the joint optimization stage, the generator optimizes the reversible mapping function of the weighted image features and the point cloud features to make the data generated from the image features close to the real data in the feature space. The discriminator determines the generated point cloud data as false data by maximizing the cosine loss between the generated point cloud data and the real point cloud data. Figure 2 and Figure 3 As shown, Figure 2 Pseudocode representation of the method for generating point clouds from images using normalized flow based on the attention mechanism; Figure 3 Pseudocode representation of the reversible linear transformation of weighted image feature distribution and point cloud feature distribution for flow model learning.

[0070] Based on the above methods and statements, the present invention provides an image-assisted point cloud generation method system, including: an acquisition module, a fusion module, a conversion module and a generation module.

[0071] Among them, the acquisition module is used to collect environmental image data and point cloud data, and use the attention mechanism based on point cloud density to obtain the attention weight of the point cloud data; the fusion module is used to perform weighted fusion of the attention weights on the environmental image data, and use VAE to simulate the distribution space of the features of the environmental image data and the features of the point cloud data after weighted fusion; the conversion module is used to use the normalized flow model NFs to learn the reversible linear transformation of the feature distribution of the environmental image data after weighted fusion to the feature distribution space of the real point cloud data, and obtain the mapping relationship between the feature distribution of the environmental image data after weighted fusion and the feature distribution space of the real point cloud data; the generation module is used to obtain the corresponding point cloud data based on the mapping relationship by taking the environmental image data as prior knowledge.

[0072] The fusion module realizes weighted fusion of environmental image data according to the point cloud density, and directionally generates missing parts of point cloud data, which is used for VAE simulation of the distribution space of weighted environmental image data and point cloud data features.

[0073] The present invention also provides a computer device, including a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of an image-assisted point cloud generation method.

[0074] In accordance with the disclosed embodiments, a computing device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth communications, etc.), or with any device (e.g., routers, modems, etc.) that enables a computing device to communicate with one or more other computing devices.

[0075] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of an image-assisted point cloud generation method are implemented.

[0076] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, such as but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0077] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. For technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. An image-assisted point cloud generation method, characterized in that: include: Collect environmental image data and point cloud data, and use the attention mechanism based on point cloud density to obtain the attention weight of point cloud data; Performing weighted fusion of attention weights on the environment image data, and using VAE to simulate the distribution space of features of the environment image data and features of the point cloud data after weighted fusion; The normalized flow model NFs is used to learn the reversible linear transformation from the feature distribution of the weighted fused environment image data to the feature distribution space of the real point cloud data, and the mapping relationship between the feature distribution of the weighted fused environment image data and the feature distribution space of the real point cloud data is obtained; According to the mapping relationship, the environmental image data is used as prior knowledge to obtain corresponding point cloud data; The method of using the attention mechanism based on point cloud density to obtain the attention weight of point cloud data includes the following steps: Use the nearest neighbor search algorithm to calculate the number of points in the neighborhood near the point cloud data and obtain the density value of the point cloud; The calculated density value is normalized by the softmax function, and the normalized density value is used as the attention score of each point; The normalized density value is used as the attention score of each point. The specific expression is: Among them, d i represents the normalized density value of point i, a is a parameter used to adjust the relationship between the density value and the attention weight, and N is the total number of points in the point cloud; After obtaining the corresponding point cloud data using the environmental image data as prior knowledge according to the mapping relationship, the Pointnet++ model is used as a discriminator to extract the features between the generated point cloud data obtained by generator reasoning and the real point cloud data. The realism of the generated point cloud data is determined by calculating the cosine distance between the features of the generated point cloud data and the features of the real point cloud data.

2. The image-assisted point cloud generation method according to claim 1, characterized in that: The environmental image data is weightedly fused with attention weights. The specific expression is as follows: Among them, Attention(d i ) represents the attention score of the i-th point, Image Feature is the image feature, and N is the total number of points in the point cloud.

3. The image-assisted point cloud generation method according to claim 1, characterized in that: The mapping relationship between the feature distribution of the environment image data after weighted fusion and the feature distribution space of the real point cloud data is specifically expressed as follows: Among them, z is a sample from the feature distribution of the weighted fusion environment image data, x' is a sample from the feature distribution of the point cloud, A is a learnable linear transformation matrix, b is a learnable bias vector, represents the composition of functions, f i is the mapping coefficient, and affine is the mapping function.

4. The image-assisted point cloud generation method according to claim 1, characterized in that: When the cosine distance between the generated point cloud data features and the real point cloud data features is calculated, the cosine distance is optimized and minimized, and the parameters of the discriminator are updated so that the discriminator can reduce the cosine distance between the two point clouds to the minimum. The specific expression is: Among them, A and B are the feature vectors of two point clouds, · represents the dot product of the vectors, ||·|| represents the norm of the vector, and Cosine Similarity is the cosine function.

5. An image-assisted point cloud generation method system, characterized in that: include: The acquisition module is used to collect environmental image data and point cloud data, and use the attention mechanism based on point cloud density to obtain the attention weight of the point cloud data; A fusion module, used for performing weighted fusion of attention weights on the environment image data, and using VAE to simulate the distribution space of features of the environment image data and features of the point cloud data after weighted fusion; A conversion module is used to use the normalized flow model NFs to learn the reversible linear transformation of the feature distribution of the weighted fused environment image data to the feature distribution space of the real point cloud data, and obtain the mapping relationship between the feature distribution of the weighted fused environment image data and the feature distribution space of the real point cloud data; A generation module, used to obtain corresponding point cloud data by using the environmental image data as prior knowledge according to the mapping relationship; The method of using the attention mechanism based on point cloud density to obtain the attention weight of point cloud data includes the following steps: Use the nearest neighbor search algorithm to calculate the number of points in the neighborhood near the point cloud data and obtain the density value of the point cloud; The calculated density value is normalized by the softmax function, and the normalized density value is used as the attention score of each point; The normalized density value is used as the attention score of each point. The specific expression is: Among them, d i represents the normalized density value of point i, a is a parameter used to adjust the relationship between the density value and the attention weight, and N is the total number of points in the point cloud; After obtaining the corresponding point cloud data using the environmental image data as prior knowledge according to the mapping relationship, the Pointnet++ model is used as a discriminator to extract the features between the generated point cloud data obtained by generator reasoning and the real point cloud data. The realism of the generated point cloud data is determined by calculating the cosine distance between the features of the generated point cloud data and the features of the real point cloud data.

6. A computer device, characterized in that: It comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of an image-assisted point cloud generation method as claimed in any one of claims 1 to 4.

7. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of an image-assisted point cloud generation method according to any one of claims 1 to 4 are implemented.