Cross-modal large aircraft point cloud completion method

By combining three-dimensional laser scanner and monocular camera data, a deep learning model of cross-modal point cloud completion was constructed, and the traditional single-modal point cloud completion method was solved, and a higher quality aircraft point cloud completion was achieved, and the performance of related tasks was improved.

CN120147190APending Publication Date: 2025-06-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510201333.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional single-mode point cloud completion method has poor point cloud completion due to the low scanning resolution of the three-dimensional sensor and the lack of texture information, which is especially uncertain in the three-dimensional reconstruction and target detection tasks of large aircraft.

Method used

A cross-modal large aircraft point cloud completion method is adopted to build a cross-modal point cloud completion deep learning model by combining the point cloud acquired by a three-dimensional laser scanner and the image data obtained by a monocular camera. The model generates a full resolution complete aircraft point cloud through point cloud feature extractor, image encoder, cross-modal feature fusion module and FoldingNet module.

Benefits of technology

By introducing monocular image assistance, more geometric details and texture information are provided, the point cloud completion effect is significantly optimized, the uncertainty of missing parts of inference is reduced, and the performance of aircraft target detection and three-dimensional reconstruction tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147190A_ABST
    Figure CN120147190A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of aeronautical survey, in particular to a cross-modal large aircraft point cloud completion method, which comprises the following steps of: acquiring a complete point cloud of a large aircraft in a static state, shooting monocular images of the large aircraft from a plurality of visual angles by fixing the pose of a camera, and establishing a cross-modal large aircraft point cloud completion data set; constructing a cross-modal aircraft point cloud completion model for automatically completing the three-dimensional incomplete point cloud of the aircraft into a complete aircraft point cloud; training the cross-modal aircraft point cloud completion model by adopting the cross-modal large aircraft point cloud completion data set; and taking the incomplete point cloud to be complemented and the monocular image of the single aircraft as the input of the cross-modal aircraft point cloud complementation model, and outputting a complete aircraft point cloud with fine geometric details. According to the method, monocular image assistance is introduced, more geometric details and texture information can be provided, complementary information can be provided for the point cloud through the images of different visual angles, and therefore the point cloud complementation effect is better optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aerial surveying, and particularly to a cross-modal large aircraft point cloud completion method. Background Art

[0002] Three-dimensional object analysis and generation tasks such as virtual assembly of large aircraft require a high-precision (geometrically faithful) point cloud model as support. However, due to factors such as the low resolution and inherent errors of three-dimensional acquisition devices, the lighting and reflection properties of the large aircraft surface, and the occlusion between aircraft components, the scanned aircraft point cloud model inevitably has missing parts, seriously affecting subsequent analysis and understanding tasks. Therefore, it is necessary to complete the incomplete aircraft point cloud to obtain a complete point cloud with correct topological structure and geometric details.

[0003] However, the traditional single-modal point cloud completion methods have the following problems: 1. Due to the low scanning resolution of three-dimensional sensors, the obtained point cloud is usually sparse, and it is difficult to determine whether the blank three-dimensional space is caused by the sparsity of the point cloud itself or its incompleteness; 2. Only relying on the input incomplete point cloud for completion, the available geometric clues are limited, resulting in great uncertainty in inferring the missing area. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a cross-modal large aircraft point cloud completion method, and its remarkable beneficial effect lies in that, compared with the existing completion methods based only on single-modal point clouds, by introducing monocular images for assistance, it can not only provide more geometric details and texture information, but also provide complementary information for the point cloud through images from different perspectives, thereby better optimizing the point cloud completion effect.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: A cross-modal large aircraft point cloud completion method, the method comprising the following steps:

[0006] S1. Obtain the complete point cloud P of the large aircraft in a stationary state i , and monocular images I of the large aircraft taken from multiple perspectives with a fixed camera pose i , and establish a cross-modal large aircraft point cloud completion data set based on the complete point cloud P i and the monocular images I i ;

[0007] S2. Construct a cross-modal aircraft point cloud completion model for automatically completing the incomplete point cloud P of the aircraft in three dimensions in into the complete aircraft point cloud P out ;

[0008] S3. Train the cross-modal aircraft point cloud completion model using the cross-modal large aircraft point cloud completion data set;

[0009] S4. Use the incomplete point cloud P to be completed in and the monocular image I of a single aircraft i as the input of the cross-modal aircraft point cloud completion model, and output a complete aircraft point cloud with fine geometric details.

[0010] Furthermore, in step S1, the specific process includes the following steps:

[0011] S11. Stick coding points at various positions on the large aircraft, and perform an omnidirectional scan while the large aircraft is stationary to obtain a complete point cloud P containing the coordinate information of the coding points i , and use it as the ground truth of the cross-modal large aircraft point cloud completion dataset, that is, the ground truth complete point cloud G;

[0012] S12. Select multiple viewpoints and fix the monocular camera respectively to take pictures of the large aircraft to obtain the monocular images I of the large aircraft at each viewpoint i ;

[0013] S13. Reconstruct the complete point cloud P i , then select multiple different viewpoints, and discard the point clouds of the occluded parts at this viewpoint. Each incomplete point cloud P in is sampled to N points;

[0014] S14. Downsample the ground truth complete point cloud G to M points, and downsample the pictures to a resolution of K×K;

[0015] S15. According to the internal parameters, external parameters and positions of the monocular camera, translate, rotate and scale the complete point cloud P i , the incomplete point cloud P in and the monocular image I i to the same coordinate system and perform data annotation. Each incomplete point cloud P in and its corresponding complete point cloud P i as well as multiple monocular images I i constitute a data pair, forming a cross-modal large aircraft point cloud completion dataset.

[0016] Furthermore, the cross-modal aircraft point cloud completion model includes:

[0017] A point cloud feature extractor and an image encoder for extracting the point cloud feature F in from the incomplete point cloud P i and the monocular image I P , the rough point cloud D and the image feature F I , and adding the position encoding PE P to the point cloud feature F I and the position encoding PE P and PE I, and perform fusion to generate the output feature F O The cross-modal feature fusion module, and based on the output feature F O and the rough point cloud D to generate the full-resolution complete aircraft point cloud P out The FoldingNet module.

[0018] Furthermore, in step S3, the specific process includes the following steps:

[0019] S31. Normalize and preprocess the monocular image I i After preprocessing, the monocular image I i is input into the pre-trained image encoder to obtain the global image feature F I , and the expression is:

[0020] F I = CLIP Encoder(I i )

[0021] In the formula, CLIP Encoder(I i ) represents performing image encoding operation on the monocular image I i ;

[0022] S32. Input the defective point cloud P in into the point cloud feature extractor to generate the point cloud feature F P , and generate the rough point cloud D from the point cloud feature F P , and the expression is:

[0023] F P = Transformer(P in )

[0024] D = MLP(F P )

[0025] In the formula, Transformer(P in ) represents performing point cloud extraction operation on the defective point cloud P in ; MLP represents the multi-layer perceptron;

[0026] S33. Add the position encoding PE P to the point cloud feature F I and PE P and PE I to obtain the point cloud feature F' P and the image feature F' I ;

[0027] S34. Use the point cloud feature F' P as the query Q, and the image feature F' ICalculate the cross-attention between the key K and the value V to obtain the fused feature F CA ;

[0028] S35. Input the fused feature F CA into a 2-layer self-attention module to enhance the internal feature correlation, and obtain the output feature F that simultaneously contains cross-modal interaction information, enhanced self-correlation features, and spatial position encoding information O , and the expression of the self-attention module is:

[0029] F O = Attention(F CA , F CA , F CA )

[0030] S36. Based on the output feature F O and the rough point cloud D, generate the full-resolution complete point cloud P through the FoldingNet module out ;

[0031] S37. Calculate the chamfer distance between the ground-truth complete point cloud G and the rough point cloud D and the output complete point cloud P out to obtain the chamfer distance L CD (P out , G) and L CD (D, G);

[0032] S38. Determine the loss function J of the cross-modal aircraft point cloud completion model according to the chamfer distances L CD (P out , G) and L CD (D, G);

[0033] S39. Update the model parameters by gradient backpropagation, train the cross-modal aircraft point cloud completion model for multiple rounds until convergence, and save the model parameters of the best result

[0034] Furthermore, in step S34, the calculation formula for calculating the cross-attention to obtain the fused feature F CA is as follows:

[0035]

[0037] F CA = Attention(F ' P , F ' I , F ' I )

[0038] where d k represents the dimension of the query, key, and value

[0039] Further, in step S36, the chamfer distance L out between the true value complete point cloud G and the complete point cloud P CD (P out , G) is calculated by the formula:

[0040]

[0041] wherein, x, respectively represent any point cloud in the complete point cloud P out and the true value complete point cloud G;

[0042] The chamfer distance L CD (D, G) between the true value complete point cloud G and the rough point cloud D is calculated by the formula:

[0043]

[0044] wherein, x ` represents any point cloud in the rough point cloud D.

[0045] Further, in step S38, the expression of the loss function J is:

[0046] J = αL CD (D, G) + βL CD (P out , G)

[0047] wherein, α and β are both hyperparameters.

[0048] By means of the above technical solution, the present invention provides a cross-modal large aircraft point cloud completion method, which at least has the following beneficial effects:

[0049] 1. Compared with the existing completion methods based only on single-modal point clouds, the present invention can not only provide more geometric details and texture information by introducing monocular images, but also provide complementary information for the point cloud through images from different perspectives, so as to better optimize the point cloud completion effect.

[0050] 2. By constructing a cross-modal point cloud completion deep learning model and performing full supervision learning, the present invention fully learns the shape information in the data, and after convergence, can infer a complete aircraft point cloud with rich details and complete structure by inputting the incomplete point cloud and monocular images, thereby improving the performance of downstream tasks such as aircraft target detection and three-dimensional reconstruction of large aircraft.

[0051] 3. The present invention obtains monocular images through a camera, providing more shape details and texture information for point cloud completion. At the same time, it combines a three-dimensional laser scanner to obtain point clouds to provide geometric information on the overall shape. This method effectively reduces the uncertainty in inferring missing parts during the point cloud completion process and improves the detail quality of the generated complete point cloud to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0053] Figure 1 is a flowchart of the cross-modal large aircraft point cloud completion method in the present invention;

[0054] Figure 2 is a network structure diagram of the cross-modal point cloud completion deep learning model in the present invention;

[0055] Figure 3 is a network structure diagram of the cross-modal feature fusion module in the present invention;

[0056] Figure 4 is an operation schematic diagram of obtaining a complete point cloud using a three-dimensional laser scanner in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Thus, the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0058] Existing point cloud completion methods only use single-modal point cloud data, and two-dimensional images are easier to obtain than three-dimensional point clouds. In addition, since images have higher resolution and texture information than point clouds, using images from different perspectives can provide complementary information for the large aircraft point cloud to better optimize the point cloud completion task. To solve the technical defects of the single-modal point cloud completion method, please refer to Figures 1-4 , this embodiment proposes a cross-modal large aircraft point cloud completion method. By introducing monocular image assistance, it can not only provide more geometric details and texture information, but also provide complementary information for the point cloud through images from different perspectives, thereby better optimizing the point cloud completion effect. As Figure 1 shown, this method includes the following steps:

[0059] S1. Obtain the complete point cloud P of the large aircraft in a stationary state i , and capture monocular images I of the large aircraft from multiple perspectives with a fixed camera pose i, and based on the complete point cloud P i and the monocular image I i to establish a cross-modal large aircraft point cloud completion dataset. The specific process includes the following steps:

[0060] S11. Paste coded points at various positions on the large aircraft, and use a 3D laser scanner to perform an omnidirectional scan when the large aircraft is stationary to obtain the complete point cloud P containing the coded point coordinate information i , as Figure 4 shown. Then use it as the ground truth of the cross-modal large aircraft point cloud completion dataset, that is, the ground truth complete point cloud G. Specifically, use a 3D laser scanner to perform an omnidirectional scan on it when the large aircraft is stationary, and the goal is to obtain a complete point cloud P with uniform distribution and containing the complete structural information of the aircraft i , and then import the complete point cloud P i into point cloud processing software (such as cloud compare), and use Gaussian filtering to filter out worthless noise and outliers, and manual removal can also be performed. Subsequently, it can be used as the ground truth of the cross-modal large aircraft point cloud completion dataset.

[0061] S12. Select multiple viewpoints and fix a monocular camera at each viewpoint to take pictures of the large aircraft to obtain the monocular image I of the large aircraft at each viewpoint i . When selecting the position and viewpoint of the monocular camera, it is necessary to ensure that the photo contains a complete picture of the large aircraft, but self-occlusion is allowed. In this example, 24 viewpoints are selected to obtain 24 monocular images;

[0062] S13. Reconstruct the complete point cloud P i , then select 8 viewpoints (allowing different viewpoints from those for monocular image acquisition), and discard the point cloud of the occluded part at this viewpoint. Each incomplete point cloud P in is sampled to N points. Specifically, for those with more than N points, downsample to N points, and for those with less than N points, upsample to N points. Thus, multiple incomplete point clouds are obtained, and each incomplete point cloud has 2048 points.

[0063] S14. Downsample the ground truth complete point cloud G to M points, and downsample the image to a resolution of K×K; specifically, downsample the ground truth complete point cloud G to 16384 points, and downsample the image to a resolution of 224×224.

[0064] S15. Translate, rotate, and scale the complete point cloud P i , the incomplete point cloud P in and the monocular image I i to the same coordinate system according to the internal parameters, external parameters, and position of the monocular camera, and perform data annotation. Each incomplete point cloud P in and its corresponding complete point cloud Pi and multiple monocular images I i constitute a data pair, based on which a cross-modal large aircraft point cloud completion dataset is constructed.

[0065] S2. Construct a cross-modal aircraft point cloud completion model for automatically in completing the three-dimensional defective point cloud P of the aircraft into the complete aircraft point cloud P out , as shown in Figure 2 , where the cross-modal aircraft point cloud completion model includes a point cloud feature extractor and an image encoder for respectively extracting point cloud feature F in from the defective point cloud P i and the monocular image I P , rough point cloud D and image feature F I , and a cross-modal feature fusion module for respectively adding positional encodings PE P to the point cloud feature F I and the image feature F P and PE I , and fusing them to generate the output feature F O , and a FoldingNet module for generating the full-resolution complete aircraft point cloud P O from the output feature F out and the rough point cloud D.

[0066] In this embodiment, by constructing a cross-modal aircraft point cloud completion model and performing fully supervised learning, the shape information in the data is fully learned, so that after convergence, the cross-modal aircraft point cloud completion model can infer a complete aircraft point cloud with rich details and structure by inputting the defective point cloud and the monocular image.

[0067] S3. Train the cross-modal aircraft point cloud completion model using the cross-modal large aircraft point cloud completion dataset; in step S3, the specific process includes the following steps:

[0068] S31. Perform standardized preprocessing on the monocular image I i , including size adjustment, normalization, and central cropping, to match the input requirements of the CLIP image encoder. The processed monocular image I i is input into the pre-trained image encoder to obtain the global image feature F I , and the expression is:

[0069] F I = CLIP Encoder(I i )

[0070] In the formula, CLIP Encoder(I i ) represents the image encoding operation on the monocular image I i .

[0071] S32. Input the defective point cloud P in into the point cloud feature extractor to generate the point cloud feature F P , and generate the rough point cloud D from the point cloud feature F P . The expression is as follows:

[0072] F P = Transformer(P in )

[0073] D = MLP(F P )

[0074] where Transformer(P in ) represents performing a point cloud extraction operation on the defective point cloud P in ; MLP represents a multi-layer perceptron;

[0075] S33. Add the position encoding PE P and PE I to the point cloud feature F P and the image feature F I respectively to obtain the point cloud feature F' P and the image feature F' I ; specifically, for the point cloud feature F P , the position encoding PE P can be constructed based on the three-dimensional coordinate positions (x, y, z) of the point cloud feature F P through sine and cosine functions, that is:

[0076]

[0077] where pos is the position index; i is the dimension index; d is the embedding dimension of the Transformer.

[0078] For the image feature F I , it is extended to the feature dimension of the Transformer through a linear transformation, and then the position encoding PE I is generated through the above sine-cosine position encoding method.

[0079] The final point cloud feature F P and the image feature F I are combined with the position encoding PE P and PE I in the following way to obtain the point cloud feature F' P and the image feature F' I , that is:

[0080] F′ P = F P + PE P

[0081] F′ I = F I + PE I

[0082] S34. Use the point cloud feature F' P as the query Q, and the image feature F' I as the key K and value V to calculate cross-attention to obtain the fused feature F CA . It effectively captures the supplementary information of the image feature F' I for the point cloud feature F' P , and at the same time combines the spatial position information. The calculation formula is as follows:

[0083]

[0085] F CA = Attention(F ' P , F ' I , F ' I )

[0086] In the formula, d k represents the dimensions of the query, key, and value.

[0087] S35. Input the fused feature F CA into a 2-layer self-attention module to enhance the internal feature correlation, and obtain the output feature F O that simultaneously contains cross-modal interaction information, enhanced self-correlation features, and spatial position encoding information. The expression of the self-attention module is:

[0088] F O = Attention(F CA , F CA , F CA )

[0089] As Figure 3 shown, the output feature of the final cross-modal feature fusion module simultaneously contains cross-modal interaction information, enhanced self-correlation features, and spatial position encoding information, providing high-quality feature representations for subsequent tasks.

[0090] S36. Based on the output feature F O and the rough point cloud D, generate the full-resolution complete point cloud P out through the FoldingNet module. The rough point cloud D can effectively capture the global structure of the point cloud, providing a reasonable initial guess for generating high-resolution point clouds and reducing the optimization difficulty.

[0091] S37. Calculate the chamfer distance L between the true complete point cloud G, the rough point cloud D, and the output complete point cloud P out CD (P out , G) and L CD (D, G). If the chamfer distance L out between the true complete point cloud G and the complete point cloud P CD (P out , G) is less than the chamfer distance L

[0092]

[0093] In the formula, x, respectively represent any point cloud in the complete point cloud P out and the true complete point cloud G;

[0094] The formula for calculating the chamfer distance L CD (D, G) between the true complete point cloud G and the rough point cloud D is:

[0095]

[0096] In the formula, x` represents any point cloud in the rough point cloud D;

[0097] S38. Determine the loss function J of the cross-modal aircraft point cloud completion model according to the chamfer distances L CD (P out , G) and L CD (D, G). The expression is:

[0098] J = αL CD (D, G) + βL CD (P out , G)

[0099] In the formula, α and β are both hyperparameters;

[0100] S39. Update the model parameters by gradient backpropagation, train the cross-modal aircraft point cloud completion model for multiple rounds until convergence, and save the model parameters with the best results.

[0101] S4. Use the incomplete point cloud P in to be completed and the monocular image I i of a single aircraft as the input of the cross-modal aircraft point cloud completion model, and output a complete aircraft point cloud with fine geometric details.

[0102] ​In this embodiment, a monocular camera and a 3D laser scanner are combined to obtain the image data and point cloud data of a large aircraft respectively; the point cloud data is preprocessed, including operations such as denoising, filtering, and sparsification. At the same time, it is aligned with the image data through camera parameters to construct a cross-modal large aircraft point cloud completion dataset; using the dataset constructed above, a cross-modal aircraft point cloud completion model is trained, and the point cloud features and image features are effectively fused through a cross-modal feature fusion module, so as to generate a large aircraft point cloud with a complete structure and rich details; after the model training is completed, the incomplete 3D aircraft point cloud and a single aircraft image are input, and a complete aircraft point cloud with fine geometric details can be generated, thereby improving the performance of downstream tasks such as aircraft target detection and 3D reconstruction of large aircraft.

[0103] Those of ordinary skill in the art can understand that all or part of the steps in implementing the method of the above embodiment can be completed by instructing relevant hardware through a program. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0104] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A cross-modal large aircraft point cloud completion method, characterized in that: The method comprises the following steps: S1. Obtain the complete point cloud P of the large aircraft in a stationary state i , and monocular images of large aircraft taken from multiple perspectives with a fixed camera position I i , and based on the complete point cloud P i and monocular image I i Establish a cross-modal large aircraft point cloud completion dataset; S2, construct a defect cloud P for the aircraft to be three-dimensional in Automatically complete the point cloud of the aircraft out Cross-modal aircraft point cloud completion model; S3, using the cross-modal large aircraft point cloud completion dataset to train the cross-modal aircraft point cloud completion model; S4, the remaining defects to be completed are clouded into P in and a single monocular image of an airplane I i As input to the cross-modal aircraft point cloud completion model, it outputs a complete aircraft point cloud with fine geometric details.

2. The cross-modal large aircraft point cloud completion method according to claim 1, characterized in that: In step S1, the specific process includes the following steps: S11, paste the coding points at various positions on the large aircraft, and perform a full-scale scan when the large aircraft is stationary to obtain a complete point cloud P containing the coding point coordinate information i , and use it as the true value of the cross-modal large aircraft point cloud completion dataset, that is, the true value complete point cloud G; S12, select multiple viewing angles and fix the monocular camera to take pictures of the large aircraft respectively, and obtain the monocular image I of the large aircraft at each viewing angle i ; S13. For the complete point cloud P i Reconstruct, then select multiple different perspectives, and discard the obscured point cloud under the perspective. Each residual point cloud P in Sample to N points; S14, downsample the true value complete point cloud G to M points, and downsample the image to a resolution of K×K; S15, complete point cloud P i 、Residual cloud P in and monocular image I i According to the internal parameters, external parameters and position of the monocular camera, the camera is translated, rotated and scaled to the same coordinate system and the data is annotated. Each residual defect cloud P in and its corresponding complete point cloud P i And multiple monocular images I i A data pair is formed to constitute a cross-modal large aircraft point cloud completion dataset.

3. The cross-modal large aircraft point cloud completion method according to claim 1, characterized in that: The cross-modal aircraft point cloud completion model includes: Used to separate residual defect clouds P in and monocular image I i Extract point cloud features F P , rough point cloud D and image features F I Point cloud feature extractor and image encoder, and point cloud feature F P and image features F I Add position code PE respectively P and PE I , and fused to generate output feature F O The cross-modal feature fusion module, and the output feature F O Generate a full-resolution complete aircraft point cloud P from the rough point cloud D out FoldingNet module.

4. The cross-modal large aircraft point cloud completion method according to claim 1, characterized in that: In step S3, the specific process includes the following steps: S31, monocular image I i After standardization preprocessing, the processed monocular image I i Input the pre-trained image encoder to obtain the global image feature F I , the expression is: F I =CLIP Encoder(I i ) In the formula, CLIP Encoder (I i ) represents the monocular image I i Perform image encoding operations; S32, the residual cloud P in Input to the point cloud feature extractor to generate point cloud features F P , and the point cloud feature F P Generate a rough point cloud D, the expression is: F P =Transformer(P in ) D=MLP(F P ) In the formula, Transformer (P in ) represents the residual defect cloud P in Perform point cloud extraction operations; MLP stands for multi-layer perceptron; S33, is the point cloud feature F P and image features F I Add position code PE respectively P and PE I Get the point cloud feature F' P and image features F' I ; S34, point cloud feature F' P As query Q, image features F' I As key K and value V, we calculate the cross attention to obtain the fused feature F CA ; S35, fusion feature F CA Input a 2-layer self-attention module to enhance the internal feature correlation and obtain the output feature F that contains cross-modal interaction information, enhanced autocorrelation features, and spatial position encoding information O , the expression of the self-attention module is: F O =Attention(F CA ,F CA ,F CA ) S36, based on output feature F O The rough point cloud D is used to generate the full-resolution complete point cloud P through the FoldingNet module. out ; S37, using chamfer distance to calculate the true value of the complete point cloud G and the rough point cloud D and the output complete point cloud P out Chamfer distance L CD (P out ,G) and L CD (D,G); S38, according to the chamfer distance L CD (P out ,G) and L CD (D, G) Determine the loss function J of the cross-modal aircraft point cloud completion model; S39. Update the model parameters by gradient back propagation. Train the cross-modal aircraft point cloud completion model for multiple rounds until convergence, and save the model parameters of the best result.

5. The cross-modal large aircraft point cloud completion method according to claim 4, characterized in that: In step S34, the cross attention is calculated to obtain the fusion feature F CA The calculation formula is as follows: F CA =Attention(F ' P ,F ' I ,F ' I ) Where, d k Dimensions representing queries, keys, and values.

6. The cross-modal large aircraft point cloud completion method according to claim 4, characterized in that: In step S36, the ground truth complete point cloud G and the complete point cloud P out Chamfer distance L CD (P out ,G) is calculated as: In the formula, x, Represent the complete point cloud P out And any point cloud in the true complete point cloud G; The chamfer distance L between the true complete point cloud G and the rough point cloud D CD The calculation formula for (D,G) is: In the formula, x ` Represents any point cloud in the rough point cloud D.

7. The cross-modal large aircraft point cloud completion method according to claim 1, characterized in that: In step S38, the loss function J is expressed as: J=αL CD (D,G)+βL CD (P out ,G) In the formula, α and β are hyperparameters.

Citation Information

Patent Citations

  • Monocular image-assisted large aircraft point cloud completion method

    CN115496881A

  • Point cloud up-sampling method and system fusing self-attention and multipath path diagram convolution

    CN116188882A

  • Multi-sensor feature fusion method based on multi-scale local cross attention

    CN119066613A