Airplane surface fastener instance segmentation method based on point cloud prompt and image fusion
By collecting and preprocessing images, point cloud data and mapping and converting them into prompt information, combined with the segmentation network of a large model, high-precision segmentation and positioning of aircraft surface fasteners is achieved, solving the problems of low detection accuracy and poor efficiency in traditional methods, and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510430325.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
When detecting aircraft surface fasteners, traditional methods have problems with low detection accuracy and poor efficiency. Especially under complex background and environmental noise interference, it is difficult to achieve high-precision segmentation and robustness.
The aircraft surface image and point cloud data are collected, preprocessed and mapped and converted into prompt information. Combined with a segmentation network based on a large model, the precise segmentation and positioning of fasteners is achieved through multimodal data fusion.
It significantly improves the example segmentation accuracy and positioning robustness of aircraft surface fasteners, improves detection efficiency and automation level, and meets the high-precision evaluation requirements for local micro geometric features.
Smart Images

Figure CN120355918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aerospace detection, and particularly to an instance segmentation method for aircraft surface fasteners based on point cloud cues and image fusion. Background Art
[0002] Aircraft surface fasteners (such as rivets, bolts, etc.) are key structural components of aircraft, and their installation quality directly affects the safety and performance of the aircraft. In the actual production and maintenance processes, how to quickly and accurately detect the status and distribution of aircraft surface fasteners has become an important technical challenge. Traditional detection methods mainly rely on manual visual inspection or single-modal image processing methods. These methods often cannot guarantee segmentation accuracy and robustness in complex backgrounds, low contrast, and under the interference of environmental noise. In addition, although some deep learning-based image instance segmentation technologies have emerged in recent years, these methods usually only rely on two-dimensional image information and lack full utilization of point cloud data that reflects the true three-dimensional geometric features of objects. On the other hand, point cloud data has accurate three-dimensional structural information. By reasonably extracting and utilizing this geometric prior information, the positioning and segmentation effects of fasteners can be significantly improved. However, there is currently a lack of an overall technical solution that can effectively fuse the geometric cues extracted from the point cloud with image information to achieve high-precision instance segmentation of aircraft surface fasteners. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present invention provides an instance segmentation method for aircraft surface fasteners based on point cloud cues and image fusion. This method aims to solve the problems of the small difference between aircraft rivets and the skin surface and the subtle local geometric features, and overcome the defects of low accuracy and poor efficiency of traditional manual detection. This method collects aircraft surface images and point cloud data and establishes a mapping, preprocesses the image and the point cloud respectively, and converts the point cloud data into cue information; inputs the image and the point cloud cue information into a segmentation network based on a large model to output a fastener segmentation mask. Through multi-modal data fusion and deep learning, precise segmentation and positioning of aircraft surface fasteners are achieved, significantly improving the detection efficiency and automation level, and providing key data support for the evaluation of the installation quality and local geometric features of aircraft skins and rivets.
[0004] To solve the above technical problems, the present invention provides the following technical solution: An instance segmentation method for aircraft surface fasteners based on point cloud cues and image fusion, comprising the following steps:
[0005] S1. Collect aircraft surface images and point cloud data, and perform mutual mapping on the image and the point cloud data;
[0006] S2. Perform preprocessing operations of grayscale conversion, adaptive denoising, and multi-scale feature enhancement on the image data to obtain a preprocessed image I with clear edge contours of the fasteners;
[0007] S3. Perform preprocessing operations of denoising, sampling, normal estimation, and plane segmentation on the point cloud data to obtain the preprocessed point cloud
[0008] S4. For the preprocessed point cloud Perform key point detection based on intrinsic shape features to obtain a set of key points K, cluster the detected set of key points, extract two key information of centroid and circle fitting according to the key points within each cluster and perform transformation, and obtain single-point prompts, box prompts, and mask prompts as point cloud prompt information P;
[0009] S5. Input the preprocessed image I and the point cloud prompt information P into a segmentation network based on a large model to output a segmentation mask of the fasteners.
[0010] Furthermore, in step S1, mutual mapping is performed on the image and point cloud data, and the specific process includes the following steps:
[0011] S11. Use a calibrated monocular structured light camera to photograph the aircraft surface to obtain image data and point cloud data;
[0012] S12. Since both the image and point cloud data are from the same camera system, utilize the pre-calibrated internal and external parameter relationships to establish a mapping between the point cloud and the image. Each point in the point cloud data is represented as p i , and establish a connection with the corresponding pixel point F(x, y) in the image through the mapping relationship:
[0013] i = x + y × w
[0014] i is the index of the point in the point cloud, x is the column number where the pixel point is located, y is the row number where the pixel point is located, and w is the width of the overall pixels of the picture.
[0015] Furthermore, the specific process of step S2 includes the following steps:
[0016] S21. For the acquired image, first perform a grayscale conversion operation on the image to convert the RGB image into a grayscale image, and at the same time maximize the retention of edge and local contrast information in the original color channels. By fusing the RGB channel information through the Laplacian operator and global contrast weights, obtain the grayscale image F1;
[0017] S22. To suppress Gaussian noise and salt-and-pepper noise while retaining the edge structure of the fasteners, use non-local means denoising and perform weighted averaging using the self-similarity of the image to obtain the denoised image F2;
[0018] S23. Apply the contrast enhancement method to the denoised image F2, using limited contrast adaptive histogram equalization, block equalization, and limiting the histogram distribution to obtain the contrast-enhanced image F3, which improves the contrast between the fastener and the background while avoiding local over-enhancement.
[0019] S24. Use the edge sharpening method, which includes gradient calculation, non-maximum suppression, double-threshold detection, and edge connection in sequence, to finally obtain the preprocessed image I with clear fastener edge contours, providing a high-precision input for subsequent segmentation.
[0020] Furthermore, the specific process of step S3 includes the following steps:
[0021] S31. For the collected point cloud data where N is the number of points, first use the statistical filtering algorithm to dynamically remove noise based on the mean and standard deviation of the neighborhood point distribution.
[0022] S32. Then, use uniform downsampling based on a voxel grid to divide the point cloud into regular voxels and retain the centroids to obtain the filtered and denoised point cloud P1.
[0023] S33. Further construct neighborhoods for the preliminarily processed point cloud P1. For each point p i , search for the point set within its neighborhood of radius r Subsequently, calculate the covariance matrix and perform eigenvalue decomposition to obtain the normal vector of the point cloud and the eigenvalues λ1, λ2, and λ3.
[0024] S34. To pre-separate the fastener body from the background skin, extract the region of interest of the fastener. Based on the plane model fitting of the random sample consensus (RANSAC), extract the point set belonging to the plane as the background, and the remaining points as the fastener region to obtain the preprocessed point cloud
[0025] Furthermore, in step S4, perform key point detection based on the intrinsic shape features on the preprocessed point cloud , and the specific process includes the following steps:
[0026] S41. First, perform significance measurement based on the eigenvalues λ1, λ2, and λ3 obtained from the normal vector calculation in the previous step, and use the eigenvalue ratio to evaluate the significance of points:
[0027]
[0028] where Saliency(p i ) is the significance of point p i , and ∈ is a very small constant to prevent the denominator from being zero.
[0029] S42. Set the threshold T sal to 0.8, and only keep the points where Saliency(p i ) > T sal . Then, perform non-maximum suppression to eliminate redundant key points and retain the locally most significant points;
[0030] S43. For each point p i , compare the saliency within its neighborhood. If the saliency of p i is the maximum within the neighborhood, mark it as a key point, and finally generate a set of key points
[0031] Furthermore, in step S4, cluster the detected set of key points, extract two key pieces of information, the centroid and circle fitting, from the key points within each cluster and transform them. The resulting single-point prompt, box prompt, and mask prompt are used as point cloud prompt information P. The specific process includes the following steps:
[0032] S44. Cluster the detected set of key points K = {k1, k2,..., k m} using the K-means clustering algorithm, and group the key points belonging to the same fastener into one cluster K i ;
[0033] S45. According to the mapping relationship between the point cloud and the image, map the key points to the pixel points in the image to obtain the key pixel points (x i , y i );
[0034] S46. Extract the centroid information as the position prompt, and calculate the centroid of the set of key points in the cluster as the center position of the fastener: Use the centroid of each key point cluster as the single-point prompt;
[0035]
[0036] x i and y i are the coordinate positions of the key pixel points respectively, and n is the number of key pixel points in a cluster;
[0037] S47. Then, use the least squares method to perform circle fitting on the key pixel points to obtain the center (a, b) and radius r of the fitted circle. The objective function used is:
[0038]
[0039] Solve for the minimum of F(a, b, r) to obtain the parameters (a, b, r);
[0040] S48. Calculate the radius r obtained by fitting a circle, and calculate the circumscribed rectangle B=(a - r, b - r, a + r, b + r), which is used as a box hint.
[0041] Generate a binary mask M(x, y) using circular information, and the generation criteria are as follows:
[0042]
[0043] Perform convex hull processing on the key pixel points, convert the polygon area into a binary mask, and generate the mask hint M(x, y) as a mask hint; the finally obtained single-point hint, box hint, and mask hint are used as hint information P.
[0044] Furthermore, in step S5, the large model-based segmentation network includes: an image encoder, which uses a pre-trained ViT model and fine-tunes the query-key-value parameters of the multi-head attention layer through the LoRA technique to adapt to specific segmentation tasks; a self-generated prompt generator, which uses a neural network to generate a prompt embedding that matches the network prompt input format from the prompt information generated by the point cloud; a mask decoder, which based on the self-attention and cross-attention mechanisms, fuses the image embedding output by the image encoder with the prompt embedding generated by the self-generated prompt generator and outputs the final instance segmentation mask.
[0045] Furthermore, in step S5, the specific process includes the following steps:
[0046] S51. Image encoding process: The preprocessed image I is used as input and fed into the image encoder. After passing through 12 Transformer layers, and in the multi-head attention module of each layer, the LoRA technique is used to perform low-rank decomposition fine-tuning on the query, key, and value matrices. Finally, after being processed by the image encoder, the image feature embedding E I , which has global semantic information and local detail representation;
[0047] S52. Input the single-point hint, box hint, and mask hint generated by the point cloud into the self-generated prompt generator and encode them to obtain prompt embedding vectors E point , E box and E mask of fixed dimensions respectively. Finally, concatenate the above three embedding vectors Concat to obtain the final prompt embedding vector E P :
[0048]
[0049] S53. The prompt information embedding E P performs internal information interaction through the self-attention layer, and the updated prompt embedding E' P can be expressed as:
[0050] It is P = SelfAttn(E P ) + E P
[0051] where SelfAttn represents the multi-head self-attention operation, which allows for feature reallocation within the prompt information and strengthens the expression of local geometric priors;
[0052] S54. Embed the image features E I and the updated prompt embedding vector E' P into the masked decoder for masked decoding and prediction, and perform fusion through multi-layer self-attention and cross-attention; the updated prompt embedding E' P is used as the query, and the image feature embedding E I is used as the key and value to input the cross-attention layer to achieve the interactive fusion of the prompt information and the image semantics;
[0053] S55. Finally, after 12 layers of Transformer decoding, the network generates multiple candidate segmentation masks, and selects the mask with the highest score as the final output to complete the instance segmentation of the fastener.
[0054] Furthermore, in step S52, the single-point prompt, box prompt, and mask prompt generated from the point cloud are input into the self-generated prompt encoder to obtain the prompt embedding vectors E point , E box , and E mask with fixed dimensions respectively. The specific process includes, for the point prompt, mapping it to a 128-dimensional embedding vector E point through a point prompt feed-forward neural network f point . f point (·) uses a two-layer fully connected network. The first layer: the fully connected layer (FC) maps the input dimension 2 to the hidden layer of 64 dimensions, and then uses ReLU activation; the second layer: the fully connected layer maps the hidden layer of 64 dimensions to the output embedding of 128 dimensions. For the box prompt, it is mapped to a 128-dimensional embedding vector E box through another box prompt feed-forward neural network f box . f box (·) is different from f point (·) in that the input end of the box prompt feed-forward neural network is modified to 4 dimensions, and the rest of the internal structure is the same. For the mask prompt, a mask prompt lightweight convolutional neural network is adopted. The input is the binary mask M(x, y), and the internal network structure is 3 convolutional layers of 3×3. After each convolutional layer, there is a batch normalization layer and a ReLU activation function. At the end of the network, a global average pooling and a fully connected layer are used to obtain a 128-dimensional embedding space E mask .
[0055] With the above technical solution, the present invention provides a method for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion, which has at least the following beneficial effects:
[0056] Through a multi-modal data fusion strategy, the present invention combines image preprocessing, point cloud prompting generation and encoding module, and a segmentation network based on a large model, realizing the deep fusion of image semantics and key geometric features of fasteners in the point cloud, thereby significantly improving the instance segmentation accuracy and positioning robustness of aircraft surface fasteners. At the same time, the present invention designs an optimized prompting information encoding module, which can accurately convert point prompts, box prompts and mask prompts into unified embedding vectors, meeting the requirements for high-precision evaluation of local micro geometric features in the quality inspection of aircraft skins and fastener installations. From data collection, preprocessing, prompting information generation to the overall construction of the segmentation network, the present invention provides a systematic technical process, overcoming the deficiencies of traditional methods in detection efficiency and accuracy, providing key technical support for the automatic detection of aircraft surface fasteners, and significantly improving the detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0058] Figure 1 It is a schematic diagram of the specific steps of the method and system for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion proposed according to the present invention;
[0059] Figure 2 It is a schematic diagram of the overall process of the method proposed by the present invention;
[0060] Figure 3 It is a schematic diagram of the process of converting point cloud data into prompting information in the present invention;
[0061] Figure 4 It is a schematic diagram of the segmentation network architecture based on a large model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Thereby, a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects can be obtained and implemented accordingly.
[0063] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0064] Please refer to Figures 1-4 , which shows a specific implementation manner of this embodiment. In this embodiment, through multi-modal data collaborative processing and deep learning technology, high-precision and automated fastener detection is achieved. The specific steps are as follows: First, use a calibrated monocular structured light camera to synchronously collect the surface image and point cloud data of the aircraft, and establish a mapping relationship between the two; Second, perform preprocessing such as grayscale conversion, adaptive denoising, and multi-scale feature enhancement on the image to improve the image quality; perform denoising, sampling, normal estimation, and plane segmentation on the point cloud data to extract effective geometric information; Then, convert the point cloud data into prompt information, including single-point prompts, box prompts, and mask prompts; Finally, input the preprocessed image and point cloud prompts into a segmentation network based on a large model, and through the collaborative work of an image encoder, a self-generated prompt generator, and a mask decoder, output an accurate segmentation mask of the fastener. The present invention solves the problems of small differences between rivets and skins and low detection accuracy in traditional methods through the deep fusion of point clouds and images, and significantly improves the detection efficiency and accuracy.
[0065] Please refer to Figure 1 and Figure 2 , this embodiment proposes an aircraft surface fastener instance segmentation method based on point cloud prompts and image fusion, and this method includes the following steps:
[0066] S1. Collect the surface image and point cloud data of the aircraft, and perform mutual mapping on the image and point cloud data;
[0067] As a preferred implementation manner of step S1, in step S1, performing mutual mapping on the image and point cloud data, the specific process includes the following steps:
[0068] S11. Use a calibrated monocular structured light camera to photograph the surface of the aircraft to obtain image data and point cloud data;
[0069] S12. Since both the image and point cloud data come from the same camera system, use the pre-calibrated internal and external parameter relationships to establish a mapping between the point cloud and the image, and each point in the point cloud data is represented as p i , and establish a connection with the corresponding pixel point F(x, y) in the image through the mapping relationship:
[0070] i = x + y × w
[0071] i is the index of the point in the point cloud, x is the column number where the pixel point is located, y is the row number where the pixel point is located, and w is the width of the overall pixels of the picture.
[0072] S2. Perform preprocessing operations of grayscale conversion, adaptive denoising, and multi-scale feature enhancement on the image data to obtain a preprocessed image I with clear edge contours of the fastener.
[0073] As a preferred implementation manner of step S2, the specific process of step S2 includes the following steps:
[0074] S21. For the acquired image, first perform a grayscale conversion operation on the image to convert the RGB image into a grayscale image, and at the same time maximize the retention of edge and local contrast information in the original color channels. By fusing the RGB channel information through the Laplacian operator and the global contrast weight, a grayscale image F1 is obtained.
[0075] S22. To suppress Gaussian noise and salt-and-pepper noise while retaining the edge structure of the fastener, non-local means denoising is adopted, and weighted averaging is performed using the self-similarity of the image to obtain a denoised image F2.
[0076] S23. Use a contrast enhancement method for the denoised image F2, limit contrast adaptive histogram equalization, perform block equalization and limit the histogram distribution to obtain a contrast-enhanced image F3, realize the improvement of the contrast between the fastener and the background, and at the same time avoid local over-enhancement.
[0077] S24. Use an edge sharpening method to perform gradient calculation, non-maximum suppression, double-threshold detection, and edge connection in sequence, and finally obtain a preprocessed image I with clear edge contours of the fastener, providing a high-precision input for subsequent segmentation.
[0078] S3. Perform preprocessing operations of denoising, sampling, normal estimation, and plane segmentation on the point cloud data to obtain a preprocessed point cloud
[0079] As a preferred implementation manner of step S3, the specific process of step S3 includes the following steps:
[0080] S31. For the acquired point cloud data N is the number of points. First, use a statistical filtering algorithm to dynamically remove noise based on the mean and standard deviation of the neighborhood point distribution.
[0081] S32. Then, use uniform downsampling based on a voxel grid to divide the point cloud into regular voxels and retain the centroid to obtain a filtered and denoised point cloud P1.
[0082] S33. The preliminarily processed point cloud P1 is further subjected to neighborhood construction. For each point p i , the point set within its neighborhood with a radius of r is searched Subsequently, the covariance matrix is calculated and eigenvalue decomposition is performed to obtain the normal vector of the point cloud, and the eigenvalues λ1, λ2, and λ3 are obtained;
[0083] S34. In order to pre-separate the fastener body from the background skin, the region of interest of the fastener is extracted. Based on the plane model fitting of the random sample consensus (RANSAC), the point set belonging to the plane is extracted as the background, and the remaining points are used as the fastener region to obtain the preprocessed point cloud
[0084] S4. For the preprocessed point cloud Intrinsic shape feature-based key point detection is performed to obtain the key point set K. The detected key point set is clustered, and two key information, namely the centroid and circle fitting, are extracted from the key points within each cluster and transformed. The obtained single-point prompt, box prompt, and mask prompt are used as the point cloud prompt information P, as Figure 3 shown;
[0085] As a preferred implementation manner of step S4, in step S4, for the preprocessed point cloud Intrinsic shape feature-based key point detection is performed. The specific process includes the following steps:
[0086] S41. First, significance measurement is performed based on the eigenvalues λ1, λ2, and λ3 obtained from the normal vector calculation in the previous step. The significance of the point is evaluated using the eigenvalue ratio:
[0087]
[0088] where Saliency(p i ) is the significance of point p i , ∈ is a minimum constant to prevent the denominator from being zero;
[0089] S42. Set the threshold T sal to 0.8, and only retain the points where Saliency(p i ) > T sal . Then, non-maximum suppression is performed to eliminate redundant key points and retain the locally most significant points;
[0090] S43. For each point p i , compare the significance within its neighborhood. If the significance of p i is the maximum within the neighborhood, it is marked as a key point, and finally the key point set is generated
[0091] More specifically, in step S4, cluster the detected set of key points, extract two key pieces of information, namely the centroid and circle fitting, based on the key points within each cluster and transform them. The obtained single-point hint, box hint, and mask hint are used as the point cloud hint information P. The specific process includes the following steps:
[0092] S44. Cluster the detected set of key points K = {k1, k2,..., k m} using the K-means clustering algorithm, and group the key points belonging to the same fastener into one cluster K i ;
[0093] S45. According to the mapping relationship between the point cloud and the image, map the key points to the pixel points in the image to obtain the key pixel points (x i , y i );
[0094] S46. Extract the centroid information as the position hint, and calculate the centroid of the set of key points in the cluster as the center position of the fastener: Use the centroid of each key point cluster as the single-point hint;
[0095]
[0096] x i and y i are the coordinate positions of the key pixel points respectively, and n is the number of key pixel points in a cluster;
[0097] S47. Then, use the least squares method to perform circle fitting on the key pixel points to obtain the center (a, b) and radius r of the fitted circle. The objective function used is:
[0098]
[0099] Solve for the minimum of F(a, b, r) to obtain the parameters (a, b, r);
[0100] S48. Calculate the circumscribed rectangle B = (a - r, b - r, a + r, b + r) through the radius r of the fitted circle. This rectangle is used as the box hint;
[0101] Generate a binary mask M(x, y) using the circular information. The generation criterion is as follows:
[0102]
[0103] Perform a convex hull operation on the key pixel points, convert the polygon area into a binary mask, and generate the mask hint M(x, y) as the mask hint; The finally obtained single-point hint, box hint, and mask hint are used as the hint information P.
[0104] In this embodiment, the present invention designs an optimized prompt information encoding module, which can accurately convert point prompts, box prompts, and mask prompts into unified embedding vectors, meeting the requirements for high-precision evaluation of local micro geometric features in aircraft skin and fastener installation quality inspection.
[0105] S5. Input the preprocessed image I and the point cloud prompt information P into the large model-based segmentation network, and output the segmentation mask of the fastener.
[0106] As a preferred implementation manner of step S5, in step S5, the large model-based segmentation network includes: an image encoder, which uses a pre-trained ViT model and fine-tunes the query-key-value parameters of the multi-head attention layer through the LoRA technique to adapt to specific segmentation tasks; a self-generated prompt generator, which uses a neural network to generate prompt embeddings that match the network prompt input format from the prompt information generated by the point cloud; a mask decoder, which fuses the image embeddings output by the image encoder and the prompt embeddings generated by the self-generated prompt generator based on the self-attention and cross-attention mechanisms, and outputs the final instance segmentation mask, as Figure 4 shown.
[0107] Specifically, the specific process of step S5 includes the following steps:
[0108] S51. Image encoding process. The preprocessed image I is input into the image encoder, passes through 12 Transformer layers, and the query, key, and value matrices are fine-tuned by low-rank decomposition in the multi-head attention module of each layer. Finally, after being processed by the image encoder, the image feature embedding E I is obtained, and this embedding has global semantic information and local detail representation;
[0109] S52. Input the single-point prompt, box prompt, and mask prompt generated by the point cloud into the self-generated prompt generator to be encoded into prompt embedding vectors E point 、E box and E mask with fixed dimensions respectively. Finally, the above three embedding vectors are concatenated by Concat to obtain the final prompt embedding vector E P :
[0110]
[0111] Specifically, input the single-point prompt, box prompt, and mask prompt generated by the point cloud into the self-generated prompt generator to be encoded into prompt embedding vectors E point 、E box and E mask with fixed dimensions respectively. The specific process includes that for the point prompt, through a point prompt feed-forward neural network f point(·) Map it to an embedding vector E of 128 dimensions point , f point (·) Adopt a two-layer fully connected network. The first layer: the fully connected layer (FC) maps the input dimension 2 to 64 dimensions in the hidden layer, and then uses ReLU activation; the second layer: the fully connected layer maps the 64 dimensions in the hidden layer to the output embedding of 128 dimensions; for the box prompt, through another box prompt feed-forward neural network f box (·) Map it to an embedding vector E of 128 dimensions box , f box (·) Different from f point (·) is that the box prompt feed-forward neural network modifies the input end to 4 dimensions, and the rest of the internal structure is the same; for the mask prompt, adopt a mask prompt lightweight convolutional neural network, the input is a binary mask M(x, y), and the internal network structure is 3 convolutional layers of 3×3. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected. The network end obtains a 128-dimensional embedding space E through global average pooling and a fully connected layer mask .
[0112] S53. Prompt information embedding E P It conducts internal information interaction through a self-attention layer, and the updated prompt embedding E' P can be expressed as:
[0113] E' P = SelfAttn(E P ) + E P
[0114] where SelfAttn represents the multi-head self-attention operation, which allows feature reallocation within the prompt information and strengthens the expression of local geometric priors;
[0115] S54. Embed the image features E I and the updated prompt embedding vector E' P into the mask decoder for mask decoding and prediction, and perform fusion through multi-layer self-attention and cross-attention; the updated prompt embedding E' P is used as the query, and the image feature embedding E I is used as the key and value to input the cross-attention layer to achieve the interaction and fusion of the prompt information and the image semantics;
[0116] S55. Finally, after 12 layers of Transformer decoding, the network generates multiple candidate segmentation masks, and selects the mask with the highest score as the final output to complete the instance segmentation of the fastener, as Figure 4 shown.
[0117] The beneficial effects of the present invention are as follows: Through the multi-modal data fusion strategy, combined with image preprocessing, point cloud prompt generation and encoding module, and a segmentation network based on a large model, the present invention realizes the deep fusion of image semantics and key geometric features of fasteners in the point cloud, thus significantly improving the instance segmentation accuracy and positioning robustness of aircraft surface fasteners. At the same time, the present invention designs an optimized prompt information encoding module, which can accurately convert point prompts, box prompts, and mask prompts into unified embedding vectors, meeting the requirements for high-precision evaluation of local microscopic geometric features in the quality inspection of aircraft skins and fastener installations. From data collection, preprocessing, prompt information generation to the overall construction of the segmentation network, the present invention provides a systematic technical process, overcomes the deficiencies of traditional methods in detection efficiency and accuracy, provides key technical support for the automated detection of aircraft surface fasteners, and significantly improves the detection efficiency and accuracy.
[0118] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0119] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices.
[0120] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An instance segmentation method for aircraft surface fasteners based on point cloud prompting and image fusion, characterized in that, It includes the following steps: S1. Collect the surface images and point cloud data of the aircraft, and perform mutual mapping on the images and point cloud data; S2. Perform preprocessing operations of grayscaling, adaptive denoising, and multi-scale feature enhancement on the image data to obtain the preprocessed image I with clear edge contours of fasteners; S3. Perform preprocessing operations of denoising, sampling, normal estimation, and plane segmentation on the point cloud data to obtain the preprocessed point cloud S4. For the preprocessed point cloud perform key-point detection based on intrinsic shape features to obtain a key-point set K, cluster the detected key-point set, extract two key pieces of information, namely the centroid and circle fitting, from the key points within each cluster and perform transformation, and use the obtained single-point prompt, box prompt, and mask prompt as point cloud prompt information P; S5. Input the preprocessed image I and the point cloud prompt information P into the segmentation network based on the large model to output the segmentation mask of the fasteners.
2. The method for instance segmentation of aircraft surface fasteners based on point cloud hint and image fusion according to claim 1, wherein: In step S1, when performing mutual mapping on the images and point cloud data, the specific process includes the following steps: S11. Use a calibrated monocular structured light camera to photograph the aircraft surface to obtain image data and point cloud data; S12. Since both the image and the point cloud data are from the same camera system, a mapping between the point cloud and the image is established using the pre-calibrated internal and external parameter relationships. Each point in the point cloud data is represented as p i , and a connection is established with the corresponding pixel point F(x, y) in the image through the mapping relationship: i = x + y×w i is the index of the point in the point cloud, x is the column number where the pixel point is located, y is the row number where the pixel point is located, and w is the width of the overall pixels of the picture.
3. The method for instance segmentation of aircraft surface fasteners based on point cloud hint and image fusion according to claim 1, wherein: The specific process of step S2 includes the following steps: S21. For the collected images, first perform grayscaling operations on the images, convert the RGB images into grayscale images, and at the same time maximize the retention of edge and local contrast information in the original color channels. By fusing the RGB channel information through the Laplacian operator and global contrast weights, obtain the grayscaled image F1; S22. In order to suppress Gaussian noise and salt-and-pepper noise while retaining the edge structure of the fasteners, use non-local means denoising, and perform weighted averaging using the self-similarity of the images to obtain the denoised image F2; S23. Use the contrast enhancement method on the denoised image F2, limit the contrast adaptive histogram equalization, perform block equalization and limit the histogram distribution to obtain the contrast-enhanced image F3, realize the improvement of the contrast between the fasteners and the background, and at the same time avoid local over-enhancement; S24. Use the edge sharpening method, perform gradient calculation, non-maximum suppression, double-threshold detection, and edge connection in sequence, and finally obtain the preprocessed image I with clear edge contours of fasteners, providing high-precision input for subsequent segmentation.
4. The method for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion according to claim 1, wherein: The specific process of step S3 includes the following steps: S31. For the collected point cloud data Where N is the number of points, first use the statistical filtering algorithm to dynamically remove noise based on the mean and standard deviation of the neighborhood point distribution; S32. Then, use uniform downsampling based on voxel grids to divide the point cloud into regular voxels and retain the centroids to obtain the filtered and denoised point cloud P1; S33. The preliminarily processed point cloud P1 is further subjected to neighborhood construction. For each point p i , the point set within its neighborhood with a radius of r is searched Subsequently, the covariance matrix is calculated and eigenvalue decomposition is performed to obtain the normal vector of the point cloud and the eigenvalues λ1, λ2, and λ3; S34. To pre-separate the fastener body from the background skin, extract the region of interest of the fastener. Based on the plane model fitting of Random Sample Consensus (RANSAC), extract the point set belonging to the plane as the background, and the remaining points as the fastener region, obtaining the preprocessed point cloud 5. The method for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion according to claim 4, wherein: In step S4, for the preprocessed point cloud perform key point detection based on intrinsic shape features. The specific process includes the following steps: S41. First, perform significance measurement according to the eigenvalues λ1, λ2, and λ3 obtained by calculating the normal vectors in the previous step, and use the eigenvalue ratio to evaluate the significance of points: Among them, Saliency (p i ) is point p i The significance of ,∈ is a very small constant to prevent the denominator from being zero; S42. Set the threshold T sal to 0.8, and only keep the points where Saliency(p i ) > T sal . Then, perform non-maximum suppression to eliminate redundant key points and retain the locally most significant points; S43. For each point p i , compare the saliency within its neighborhood. If the saliency of p i is the maximum within the neighborhood, mark it as a key point, and finally generate a set of key points 6. The method for instance segmentation of aircraft surface fasteners based on point cloud hint and image fusion according to claim 1, wherein: In step S4, cluster the detected set of key points, extract two key information of the centroid and circle fitting according to the key points within each cluster and perform transformation, and use the obtained single-point prompt, box prompt, and mask prompt as the point cloud prompt information P. The specific process includes the following steps: S44. Cluster the detected set of key points \(K = \{k_1, k_2, \ldots, k\) m \}, using the K - means clustering algorithm, and group the key points belonging to the same fastener into one cluster \(K\) i ; S45. According to the mapping relationship between the point cloud and the image, map the key points to the pixel points in the image to obtain the key pixel points (x i , y i ); S46. Extract the centroid information as the position prompt, calculate the centroid of the set of key points in the cluster as the center position of the fastener: use the centroid (x, y) of each key point cluster as the single-point prompt; x i and y i are the coordinate positions of the key pixel points respectively, and n is the number of key pixel points in a cluster; S47. Then, use the least squares method to perform circle fitting on the key pixel points to obtain the center a, b and radius r of the fitted circle, and the objective function used is: Solve for minimizing F(a, b, r) to obtain the parameters a, b, r; S48. Calculate the circumscribed rectangle B = (a - r, b - r, a + r, b + r) using the radius r obtained by fitting a circle, and use this rectangle as a box prompt. Generate a binary mask M(x, y) using circular information, and the generation criteria are as follows: Perform a convex hull operation on the key pixel points, convert the polygon area into a binary mask, and generate a mask prompt M(x, y) as the mask prompt. The finally obtained single-point prompt, box prompt, and mask prompt are used as the prompt information P.
7. The method for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion according to claim 1, wherein: In step S5, the large model-based segmentation network includes: an image encoder, which uses a pre-trained ViT model and fine-tunes the query-key-value parameters of the multi-head attention layer through the LoRA technique to adapt to specific segmentation tasks; a self-generated prompt generator, which uses a neural network to generate prompt embeddings that match the network prompt input format from the prompt information generated by the point cloud; a mask decoder, which based on the self-attention and cross-attention mechanisms, fuses the image embeddings output by the image encoder with the prompt embeddings generated by the self-generated prompt generator, and outputs the final instance segmentation mask.
8. The method for instance segmentation of aircraft surface fasteners based on point cloud hint and image fusion according to claim 7, wherein: The specific process of step S5 includes the following steps: S51. Image encoding process. The preprocessed image I is fed into the image encoder as input, passes through 12 Transformer layers, and the query, key, and value matrices are fine-tuned by low-rank decomposition in the multi-head attention module of each layer. Finally, after being processed by the image encoder, the image feature embedding E is obtained. I , which has global semantic information and local detail representation. S52. Input the single-point prompt, box prompt, and mask prompt generated from the point cloud into the self-generated prompt encoder to respectively obtain prompt embedding vectors \(E\) of a fixed dimension point , \(E\) box and \(E\) mask . Finally, concatenate the above three embedding vectors Concat to obtain the final prompt embedding vector \(E\) P : S53. Prompt information embedding E P The self conducts internal information interaction through the self-attention layer, and the updated prompt embedding E' P can be expressed as: It is P = SelfAttn(E P ) + E P Where SelfAttn represents the multi-head self-attention operation, which allows for feature reallocation within the prompt information and strengthens the expression of local geometric priors. S54. Embed the image features into E I and the updated prompt embedding vector E' P Input the mask decoder to perform mask decoding and prediction, and perform fusion through multi-layer self-attention and cross-attention; the updated prompt embedding E' P As the query, the image feature embedding E I Is input into the cross-attention layer as the key and value to achieve the interactive fusion of prompt information and image semantics; S55. Finally, after 12 layers of Transformer decoding, the network generates multiple candidate segmentation masks, and selects the mask with the highest score as the final output to complete the instance segmentation of the fastener.
9. The method for instance segmentation of aircraft surface fasteners based on point cloud prompting and image fusion according to claim 8, wherein: In step S52, the single-point prompt, box prompt, and mask prompt generated from the point cloud are input into the self-generated prompt encoder to obtain prompt embedding vectors E point , E box and E mask , and the specific process includes: for the point prompt, it is mapped to a 128-dimensional embedding vector E point through a point prompt feed-forward neural network f point , and f point (·) adopts a two-layer fully connected network. The first layer: the fully connected layer FC maps the input dimension 2 to the hidden layer with 64 dimensions, and then uses ReLU activation; the second layer: the fully connected layer maps the 64-dimensional hidden layer to the output embedding with 128 dimensions; for the box prompt, it is mapped to a 128-dimensional embedding vector E box through another box prompt feed-forward neural network f box , and the difference between f box (·) and f point (·) is that the box prompt feed-forward neural network modifies the input end to 4 dimensions, and the rest of the internal structure is the same; for the mask prompt, a mask prompt lightweight convolutional neural network is adopted. The input is the binary mask M(x, y), and the internal network structure is 3 convolutional layers of 3×3. After each convolutional layer, there is a batch normalization layer and a ReLU activation function. At the end of the network, a 128-dimensional embedding space E mask is obtained through global average pooling and a fully connected layer.