Aircraft surface mark point coarse positioning method based on deep learning and medium

Through a deep learning-based method, combined with hollow convolution and edge-aware convolution for feature extraction, the problems of inflexible threshold selection and insufficient robustness in the coarse positioning method of surface marking points in the existing aircraft are solved, and more efficient automated positioning process and higher detection accuracy are achieved.

CN120088305AInactive Publication Date: 2025-06-03LOW SPEED AERODYNAMIC INST OF CHINESE AERODYNAMIC RES & DEV CENT

Patent Information

Application Number
CN202510562111.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing method of coarse positioning of surface marking points in aircraft, the threshold selection is inflexible, the adaptability of lighting and texture changes is poor, the robustness is insufficient, and manual parameter adjustment is required, which increases complexity and time cost.

Method used

A deep learning-based method is adopted to perform feature extraction and residual processing through the backbone network, and feature extraction is performed by combining multiple expansion rates of hollow convolution, horizontal edge-aware convolution and vertical edge-aware convolution to integrate spatial information and edge information to reduce false detection and missed detection.

Benefits of technology

Improve the accuracy and robustness of marker point detection, reduce the need for manual parameter adjustment and manual intervention, and achieve a more efficient automated positioning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088305A_ABST
    Figure CN120088305A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning-based aircraft surface mark point coarse positioning method and a medium, and relates to the technical field of aviation. According to the scheme, the surface image information of the aircraft is collected, the image information is preprocessed and then input into the mark point coarse positioning neural network for mark point recognition, feature extraction and residual processing are carried out through the backbone network to obtain four feature maps, and then the four feature maps serve as input of the neck network; and after feature fusion processing of the feature pyramid network, a mark point identification feature map is obtained through detection of a detection head. And finally, based on the final detection result of the mark point, carrying out coarse positioning on the mark point. Due to the fact that feature extraction is carried out by combining cavity convolution of multiple expansion rates, horizontal edge perception convolution processing and vertical edge perception convolution processing, space information and edge information can be integrated in feature extraction, and the phenomena of false detection and missing detection of mark points in aircraft surface mark point coarse positioning based on deep learning are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aviation technology, and particularly to a method and medium for rough positioning of marking points on the surface of an aircraft based on deep learning. Background Art

[0002] The pressure-sensitive paint technology is a non-contact full-field pressure measurement method based on image processing. By uniformly coating a pressure-sensitive paint on the surface of a model, irradiating the surface of the model with a light source such as a laser or an ultraviolet LED, and using the oxygen quenching reaction that occurs when the luminescent molecules in the pressure-sensitive paint come into contact with oxygen, the change in the luminescence intensity is measured to infer the pressure distribution. Currently, the pressure-sensitive paint technology has been widely used in the pressure measurement on the surface of aerospace aircraft.

[0003] In the PSP technology, it is necessary to arrange image marking points to carry out image correction and registration. The transformation parameters of the two images are obtained mainly by matching the marking points on the windless reference image and the windy working image taken in the aircraft pressure-sensitive paint test for registration, which includes two steps: rough positioning and precise positioning. In the existing methods for positioning marking points on the surface of an aircraft model, rough positioning usually adopts image processing technology based on thresholds. However, this method has the following problems: First, the threshold selection has a significant impact on the positioning result. Improper setting may lead to inaccurate positioning, and the threshold can only be set uniformly, lacking flexibility. Second, the threshold method has poor adaptability to changes in illumination, the texture of the aircraft surface, and the shape of the marking points. In addition, when there is noise or the image quality is poor, its robustness is insufficient and the positioning accuracy decreases. Finally, the threshold method requires manual parameter adjustment, which increases the complexity and time cost and is difficult to achieve automation.

[0004] The rapid development of deep learning has greatly promoted the progress of image processing and computer vision technologies. As one of the core tasks in this field, object detection has also made significant breakthroughs. In the field of small object detection, object detection algorithms based on deep learning have gradually become the mainstream. Under the CNN architecture, mainstream object detection algorithms can be further divided into two-stage and single-stage methods. Typical two-stage detectors (such as the RCNN series) first generate candidate regions using selective search or region proposal networks, and then further perform object classification and positioning. Although such algorithms perform well in detection accuracy, their computational overhead is large, which limits their popularity in real-time applications. In contrast, single-stage detection algorithms (such as the YOLO series, SSD, etc.) omit the generation of candidate regions and directly perform object prediction on the input image, significantly improving the detection speed, while the accuracy is gradually approaching or even exceeding that of two-stage methods. Considering the requirements of real-time performance and computational resource balance in the usage scenario, YOLOV8 is used as the baseline model for algorithm improvement to improve the detection accuracy of small marking points while maintaining the detection efficiency and computational consumption. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide a method and medium for rough positioning of marked points on the surface of an aircraft based on deep learning, which has the characteristics of reducing the misdetection and missed detection of marked points and reducing the manual adjustment of parameters and manual intervention in the positioning process.

[0006] In a first aspect, in one embodiment, a method for rough positioning of marked points on the surface of an aircraft based on deep learning is provided, including: Collecting the surface image information of the aircraft, preprocessing the image information, and inputting it into a neural network for rough positioning of marked points based on deep learning for marked point recognition, including: Taking the preprocessed surface image information as input, and obtaining a second feature map, a third feature map, a fourth feature map, and a fifth feature map after feature extraction and residual processing by the backbone network; Taking the second feature map, the third feature map, the fourth feature map, and the fifth feature map as inputs of the neck network respectively, after multi-scale feature fusion processing by the feature pyramid network, and then obtaining the final detection result of the marked points through detection by the detection head; Based on the final detection result of the marked points, rough positioning of the marked points is performed; Among them, taking the preprocessed surface image information as input, and obtaining a second feature map, a third feature map, a fourth feature map, and a fifth feature map after feature extraction and residual processing by the backbone network, including: taking the preprocessed image information as input for edge perception processing to obtain a first feature map, and performing four-level residual processing on the first feature map in sequence, corresponding to obtaining the second feature map, the third feature map, the fourth feature map, and the fifth feature map in sequence; The above-mentioned taking the preprocessed image information as input and first performing edge perception processing to obtain a first feature map includes: Performing ordinary convolution processing on the preprocessed image information to obtain a first preliminary feature map; Performing atrous convolution processing on the first preliminary feature map with multiple different dilation rates to obtain corresponding multiple second preliminary feature maps, and performing feature fusion processing on the multiple second preliminary feature maps and then performing normalization activation to obtain a third preliminary feature map; Performing vertical edge perception convolution processing and horizontal edge perception convolution processing on the first preliminary feature map respectively to obtain a fourth preliminary feature map and a fifth preliminary feature map correspondingly, and performing element-wise addition on the fourth preliminary feature map and the fifth preliminary feature map to obtain a sixth preliminary feature map; Performing feature fusion on the third preliminary feature map and the sixth preliminary feature map, and then performing convolution processing to obtain a first feature map.

[0007] In one embodiment, the method includes: performing atrous convolutions with multiple different dilation rates on the first preliminary feature map respectively to obtain corresponding multiple second preliminary feature maps, performing feature fusion on the multiple second preliminary feature maps, and then performing normalization activation to obtain a third preliminary feature map. Among them, represents the first preliminary feature map, , and represent three second preliminary feature maps obtained by atrous convolutions with three different dilation rates; represents an atrous convolution with a convolution kernel size of k and a dilation rate of e; represents the third preliminary feature map, BN represents normalization processing, Relu represents the Relu activation function, and Concat represents channel concatenation processing.

[0008] In one embodiment, the method includes: performing vertical edge-aware convolution processing and horizontal edge-aware convolution processing on the first preliminary feature map respectively to obtain a fourth preliminary feature map and a fifth preliminary feature map correspondingly, and performing element-wise addition on the fourth preliminary feature map and the fifth preliminary feature map to obtain a sixth preliminary feature map. Among them, represents the Sobel operator, represents the first preliminary feature map, represents the fourth preliminary feature map, represents the fifth preliminary feature map, represents the sixth preliminary feature map, T represents transpose, represents a convolution operation, and Sum represents element-wise addition.

[0009] In one embodiment, the method includes: performing feature fusion on the third preliminary feature map and the sixth preliminary feature map, and then performing convolution processing to obtain a first feature map. Among them, represents the first feature map, represents a convolution with a convolution kernel of 1×1, The convolution with a convolution kernel of 3×3 is denoted as Convolution, and Concat represents the channel concatenation process. denotes the third preliminary feature map. denotes the sixth preliminary feature map.

[0010] In a second aspect, an embodiment provides a computer-readable storage medium, characterized in that a program is stored in the medium, and the program can be loaded and executed by a processor to perform the method for rough positioning of aircraft surface marking points according to any one of the above embodiments.

[0011] The beneficial effects of the present invention are as follows: Since in the feature extraction by the backbone network, the dilated convolution with multiple dilation rates, the horizontal edge-aware convolution processing, and the vertical edge-aware convolution processing are combined for feature extraction, the spatial information and edge information can be integrated in the feature extraction, reducing the false detection and missed detection of marking points in the rough positioning of aircraft surface marking points based on deep learning, thereby reducing the manual adjustment of parameters and the manual intervention in the positioning process. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic diagram of the network architecture for rough positioning of aircraft surface marking points based on deep learning according to an embodiment of the present application; Figure 2 is a schematic diagram of the method flow for inputting the preprocessed image information into the neural network for rough positioning of marking points based on deep learning to identify marking points according to an embodiment of the present application; Figure 3 is a schematic diagram of the method flow for obtaining the first feature map after edge-aware processing by using the preprocessed image information as input according to an embodiment of the present application; Figure 4 is a schematic diagram of the method flow for multi-scale feature fusion processing according to an embodiment of the present application; Figure 5 is a schematic diagram of the method flow for the detection head detection according to an embodiment of the present application; Figure 6 is a schematic diagram of the visualization result of the aircraft marking point data target set in a dense scene obtained by the detection result of the baseline model YOLOv8; Figure 7 is a schematic diagram of the visualization result of the aircraft marking point data target set in a dense scene obtained by the detection result of the method of the present application; Figure 8 is a schematic diagram of the visualization result of the aircraft marking point data set in a scene with texture interference obtained by the detection result of the baseline model YOLOv8; Figure 9 is a schematic diagram of the visualization result of the aircraft marking point data set in a scene with texture interference obtained by the detection result of the method of the present application. Detailed Implementation Modes

[0013] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific implementation modes. Similar elements in different implementation modes adopt related similar element numbers. In the following implementation modes, many detailed descriptions are provided to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification, in order to avoid the core part of the present application being overwhelmed by excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the descriptions in the specification and the general technical knowledge in the art.

[0014] In addition, the features, operations, or characteristics described in the specification can be combined in any appropriate manner to form various implementation modes. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for clearly describing a certain embodiment, and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0015] The serial numbers assigned to the components herein, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meanings.

[0016] For the convenience of explaining the inventive concept of the present application, the following briefly describes the technology for identifying marked points on the surface of an aircraft.

[0017] In the current marked point recognition technology based on deep learning, a convolutional attention module is proposed to adaptively select and fuse the most representative features, while enhancing the ability to extract important information from different regions of the image, thereby improving the spatial detail perception effect.

[0018] However, the applicant found in the research that the marked points on the surface of the aircraft have strong edge features. However, in the baseline model and related optimization methods, if only focusing on spatial information and not specifically highlighting the importance of edge features, it will lead to the loss of edge features in the low-level information, resulting in false detection and missed detection.

[0019] In view of this, in the embodiments of the present application, a method and medium for rough positioning of marked points on the surface of an aircraft based on deep learning are provided. The surface image information of the aircraft is collected, and after preprocessing the image information, it is input into a neural network for rough positioning of marked points based on deep learning for marked point recognition, including obtaining four feature maps after feature extraction and residual processing through a backbone network, and then using these four feature maps as the input of the neck network. After feature fusion processing by a feature pyramid network, a marked point recognition feature map is obtained through detection by a detection head. Finally, based on the final detection result of the marked points, rough positioning of the marked points is performed. Since in the feature extraction through the backbone network, atrous convolution with multiple dilation rates, horizontal edge perception convolution processing, and vertical edge perception convolution processing are combined for feature extraction, spatial information and edge information can be integrated during feature extraction, reducing the false detection and missed detection of marked points in the rough positioning of marked points on the surface of an aircraft based on deep learning, thereby reducing manual parameter adjustment and manual intervention in the positioning process.

[0020] To fully illustrate a method for rough positioning of marked points on the surface of an aircraft based on deep learning in the embodiments of the present application, the following first describes a partial training process of a neural network for rough positioning of marked points required during the rough positioning process.

[0021] First is data collection and preparation: In the pressure-sensitive paint experiment, image data of the aircraft model can be captured using a high-speed camera, including reference images under windless conditions and working images under windy conditions. Among them, 256 pictures under different angles of attack and wind speeds are selected. The surface of the aircraft model is coated with a pressure-sensitive paint coating, and multiple marked points are evenly distributed at its edge contour.

[0022] The marked points in the original dataset are manually annotated to generate xml files with position coordinates and attribute names, and then converted into txt format. The dataset is expanded to 1200 for training using image processing techniques, including adding noise, rotation, scaling, etc., and divided into a training set, a test set, and a validation set according to a ratio of 7:2:1.

[0023] Construct a preliminary network architecture, including a backbone network, a neck network, and a detection head. In one embodiment, please refer to the network architecture Figure 1 . Among them, the backbone network includes an edge perception module and a multi-stage residual module. The neck network uses a feature pyramid network to receive multi-stage outputs of the multi-stage residual module, and after processing, outputs them to the detection head. Figure 1 In, 1440×1440 is the input image after data preprocessing, ESFM is the edge perception network, Cf2 represents convolution, Conv represents convolution, SPPF represents pooling, MFH is the multi-scale feature fusion network, UpSample represents upsampling, Concat represents concatenation, and LDEDH is the detection head.

[0024] In current rough positioning and recognition of marking points based on deep learning, the YOLOv8 network model is usually used for object detection. However, the applicant found in the research that the YOLOv8 network model generally uses two ordinary convolutional layers to extract preliminary features. The standard convolution adopted mainly focuses on local spatial information and does not specifically optimize the convolutional kernel to highlight edge features.

[0025] In view of this, in the embodiments of the present application, for preliminary feature extraction, an edge perception module is proposed. By strengthening edge information, the geometric contour and details of the target are captured, and at the same time, the global perception ability is improved by combining spatial information, so as to make up for the deficiencies of using the YOLOv8 network model.

[0026] The constructed preliminary network architecture is placed in the configured training environment, the corresponding parameter file is modified, and the network model is trained using the pre-divided training set and validation set to obtain a trained rough positioning neural network for marking points. Considering the situation that the number of self-made dataset pictures is small and the picture resolution is high, to ensure the training quality, help converge faster on the small dataset and avoid falling into local minima, the Adam optimizer is used for training, and the input image resolution is 1440×1440.

[0027] Based on the above-trained rough positioning neural network for marking points, the method for rough positioning of marking points on the aircraft surface based on deep learning in the embodiments of the present application includes: Collect the surface image information of the aircraft, preprocess the image information, and input it into the rough positioning neural network for marking points based on deep learning for marking point recognition. Please refer to Figure 2 , specifically, it may include: Step S10, using the preprocessed surface image information as the input, after feature extraction and residual processing by the backbone network, the second feature map, the third feature map, the fourth feature map, and the fifth feature map are obtained.

[0028] Specifically, using the preprocessed image information as the input for edge perception processing to obtain the first feature map, and performing four-level residual processing on the first feature map in sequence to obtain the second feature map, the third feature map, the fourth feature map, and the fifth feature map in sequence.

[0029] Please refer to Figure 3 , using the preprocessed image information as the input for edge perception processing to obtain the first feature map, including: Step S100, performing ordinary convolution processing on the preprocessed image information to obtain the first preliminary feature map.

[0030] In one embodiment, the preprocessed image information is first subjected to a 3×3 convolution to obtain a first preliminary feature map, and then the first preliminary feature map is input into two parallel branches to separately extract target spatial information and edge information.

[0031] Step S200: Perform atrous convolutions with multiple different dilation rates on the first preliminary feature map respectively to obtain corresponding multiple second preliminary feature maps, and perform feature fusion processing on these multiple second preliminary feature maps and then perform normalization activation to obtain a third preliminary feature map.

[0032] In one embodiment, to better extract target spatial information, atrous convolutions with three different dilation rates are respectively used to perform atrous convolution processing on the first preliminary feature map, which can be expressed as: where represents the first preliminary feature map, , and represent three second preliminary feature maps obtained by atrous convolution processing with three different dilation rates; represents an atrous convolution with a convolution kernel size of k and a dilation rate of e.

[0033] In one embodiment, then perform feature fusion processing on multiple second preliminary feature maps and then perform normalization activation to obtain a third preliminary feature map, which can be expressed as: where represents the third preliminary feature map, BN represents normalization processing, Relu represents the Relu activation function, and Concat represents channel concatenation processing.

[0034] In this embodiment, by inserting holes between adjacent elements of the convolution kernel, the purpose of increasing the receptive field of the convolution kernel without increasing the number of parameters and computational complexity is achieved. These convolution results complement each other, fully fuse local and global features, and thus retain rich spatial details.

[0035] Step S300: Perform vertical edge-aware convolution processing and horizontal edge-aware convolution processing on the first preliminary feature map respectively to obtain a fourth preliminary feature map and a fifth preliminary feature map correspondingly, and perform element-wise addition on the fourth preliminary feature map and the fifth preliminary feature map to obtain a sixth preliminary feature map.

[0036] To better extract edge information, it is implemented in combination with the Sobel edge filter. First, the Sobel operator is defined. In one embodiment, the Sobel operator is a convolution kernel with the same number of channels as the input feature map, used to detect edges in the vertical direction, and the transposed one is used to detect edges in the horizontal direction. Then step S300 can be expressed as: Among them, represents the Sobel operator, represents the first preliminary feature map, represents the fourth preliminary feature map, represents the fifth preliminary feature map, represents the sixth preliminary feature map, T represents transpose, represents the convolution operation, and Sum represents element-wise addition.

[0037] In this way, through the convolution operation with the predefined fixed Sobel convolution kernel, the gradient information in the vertical and horizontal directions is calculated to output an image representation with edge features.

[0038] In step S400, after fusing the features of the third preliminary feature map and the sixth preliminary feature map, convolution processing is performed to obtain the first feature map.

[0039] To integrate the spatial information and edge information, first, the feature maps obtained from the two branches are concatenated in the channel dimension. In one embodiment, the concatenated feature map is first processed by a 3×3 convolution kernel to enhance the feature expression ability of the spatial edge information, enabling the model to more comprehensively understand the image content. Then a 1×1 convolution kernel is used to adjust the output channels so that the output high-quality feature map can directly connect to the subsequent tasks. Then step S400 can be expressed as: Among them, represents the first feature map, represents the convolution with a 1×1 convolution kernel, represents the convolution with a 3×3 convolution kernel, Concat represents channel concatenation processing, represents the third preliminary feature map, represents the sixth preliminary feature map.

[0040] Based on the obtained first feature map, after four-level residual processing, the second feature map, the third feature map, the fourth feature map, and the fifth feature map are obtained in sequence.

[0041] In one embodiment, the first feature map is subjected to a first-level residual process to obtain a second feature map, the second feature map is successively subjected to convolution and residual processes to obtain a third feature map, the third feature map is successively subjected to convolution and residual processes to obtain a fourth feature map, and the fourth feature map is successively subjected to convolution, residual, and pooling processes to obtain a fifth feature map.

[0042] Step S20: Respectively take the second feature map, the third feature map, the fourth feature map, and the fifth feature map as the inputs of the neck network. After multi-scale feature fusion processing by the feature pyramid network, the final detection result of the marked points is obtained through detection by the detection head.

[0043] In the application scenario of the present application, the target detection image often contains targets of different scales. Therefore, an efficient feature fusion strategy is crucial for improving the detection performance. The feature differences in the image caused by the scale differences of the targets are a major difficulty in multi-scale target detection. Single-scale features are difficult to capture global information and local details simultaneously, thus affecting the detection effect.

[0044] However, the applicant found in the research that current advanced FPN (Feature Pyramid Networks) architecture networks such as BiFPN (Weighted Bidirectional Feature Pyramid Network), AFPN (Asymptotic Feature Pyramid Network), etc. lack intermediate layers for aggregating far-layer information, and feature interaction is limited to adjacent layers, restricting the effect of feature fusion. Due to the lack of intermediate layer aggregation during the transmission of far-layer features, it is prone to gradual attenuation or loss layer by layer, making it difficult for high-level semantics to directly affect low-level expressions. In addition, relying solely on adjacent layer fusion will weaken the ability to capture long-range dependencies, restricting the understanding of the relevance of targets of different scales, thus resulting in unbalanced multi-scale target detection effects.

[0045] In view of this, in one embodiment of the present application, a new feature pyramid network is provided to implement the specific process of respectively taking the second feature map, the third feature map, the fourth feature map, and the fifth feature map as the inputs of the neck network and performing multi-scale feature fusion processing by the feature pyramid network in step S20, including: Step S2011: Upsample the fifth feature map and perform feature fusion with the fourth feature map to obtain a first fusion feature map.

[0046] Step S2012: Perform a residual process on the first fusion feature map and perform a first-level multi-scale feature fusion process with the second feature map and the third feature map to obtain a first multi-scale fusion feature map.

[0047] Step S2013: Respectively perform upsampling processing and convolution processing on the first multi-scale fusion feature map to obtain a corresponding first upsampled feature map and a first convolutional feature map.

[0048] In step S2014, the first upsampled feature map and the second feature map are subjected to feature fusion and then residual processing to obtain a second fused feature map; after the first convolutional feature map and the first fused feature map are subjected to feature fusion, residual processing is performed to obtain a third fused feature map.

[0049] In step S2015, the second fused feature map, the third fused feature map, and the first multi-scale fused feature map are subjected to second-level multi-scale feature fusion processing to obtain a second multi-scale fused feature map.

[0050] In step S2016, the second multi-scale fused feature map is respectively subjected to upsampling processing and convolutional processing to obtain a corresponding second upsampled feature map and a second convolutional feature map.

[0051] In step S2017, the first upsampled feature map, the second upsampled feature map, and the second fused feature map are subjected to feature fusion and then residual processing to obtain a sixth feature map; after the first convolutional feature map, the second convolutional feature map, and the third fused feature map are subjected to feature fusion, residual processing is performed to obtain a seventh feature map.

[0052] Through the first-level multi-scale feature fusion processing, the second feature map, the third feature map, and the fourth feature map are subjected to feature aggregation and then diffusion, and then through the second-level multi-scale feature fusion processing, the diffused features are subjected to aggregation processing again, so that the feature pyramid network introduces an intermediate layer as a bridge for far-layer feature interaction, effectively shortening the information transmission path, and improving the global perception ability and detection accuracy.

[0053] The sixth feature map is used as the input of the detection head of the P2 layer, the seventh feature map is used as the input of the detection head of the P4 layer, and the second multi-scale fused feature map is used as the input of the detection head of the P3 layer. In this way, the high-resolution detection head of the P2 layer is increased, and the detection head of the P5 layer is removed to optimize the structure, retaining more fine-grained information such as edges and textures to enhance the recognition of small targets.

[0054] In one embodiment, in order to reduce the information loss caused by layer-by-layer transmission, a new multi-scale feature fusion processing method is given, which is applied to the multi-scale feature fusion processing of any of the above levels. Please refer to Figure 4 , and this processing method includes: In step S1000, among the three inputs of the multi-scale feature fusion processing, the input corresponding to the detection head of the P2 layer is subjected to upsampling processing and then convolutional processing to obtain a first multi-scale feature map, the input corresponding to the detection head of the P3 layer is subjected to convolutional processing to obtain a second multi-scale feature map, and the input corresponding to the detection head of the P4 layer is subjected to downsampling processing to obtain a third multi-scale feature map.

[0055] Please refer to Figure 1, for the first-level multi-scale feature fusion process, the input corresponding to the P2 layer detection head is the second feature map, the input corresponding to the P3 layer detection head is the first multi-scale fusion feature map, and the input corresponding to the P4 layer detection head is the feature map obtained by performing residual processing on the first fusion feature map.

[0056] For the second-level multi-scale feature fusion process, the input corresponding to the P2 layer detection head is the second fusion feature map, the input corresponding to the P3 layer detection head is the third feature map, and the input corresponding to the P4 layer detection head is the third fusion feature map.

[0057] Step S2000, perform feature fusion on the first multi-scale feature map, the second multi-scale feature map, and the third multi-scale feature map to obtain the first primary multi-scale fusion feature map.

[0058] In one embodiment, step S2000 can be expressed as: Among them, represents the first primary multi-scale fusion feature map, SiLu represents the SiLu activation function, Adown represents downsampling, P 2 , P 3 and P 4 respectively represent the inputs corresponding to the P2 layer detection head, the P3 layer detection head, and the P4 layer detection head.

[0059] Through step S2000, the sizes and channel numbers of feature maps at different levels are unified.

[0060] Step S3000, perform depthwise convolutions with three different convolutional kernel sizes on the first primary multi-scale fusion feature map to obtain the corresponding fourth multi-scale feature map, fifth multi-scale feature map, and sixth multi-scale feature map.

[0061] In one embodiment, the depthwise convolutions with three different convolutional kernel sizes include a 1×1 depthwise convolution, a 3×3 depthwise convolution, and a 5×5 depthwise convolution. The 1×1 depthwise convolution captures local detailed features, the 3×3 depthwise convolution balances local features and context information, and the 5×5 depthwise convolution extracts global features with a larger receptive field. In this way, it is ensured that the network can simultaneously focus on the details of small-scale targets and the overall information of large-scale targets. In addition, the depthwise convolutions adopted here further reduce the computational amount and parameter scale, enabling the network to support information processing at more scales with limited resources.

[0062] Step S4000, add the fourth multi-scale feature map, the fifth multi-scale feature map, and the sixth multi-scale feature map element-wise to obtain the second primary multi-scale fusion feature map.

[0063] The features of each branch are fused through element-wise addition to integrate information from different receptive fields.

[0064] Step S5000: After performing pointwise convolution on the second primary multi-scale fusion feature map, it is added element-wise to the first primary multi-scale fusion feature map to complete the multi-scale feature fusion process.

[0065] In one embodiment, 1×1 pointwise convolution is used for channel compression and feature weighting, reducing redundancy and enhancing the correlation between features.

[0066] Finally, by adding the fused feature maps processed by the three branches to the original first primary multi-scale fusion feature map element-wise, while retaining the original information, the feature expression is strengthened. In this way, the multi-scale feature fusion process used for feature aggregation provides a more comprehensive feature representation through multi-branch processing, multi-scale feature fusion, and residual connection.

[0067] In one embodiment, the process from step S3000 to step S5000 can be expressed as: Where, represents the output of the multi-scale feature fusion process, PWConv represents pointwise convolution; DWConv represents depthwise convolution, and its subscript is the size of the convolution kernel; represents the output after pointwise convolution.

[0068] In the feature diffusion stage, the fused multi-scale information is re-transmitted to each layer, enabling each layer to retain its own features while fusing the information of other layers, thereby enhancing the generalization ability of object detection.

[0069] The detection head simultaneously completes the classification and regression tasks through a single convolutional neural network. The goal of the classification task is to classify the detected objects into one of the predefined categories, and the goal of the regression task is to predict the precise location and size of the objects.

[0070] The applicant found in the research that due to the scene where it is difficult to distinguish the target edge and surface texture information in the rough positioning of the marked points on the aircraft surface, the generalization ability of the original detection head is significantly limited. The reason is that the detection ability of small targets is restricted.

[0071] The applicant also found in the research that the shallow feature map is rich in detailed information and is crucial for small target detection. To improve the recognition ability, the P2 layer detection head can be introduced during the detection process, and the P5 layer detection head can be removed to improve the recognition ability of small targets. However, since the convolutional computational load is positively correlated with the resolution of the input feature map, performing the detection task at the P2 layer will increase the computational burden.

[0072] In view of this, in order to reduce the scale of model parameters and optimize the detection performance while controlling the computing cost, the present application proposes a new detection head detection method. Please refer to Figure 5 , including: Step S2021: For any detection head, the input of the detection head is subjected to convolution processing to obtain an input feature map.

[0073] Step S2022: The input feature map is subjected to at least one level of detailed feature enhancement processing to obtain an output feature map.

[0074] In one embodiment, to improve the expressive ability, the detailed feature enhancement processing includes: respectively performing ordinary convolution processing, central difference convolution processing, corner difference convolution processing, vertical difference convolution processing, and horizontal difference convolution processing on the input feature map, and then performing element-wise addition on the five convolution processing results to obtain a detailed enhancement feature map.

[0075] In one embodiment, the detailed feature enhancement processing can be expressed as: where, represents the detailed enhancement feature map, DEConv represents the detailed enhancement convolution, represents the feature map input for the detailed feature enhancement processing; represents the convolution kernel of the i-th convolution processing among the five convolution processings, 1 ≤ i ≤ 5.

[0076] In the above solution, ordinary convolution is used to enhance intensity information, and difference convolution strengthens gradient features. Finally, the output is obtained through addition fusion. The core lies in that on the one hand, ordinary convolution is used to integrate prior information, enhance feature representation and generalization ability. On the other hand, combined with four difference convolutions, the ability to extract target edge and surface texture information can be improved. On the third hand, with the help of the reparameterization technology of addition fusion, it is equivalently converted into ordinary convolution, so that no additional computational overhead is required during the inference stage.

[0077] In one embodiment, in order to measure the balance between accuracy and lightweight requirements, the input feature map is subjected to two levels of detailed feature enhancement processing to obtain an output feature map.

[0078] For the processing of two branches, the first branch adopts ordinary convolution processing, and the second branch adopts four parallel difference convolution processings; the feature maps after the four difference convolution processings are merged, and the merged feature map is added element-wise to the feature map after the ordinary convolution processing of the first branch to obtain a detailed enhancement feature map.

[0079] Step S2023: The output feature map is subjected to classification and regression processing, and then mapped to output to obtain the detection result.

[0080] In one embodiment, after classifying and regressing the output feature map, the detection result is mapped and output, including: performing a 1×1 two-dimensional convolution on the output feature map in the classification branch to obtain a classification loss, and performing a 1×1 two-dimensional convolution on the output feature map in the regression branch and then performing an adaptive scaling process to obtain a regression loss; performing an image mapping on the obtained classification feature map and regression feature map to obtain the detection result.

[0081] Step S2024, mapping the detection results of all detection heads to obtain the final detection result.

[0082] Finally, map the detection results of the P2 detection head, P3 detection head, and P4 detection heads to obtain the final detection result. In the embodiment of the present application, the classification and regression predictions share three convolution weights, which are merged through a weight sharing mechanism, reducing redundant calculations, and introducing a Scale layer in the regression branch for adaptive scaling to eliminate the target scale difference.

[0083] Step S30, based on the obtained final detection result, perform rough positioning on the marked points.

[0084] Based on the above-mentioned deep learning-based method for rough positioning of aircraft surface marked points in any one of the embodiments, during the feature extraction by the backbone network, by combining dilated convolutions with multiple dilation rates, horizontal edge-aware convolution processing, and vertical edge-aware convolution processing for feature extraction, spatial information and edge information can be integrated during feature extraction, reducing the misdetection and missed detection phenomena of marked points in the deep learning-based rough positioning of aircraft surface marked points, thereby reducing manual parameter adjustment and manual intervention in the positioning process.

[0085] To verify the effectiveness of the above-mentioned deep learning-based method for rough positioning of aircraft surface marked points in the embodiments, classic object detection evaluation metrics are used for evaluation. The mAP based on the mean average precision (mAP) 0.5 and the mAP 0.5:0.95 As the core comprehensive metrics, the performance of the model at different detection thresholds is evaluated by combining Precision and Recall.

[0086] Please refer to Figures 6 to 9 (In the figure, the dark blue content is the pressure hole label, and the sky blue content is the marked point label), Figure 6 and Figure 8 represent the detection results of the baseline model YOLOv8, Figure 7 and Figure 9 represent the detection results of the method of the present application, Figure 6 and Figure 7 are the visualization results in the dense scenario of the aircraft marked point data target set, Figure 8 and Figure 9Visualization results in the scenario of texture interference in the aircraft marker point dataset. It can be seen that by enhancing the extraction of the edge and shape features of the pressure holes, the improved algorithm can overcome the interference of surface pits and texture backgrounds on the aircraft and detect the pressure holes missed by the baseline algorithm.

[0087] In an embodiment of the present application, a computer-readable storage medium is provided. A program is stored on the storage medium, and the stored program includes the method in any one of the above embodiments that can be loaded and processed by a processor.

[0088] Those skilled in the art can understand that all or part of the functions of the above methods can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions are implemented by a computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, the above all or part of the functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, downloaded or copied and saved to the memory of the local device, or the system of the local device is updated. When the processor executes the program in the memory, the above all or part of the functions in the embodiments can be implemented.

[0089] The above uses specific examples to elaborate on the present invention, which is only for helping to understand the present invention and is not used to limit the present invention. For those skilled in the art of the present invention, several simple deductions, deformations or substitutions can be made according to the idea of the present invention.

Claims

1. A method for rough positioning of aircraft surface markers based on deep learning, characterized in that: include: Collect the surface image information of the aircraft, and input the image information into the deep learning-based marker point coarse positioning neural network for marker point recognition after preprocessing, including: The preprocessed surface image information is used as input, and the backbone network performs feature extraction and residual processing to obtain the second feature map, the third feature map, the fourth feature map and the fifth feature map; The second feature map, the third feature map, the fourth feature map and the fifth feature map are respectively used as the input of the neck network, and after the multi-scale feature fusion processing of the feature pyramid network, the final detection result of the marker point is obtained by the detection head; Based on the final detection results of the marker points, the marker points are roughly positioned; The method uses the preprocessed surface image information as input, performs feature extraction and residual processing through a backbone network, and obtains a second feature map, a third feature map, a fourth feature map, and a fifth feature map, including: using the preprocessed image information as input and performing edge perception processing to obtain a first feature map, and sequentially performing four-level residual processing on the first feature map to obtain a second feature map, a third feature map, a fourth feature map, and a fifth feature map in sequence; The method of taking the preprocessed image information as input and performing edge sensing processing to obtain the first feature map includes: Performing ordinary convolution processing on the preprocessed image information to obtain a first preliminary feature map; Performing a plurality of dilation rate atrous convolution processing on the first preliminary feature map to obtain a plurality of corresponding second preliminary feature maps, and performing feature fusion processing on the plurality of second preliminary feature maps and then performing normalization activation to obtain a third preliminary feature map; Performing vertical edge-aware convolution processing and horizontal edge-aware convolution processing on the first preliminary feature map respectively to obtain a fourth preliminary feature map and a fifth preliminary feature map respectively, and performing element-by-element addition of the fourth preliminary feature map and the fifth preliminary feature map to obtain a sixth preliminary feature map; After the third preliminary feature map and the sixth preliminary feature map are feature fused, a convolution process is performed to obtain a first feature map.

2. The method for coarse positioning of aircraft surface marker points according to claim 1, characterized in that: The method of performing a plurality of dilation rate-different dilation convolution processes on the first preliminary feature map to obtain a plurality of corresponding second preliminary feature maps, and performing feature fusion processing on the plurality of second preliminary feature maps and then performing normalization activation to obtain a third preliminary feature map includes: in, represents the first preliminary feature map, , and Represents three second preliminary feature maps obtained by dilated convolution with three different dilation rates; Represents a dilated convolution with a kernel size of k and a dilation rate of e; represents the third preliminary feature map, BN represents normalization, Relu represents Relu activation function, and Concat represents channel concatenation.

3. The method for coarse positioning of aircraft surface marker points according to claim 1, characterized in that: The method of performing vertical edge-aware convolution processing and horizontal edge-aware convolution processing on the first preliminary feature map to obtain a fourth preliminary feature map and a fifth preliminary feature map respectively, and performing element-by-element addition of the fourth preliminary feature map and the fifth preliminary feature map to obtain a sixth preliminary feature map includes: in, represents the Sobel operator, represents the first preliminary feature map, represents the fourth preliminary feature map, represents the fifth preliminary feature map, represents the sixth preliminary feature map, T represents transposition, Represents the convolution operation, and Sum represents element-by-element addition.

4. The method for coarse positioning of aircraft surface marker points according to claim 1, characterized in that: The method of fusing the third preliminary feature map and the sixth preliminary feature map and then performing convolution processing to obtain the first feature map includes: in, represents the first feature map, Indicates a convolution with a convolution kernel of 1×1. Indicates that the convolution kernel is 3×3 convolution, Concat indicates channel concatenation processing, represents the third preliminary feature map, Represents the sixth preliminary feature map.

5. The method for coarse positioning of aircraft surface marker points according to claim 1, characterized in that: The second feature map, the third feature map, the fourth feature map and the fifth feature map are respectively used as the input of the neck network, and are subjected to multi-scale feature fusion processing of the feature pyramid network, including: Performing feature fusion on the fifth feature map after upsampling the fifth feature map and the fourth feature map to obtain a first fused feature map; Performing residual processing on the first fused feature map, performing first-level multi-scale feature fusion processing on the second feature map and the third feature map to obtain a first multi-scale fused feature map; The first multi-scale fusion feature map is upsampled and convolved to obtain a corresponding first upsampled feature map and a first convolution feature map; The first upsampled feature map is subjected to feature fusion with the second feature map, and then residual processing is performed to obtain a second fused feature map; the first convolution feature map is subjected to feature fusion with the first fused feature map, and then residual processing is performed to obtain a third fused feature map; Perform a second-level multi-scale feature fusion process on the second fused feature map, the third fused feature map, and the first multi-scale fused feature map to obtain a second multi-scale fused feature map; The second multi-scale fusion feature map is upsampled and convolved to obtain a corresponding second upsampled feature map and a second convolution feature map; The first up-sampled feature map, the second up-sampled feature map and the second fused feature map are feature fused and then residual processed to obtain the sixth feature map; the first convolution feature map, the second convolution feature map and the third fused feature map are feature fused and then residual processed to obtain the seventh feature map.

6. The method for coarse positioning of aircraft surface marker points according to claim 5, characterized in that: The detection head includes a P2 layer detection head, a P3 layer detection head and a P4 layer detection head. The sixth feature map is used as an input of the P2 layer detection head, the seventh feature map is used as an input of the P4 layer detection head, and the second multi-scale fusion feature map is used as an input of the P3 layer detection head.

7. The method for coarse positioning of aircraft surface marker points according to claim 6, characterized in that: Multi-scale feature fusion processing at any level, including: Among the three inputs of the multi-scale feature fusion processing, the input corresponding to the P2 layer detection head is upsampled and then convolved to obtain a first multi-scale feature map, the input corresponding to the P3 layer detection head is convolved to obtain a second multi-scale feature map, and the input corresponding to the P4 layer detection head is downsampled to obtain a third multi-scale feature map; Performing feature fusion on the first multi-scale feature map, the second multi-scale feature map, and the third multi-scale feature map to obtain a first primary multi-scale fused feature map; Performing depthwise convolutions of three different convolution kernel sizes on the first primary multi-scale fusion feature map to obtain corresponding fourth multi-scale feature map, fifth multi-scale feature map and sixth multi-scale feature map; Add the fourth multi-scale feature map, the fifth multi-scale feature map, and the sixth multi-scale feature map element by element to obtain a second primary multi-scale fusion feature map; The second primary multi-scale fusion feature map is convolved point by point and then added element by element with the first primary multi-scale fusion feature map to complete the multi-scale feature fusion processing.

8. The method for coarse positioning of aircraft surface marker points according to claim 1, characterized in that: The detection method of the detection head includes: For any detection head, the input of the detection head is convolved to obtain the input feature map; The input feature map is subjected to at least one level of detail feature enhancement processing to obtain an output feature map; After classification and regression processing of the output feature map, the output is mapped to obtain the detection result; The detection results of all detection heads are mapped and processed to obtain the final detection result.

9. The method for coarse positioning of aircraft surface marker points according to claim 8, characterized in that: The detail feature enhancement processing includes: performing ordinary convolution processing, center difference convolution processing, angle difference convolution processing, vertical difference convolution processing and horizontal difference convolution processing on the input feature map, and then adding the five convolution processing results element by element to obtain a detail enhanced feature map.

10. A computer-readable storage medium, characterized in that: The medium stores a program, which can be loaded by a processor and execute the method for coarse positioning of aircraft surface marker points as claimed in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Social network image tampering positioning method based on multi-scale feature intelligent perception

    CN115063373A

  • Training method of image enhancement model, and image enhancement method and device

    CN118154445A

  • Method for detecting defects of copper strip

    CN118505671A

  • Steel surface defect detection method and device

    CN119399125A

Cited By

  • Rotation detection system and method based on multi-scale feature fusion

    CN121033678A

  • Rotation Detection System and Method Based on Multi-Scale Feature Fusion

    CN121033678B