Aircraft take-off and landing scene identification method

By building a semantic segmentation network, using the fusion technology of feature retention and multi-scale attention modules, the accuracy and real-time problems of aircraft take-off and landing scene recognition in the existing technology are solved, pixel-level prediction and recognition of multi-category targets are achieved, and the automation level of automatic flight systems is improved.

CN120374982APending Publication Date: 2025-07-25BEIJING AERONAUTIC SCI & TECH RES INST OF COMAC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510490726.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing aircraft take-off and landing scenario recognition methods are difficult to achieve fast, accurate and comprehensive segmentation in complex environments. The existing network architecture is simple and has poor real-time performance, so it is impossible to effectively analyze multiple categories of aircraft take-off and landing scenarios.

Method used

A semantic segmentation network architecture is built, and multi-layer feature fusion is carried out through feature preservation attention modules and multi-scale feature attention modules, combined with encoders and decoders, to achieve pixel-level prediction and recognition of multi-category targets.

Benefits of technology

It improves the accuracy and versatility of aircraft take-off and landing scene recognition, optimizes the inference speed, adapts to aircraft take-off and landing scene images of various sizes, and enhances the stability and robustness of the automatic flight system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374982A_ABST
    Figure CN120374982A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an aircraft take-off and landing scene recognition method, and the method comprises the steps: building an aircraft take-off and landing scene data set, and constructing a multi-class semantic segmentation data set through data annotation, data enhancement and other operations; constructing a semantic segmentation network architecture, and performing training and verification on the data set to obtain an optimal aircraft take-off and landing scene segmentation model; and performing quantization processing on the optimal aircraft take-off and landing scene segmentation model to obtain a to-be-deployed model. Compared with the prior art, the method has the advantages that the attention mechanism and the multi-scale fusion method are used, the detection precision of the takeoff and landing runway of the airplane is improved, the reasoning speed is optimized on the basis of ensuring the segmentation effect, and the method has more advantages in practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a method for recognizing aircraft takeoff and landing scenarios. Background Art

[0002] In the modern aviation field, the automatic flight system has become a key component in ensuring flight safety and improving efficiency. Especially during the takeoff and landing phases, the visual assistance module plays a crucial role, significantly enhancing the efficiency of the aircraft during taxiing, takeoff, and landing guidance. However, due to the environmental complexity during aircraft takeoff and landing (such as variable weather conditions, interference from runway reflected light, etc.), existing visual assistance systems face many challenges in accurately identifying runways. Currently, the methods with better effects mostly adopt object detection technology, and label the runway position as a rectangular box. However, this method cannot be accurate to the specific pixel-level details of the runway. Semantic segmentation can combine object detection and image classification for overall environmental perception. When attempting semantic segmentation technology, existing network architectures are relatively simple, and usually only perform single-class segmentation on the runway, making it difficult to comprehensively display the aircraft takeoff and landing scenarios. Therefore, there is an urgent need for a segmentation method that can quickly, accurately, and comprehensively analyze aircraft takeoff and landing scenarios in various complex environments, so as to enhance the performance of the visual assistance system in the automatic flight system.

[0003] Chinese patent document CN116597315A discloses a runway detection method for infrared remote sensing images based on YOLO v3-tiny. This document obtains an accurate runway contour line through a multi-stage method, that is, through methods such as object detection, line segment detection, sorting and fitting, etc., and finally obtains the runway contour line. This method of combining multiple stages to obtain the runway contour has low efficiency and poor real-time performance.

[0004] Chinese patent document CN113052106A discloses an aircraft takeoff and landing runway recognition method based on the PSPNet network. This document only detects the airport runway and fails to comprehensively analyze the runway image scene, which will be restricted to a certain extent in practical applications. In addition, the network structure of this document is mainly improved based on PSPNet, with poor performance in capturing local details and boundary information, and the real-time performance of image inference is poor due to the high model complexity. In addition, this document only detects the airport runway and fails to comprehensively analyze the runway image scene, which will be restricted to a certain extent in practical applications.

[0005] Chinese Patent Document CN116778315A discloses a method for extracting airport runways from remote sensing images based on joint learning of boundaries and context. This document is mainly an improvement based on HRNet. Although it enhances the retention of details, the designed network structure also significantly increases the computational complexity and memory requirements. In addition, this document is designed for specific remote sensing image tasks and may perform poorly in other types of tasks, such as runway image segmentation tasks, lacking generality. Summary of the Invention

[0006] An embodiment of this specification provides a method for recognizing aircraft takeoff and landing scenarios to solve the technical problem of how to improve the recognition effect of aircraft takeoff and landing scenarios.

[0007] To solve the above technical problems, the embodiments of this specification provide the following technical solutions:

[0008] An embodiment of this specification provides a method for recognizing aircraft takeoff and landing scenarios, and the method includes:

[0009] Establish an aircraft takeoff and landing scenario dataset, construct a semantic segmentation network architecture, train and verify the model performance, and perform quantization processing and deployment testing;

[0010] Among them, constructing a semantic segmentation network architecture includes:

[0011] Encode the image to be recognized to obtain an encoded feature map;

[0012] Use a feature retention attention module to determine a global feature map based on the encoded feature map;

[0013] Decode the global feature map to obtain a decoded feature map;

[0014] Use a multi-scale feature attention module to fuse the feature maps obtained by upsampling each layer during the decoding process with the global feature map, and splice them to obtain a fused feature map;

[0015] Add the decoded feature map and the fused feature map to finally obtain the recognition result of the aircraft takeoff and landing scenario.

[0016] Optionally, establishing an aircraft takeoff and landing scenario dataset includes:

[0017] Collect initial images, annotate the initial images to obtain label images corresponding to the initial images;

[0018] Perform data augmentation operations on the initial images and their corresponding labels to obtain a data-augmented dataset.

[0019] Perform data preprocessing operations on the data-augmented dataset to obtain a final multi-class semantic segmentation dataset.

[0020] Optionally, the semantic segmentation network architecture includes an encoder and a decoder;

[0021] The encoder includes multiple layers of residual modules and downsampling;

[0022] The decoder includes multiple layers of convolutional blocks and upsampling, and each layer is connected to the encoder through a skip link with convolution;

[0023] The semantic segmentation network further includes a feature retention attention module;

[0024] The semantic segmentation network further includes a multi-scale feature attention module.

[0025] Optionally, training and validating the model performance includes:

[0026] Dividing the multi-class semantic segmentation dataset according to a ratio to obtain a training set and a validation set, where the RGB images and the label images correspond one by one.

[0027] During the training process, a mixed loss function is used for training supervision to obtain multiple aircraft takeoff and landing scene segmentation models.

[0028] During the validation process, by calculating the segmentation metrics of each category of the validation set images, the optimal aircraft takeoff and landing scene segmentation model is obtained.

[0029] Optionally, quantization processing and deployment testing include:

[0030] Performing quantization processing with layer-by-layer calibration on the optimal aircraft takeoff and landing scene segmentation model obtained by validation to obtain a model to be deployed, and applying it to the target device for testing.

[0031] Optionally, encoding the image to be recognized to obtain an encoded feature map includes:

[0032] Passing the image to be recognized through multiple layers of residual modules and downsampling in sequence to obtain an encoded feature map.

[0033] Optionally, using the feature retention attention module to determine the global feature map according to the encoded feature map includes:

[0034] Performing max pooling operations with different pooling kernel sizes on the encoded feature map to obtain context feature maps of different scales;

[0035] Performing an upsampling operation on the context feature maps of different scales to restore them to the same scale as the encoded feature map;

[0036] Performing channel merging on the restored context feature map and the encoded feature map;

[0037] Split the feature map after channel merging by channels, and add the split result to the encoded feature map to obtain a global feature map.

[0038] Optionally, decode the global feature map to obtain a decoded feature map, including:

[0039] Decode the global feature map through multiple layers of convolutional blocks and upsampling. Among them, when performing convolutional blocks and upsampling for each layer, introduce the information of the same layer obtained during the encoding process into the feature map obtained by upsampling in a skip connection manner with convolution to obtain a decoded feature map.

[0040] Optionally, use a multi-scale feature attention module to perform a fusion operation on the feature map obtained by upsampling for each layer during the decoding process and the global feature map, and splice them to obtain a fused feature map, including:

[0041] Use a multi-scale feature attention module to adjust the number of channels of the global feature map through a convolutional layer, and perform upsampling using linear interpolation;

[0042] Take the result obtained after upsampling as a weight, perform dot product operations with the feature maps obtained by upsampling for each layer during the decoding process respectively, and then perform upsampling operations and convolutional layer adjustments to obtain the fused feature map for each layer.

[0043] Splice the results of fusing each layer with the global feature map to obtain the final fused feature map.

[0044] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:

[0045] The above technical solution can realize pixel-level prediction and recognition of multi-category targets in the aircraft takeoff and landing scenarios, and can achieve better aircraft takeoff and landing scenario recognition effects compared with existing models. Moreover, it can predict and recognize aircraft takeoff and landing scenario images of various different sizes, with stronger versatility. In addition, this method optimizes the inference speed on the basis of ensuring segmentation accuracy and has more advantages in practical applications. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments of this specification or the prior art. Obviously, the following only shows the drawings required for a part of the embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 It is the flowchart of model training, verification and testing in the first embodiment of this specification.

[0048] Figure 2 It is the semantic segmentation network architecture diagram in the first embodiment of this specification.

[0049] Figure 3 It is the network structure diagram of the feature retention attention module in the first embodiment of this specification.

[0050] Figure 4 It is the quantitative comparison table of the image test results in the first embodiment of this specification.

[0051] Figure 5 It is the visualization effect diagram of the image test results in the first embodiment of this specification.

[0052] Figure 6 It is the visualization effect diagram of the video test results in the first embodiment of this specification. Specific implementation manners

[0053] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments involved in the specific implementation manners are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the specific implementation manners without creative efforts shall fall within the protection scope of this application.

[0054] The first embodiment of this specification (hereinafter referred to as "Embodiment 1") provides an aircraft takeoff and landing scene recognition method. The execution subject of Embodiment 1 includes but is not limited to an airborne terminal, a server, an operating system, or an application program. That is, the execution subject can be various and can be set, used, or transformed according to needs. In addition, a third-party application program can assist the execution subject in executing Embodiment 1. For example, the method provided in Embodiment 1 can be executed by a server, and a corresponding application program can be installed on a terminal (the terminal can be held by a user). Data can be transmitted between the terminal or the application program and the server, so as to assist the server in executing the method provided in Embodiment 1.

[0055] The aircraft takeoff and landing scene recognition method provided in Embodiment 1 includes:

[0056] Establishing an aircraft takeoff and landing scene data set, constructing a semantic segmentation network architecture, training and validating the model performance, and performing quantization processing and deployment testing.

[0057] Among them, constructing the semantic segmentation network architecture includes:

[0058] Encoding the image to be recognized to obtain an encoded feature map;

[0059] Use the feature-preserving attention module to determine the global feature map based on the encoded feature map;

[0060] Decode the global feature map to obtain the decoded feature map;

[0061] Use the multi-scale feature attention module to perform a fusion operation on the feature maps obtained by upsampling each layer during the decoding process and the global feature map, and splice them to obtain the fused feature map;

[0062] Add the decoded feature map and the fused feature map to finally obtain the recognition result of the aircraft takeoff and landing scenario.

[0063] In the first embodiment, establishing an aircraft takeoff and landing scenario dataset may include:

[0064] Collect the initial images, annotate the initial images to obtain the label images corresponding to the initial images;

[0065] Perform data augmentation operations on the initial images and their corresponding labels to obtain the dataset after data augmentation.

[0066] Perform data preprocessing operations on the dataset after data augmentation to obtain the final multi-class semantic segmentation dataset.

[0067] The following further illustrates how to obtain the multi-class semantic segmentation dataset:

[0068] For example, the camera can be installed at the tail of the aircraft, and images can be taken from the forward view perspective. The taken images are used as the initial images. The content included in the initial images can be selected according to needs. For example, the initial images can include part of the airport runway. For any initial image, the runway, sky, fuselage, land, etc. in the initial image are segmented and annotated, and the annotated image is the label image corresponding to the initial image.

[0069] Data augmentation operations can be performed on the initial images of the training set and the validation set and their corresponding label images. The data augmentation operations include but are not limited to random rotation, random flipping, random cropping, random scaling, and random color gamut transformation.

[0070] Next, preprocess each image after the data augmentation operation, which may include one or more operations such as normalization, channel transformation, and equal-scale scaling and padding. The purpose is to ensure that images of different sizes can be accurately processed subsequently (such as model training) without distortion. Specifically, for any image, normalization includes adjusting the range of pixel values, for example, scaling the pixel values to [0, 1] or [-1, 1] to meet the model calculation requirements. Channel transformation includes adjusting the channel order or color channel format of the image to match the requirements of the model framework. Equal-scale scaling and padding include: first, backing up the original image (i.e., the image before preprocessing) and calculating its size, then adjusting the image size by adding gray padding areas to maintain the original ratio, then normalizing the image and adding the batch_size dimension, and then performing a transpose operation to adapt to the required input format, for example, generating an image with dimensions of 3×H×W. Where H represents the height and W represents the width, and both can be set according to needs. This image preprocessing method can ensure that images of aircraft takeoff and landing scenes of different sizes can all be applicable to this aircraft takeoff and landing scene recognition method, making it more general.

[0071] In the first embodiment, training and validating the model performance may include: First, divide the multi-class semantic segmentation dataset according to a ratio to obtain a training set and a validation set, where the RGB images and the label images correspond one by one. Then, use the dataset to train in the above-mentioned semantic segmentation network. The training process uses the Adam optimizer for optimization learning, adjusts the learning rate according to the cosine annealing decay strategy, and uses a mixed loss function composed of the cross-entropy loss function and the focal loss function for training supervision to obtain multiple aircraft takeoff and landing scene segmentation models. During the validation process, the optimal aircraft takeoff and landing scene segmentation model is obtained by calculating the segmentation metrics of each category of the validation set images.

[0072] The following further illustrates how to train and validate the model performance to obtain the optimal aircraft takeoff and landing scene segmentation model:

[0073] Reference Figure 1 , a semantic segmentation network model can be pre-constructed in advance, and the semantic segmentation network model is trained using training samples and the trained model is validated using validation samples. Reference Figure 2 , the semantic segmentation network architecture includes an encoder and a decoder; the encoder includes multiple layers of residual modules and downsampling; the decoder includes multiple layers of convolutional blocks and upsampling, and each layer is connected to the encoder through a skip link with a convolutional block; the semantic segmentation network also includes a feature retention attention module; the semantic segmentation network also includes a multi-scale feature attention module.

[0074] During the training process, the Adam optimizer is used to optimize and learn the semantic segmentation network model. Among them, the initial learning rate can be set to 2×10-4 (For example only), the learning rate can be adjusted according to the cosine annealing decay strategy. For example, the learning rate is set to 0.5. And, the cross-entropy loss function L CE and the focal loss function L focal are combined to form a hybrid loss function L MixedLoss for training and supervision of the semantic segmentation network model to solve the problem of uneven pixel ratio of various categories in the aircraft takeoff and landing scenario, and at the same time perform constraints at the pixel level and the image level to enhance the model learning effect.

[0075] The formula of the above loss function is as follows:

[0076] L MixedLoss = L CE + L focal

[0077]

[0078] where N is the number of pixels to be predicted, M is the number of categories, y ic is a sign function that takes values of 0 or 1, where 1 means the predicted class of sample i is equal to c, 0 means the predicted class of sample i is not equal to c, log(.) is the natural logarithm with base e, and p ic is the predicted probability that the predicted sample i belongs to class c. γ is the focal parameter.

[0079] After training, multiple trained models are obtained. Then, each trained model is verified on the validation set of the above-mentioned population to determine the optimal model.

[0080] Among them, the verification metrics can be the intersection over union (IoU) and the pixel accuracy (PA). The intersection over union is the main reference metric, and the pixel accuracy is the secondary reference metric. The specific formulas of the verification metrics are as follows:

[0081]

[0082] where true positive (TP) represents the number of pixels correctly predicted to belong to a given class, true negative (TN) represents the number of pixels correctly identified as not belonging to a given class, false positive (FP) represents the number of pixels incorrectly predicted to belong to a given class, false negative (FN) represents the number of pixels incorrectly identified as not belonging to a given class, and k represents the number of segmentation categories.

[0083] The larger the average intersection over union of each category and the higher the average pixel accuracy, the better the model's performance for the aircraft takeoff and landing scenario where the dataset is collected. Therefore, the model with the largest calculated average intersection over union and the highest average pixel accuracy is the optimal model for the aircraft takeoff and landing scenario where the dataset is located.

[0084] In the first embodiment, the quantization process and deployment test may include:

[0085] Collect correction data to provide a reference for post-training quantization;

[0086] Use the post-training quantization method after quantization-aware training, and adopt a layer-by-layer correction method to determine the quantization parameters by calculating the range of activation values of each layer, and obtain the finally quantized aircraft takeoff and landing scene segmentation model.

[0087] Deploy the above-mentioned quantized aircraft takeoff and landing scene segmentation model to the target device for on-site testing.

[0088] In practical applications, first, by calculating and selecting appropriate quantization ranges (such as minimum and maximum values), ensure that the performance of the model after quantization is as close as possible to the original model. Second, by collecting example data (which can be training data or correction data sets) and obtaining statistical information through forward propagation, so that appropriate quantization parameters can be obtained for each layer during the quantization process. Then, through the layer-by-layer correction quantization method, obtain the quantized aircraft takeoff and landing scene segmentation model. Finally, deploy the quantized aircraft takeoff and landing scene segmentation model to execution entities such as airborne terminals, servers, operating systems, or applications for on-site testing.

[0089] In the first embodiment, encoding the to-be-recognized image to obtain an encoded feature map may include:

[0090] Pass the to-be-recognized image through multiple layers of residual modules and downsampling in sequence to obtain an encoded feature map.

[0091] In the first embodiment, using the feature retention attention module to determine the global feature map based on the encoded feature map may include:

[0092] Perform max-pooling operations with different pooling kernel sizes on the encoded feature map to obtain context feature maps of different scales; perform upsampling operations on the context feature maps of different scales to restore them to the same scale as the encoded feature map; perform channel merging on the restored context feature map and the encoded feature map; split the feature map after channel merging by channels, and add the split results to the encoded feature map to obtain the global feature map.

[0093] In the first embodiment, decoding the global feature map to obtain a decoded feature map includes:

[0094] Decode the global feature map through multiple layers of convolutional blocks and upsampling. Among them, when performing convolutional blocks and upsampling for each layer, introduce the information of the same layer obtained during the encoding process into the feature map obtained by upsampling in a skip connection manner with convolution to obtain the decoded feature map.

[0095] In the first embodiment, the multi-scale feature attention module is used to fuse the feature maps obtained by upsampling each layer in the decoding process with the global feature map, and after splicing, a fused feature map is obtained, including:

[0096] The multi-scale feature attention module adjusts the number of channels of the global feature map through a convolutional layer and performs upsampling using linear interpolation; the result obtained after upsampling is used as a weight, and is respectively dot-multiplied with the feature maps obtained by upsampling each layer in the decoding process, and then through an upsampling operation and convolutional layer adjustment, the fused feature map of each layer is obtained; the results of fusing each layer with the global feature map are spliced to obtain the final fused feature map.

[0097] In practical applications, images of the aircraft takeoff and landing scene can be collected by installing image acquisition devices such as cameras on the aircraft, or images of the aircraft takeoff and landing scene can be obtained through other means. The images of the aircraft takeoff and landing scene can be backed up first, and the height and width of the images of the aircraft takeoff and landing scene are calculated, and then in cases where it is necessary (such as cases where it is necessary to adjust the size or resolution of the images of the aircraft takeoff and landing scene), data augmentation operations and preprocessing are performed on the images of the aircraft takeoff and landing scene, including: performing a distortion-free resize on the images of the aircraft takeoff and landing scene, that is, adjusting the size of the images of the aircraft takeoff and landing scene by adding gray bars, and normalizing the resized images of the aircraft takeoff and landing scene, dividing all pixel points by 255, adding a batch_size dimension, and then performing a transpose operation. The preprocessed images can be used as the images to be recognized for aircraft takeoff and landing scene recognition. If it is not necessary to process the images of the aircraft takeoff and landing scene, the images of the aircraft takeoff and landing scene can be used as the images to be recognized for aircraft takeoff and landing scene recognition.

[0098] The content of the encoding operation is further described below:

[0099] After inputting the image to be recognized into the semantic segmentation network, the aircraft takeoff and landing scene recognition model inputs the image to be recognized into the encoder, and extracts feature maps of different resolutions through multiple layers of residual modules and downsampling.

[0100] As an example, there can be four layers of residual modules and downsampling. The image to be recognized passes through these four layers of residual modules and downsampling in sequence, and the output (the output is a feature map) of the previous layer of residual module and downsampling is used as the input of the next layer of residual module and downsampling. After passing through each layer of residual module and downsampling, the height and width of the image are halved, thereby obtaining four feature maps of different resolutions. That is to say, after the image to be recognized 3×H×W passes through the first layer of residual module and downsampling, the output is feature map (denoted as FM1); The feature map of The feature map (denoted as FM2); After the feature map passes through the third residual module and downsampling, the output is The feature map (denoted as FM3); After the feature map passes through the fourth residual module and downsampling, the output is The feature map (denoted as FM4).

[0101] Through the above operations, the image to be recognized is encoded, and feature maps with different resolutions are obtained.

[0102] Next, it is further explained how to obtain the global feature map:

[0103] The encoded feature map (i.e., the feature map output by the last residual module and downsampling, Figure 2 in which is Feature Map4, i.e., FM4) is input into the Feature Preservation Attention Module (FPAM), and FPAM adaptively extracts the multi-scale context features of the encoded feature map. Specifically, referring to Figure 3 , FPAM first obtains the context feature maps of different scales of the encoded feature map through max pooling operations ( Figure 3 two context feature maps are obtained in it. To obtain multiple context feature maps, different pooling kernels can be used), and then restores the feature maps of different scales to the same scale as the encoded feature map through the upsampling operation of bilinear interpolation, so as to fully retain the different context feature information while fusing with the original feature map (i.e., FM4).

[0104] Then, the channel merging of the context feature map and the original feature map is realized through a concatenation operation. For example Figure 3 two context feature maps of different scales are obtained in it, and both of these two context feature maps are restored to the same scale as the encoded feature map, then the two restored context feature maps are merged with the encoded feature map in channels, that is, a concat operation is performed on the two restored context feature maps and the encoded feature map.

[0105] The feature map obtained after channel merging is successively passed through a 3×3 convolutional layer, a ReLU activation layer, a 1×1 convolutional layer, and a Sigmoid activation layer to generate corresponding spatial weights for each context feature map and the encoded feature map, and the spatial weights are used to obtain local feature maps in subsequent decoding operations.

[0106] The feature map obtained by merging channels through the split operation is split by channel, and three feature maps with the same scale are split out. The three split feature maps are added to the encoded feature map to obtain a global feature map (Global Feature Map, GFM).

[0107] As above, the global feature map is generated by FPAM. In the above operations, FPAM aggregates the global information of the scene of the image to be recognized into each pixel of the global feature map, so that the global feature map has rich multi-scale context information.

[0108] The following further explains how to obtain the decoded feature map:

[0109] Perform a decoding operation on the above global feature map to restore the image resolution of the global feature map to the resolution H×W of the image to be recognized.

[0110] Specifically, the decoding operation may include:

[0111] The aircraft takeoff and landing scene recognition model includes a decoder, and the decoder can include multiple layers, each layer being a convolutional block and an upsampling. The global feature map is input into the decoder, and the global feature map is decoded through multiple layers of convolutional blocks and upsampling to restore it to the original resolution (i.e., the resolution of the image to be recognized). Corresponding to the encoder, the decoder includes four layers of convolutional blocks and upsampling. The global feature map passes through these four layers of convolutional blocks and upsampling in sequence, and the output of the previous layer of convolutional block and upsampling is used as the input of the next layer of convolutional block and upsampling. When passing through the convolutional block and upsampling, the local features are effectively extracted through the aforementioned spatial weights, so as to retain the spatial information of the data and enhance the attention of the aircraft takeoff and landing scene recognition model to the key areas.

[0112] Specifically, when performing the convolutional block and upsampling for each layer, the information of the same layer obtained in the encoding operation is added to the feature maps obtained from the convolutional block and upsampling of that layer in the form of a skip connection with convolution. That is to say, in the first convolutional block and upsampling, the information obtained from the first residual module and downsampling in the aforementioned encoding operation is introduced into the feature maps obtained from the first convolutional block and upsampling through the concat operation to obtain the feature map FM5; in the second convolutional block and upsampling, the information obtained from the second residual module and downsampling in the aforementioned encoding operation is introduced into the feature maps obtained from the second convolutional block and upsampling through the concat operation to obtain the feature map FM6; in the third convolutional block and upsampling, the information obtained from the third residual module and downsampling in the aforementioned encoding operation is introduced into the feature maps obtained from the third convolutional block and upsampling through the concat operation to obtain the feature map FM7; in the fourth convolutional block and upsampling, the information obtained from the fourth residual module and downsampling in the aforementioned encoding operation is introduced into the feature maps obtained from the fourth convolutional block and upsampling through the concat operation to obtain the feature map FM8.

[0113] Since the feature maps generated by the shallower layers (closer to the input layer) of the encoder have a larger spatial size and retain more spatial detail information, through the above decoding operation, the high-resolution feature maps obtained by the encoder are fused with the upsampled feature maps obtained by the decoder through skip connections, realizing the addition of the information of the same layer in the shallower layers of the encoder to the corresponding layers of the decoder, that is, directly passing the high-resolution feature maps of the shallower layers in the encoder to the corresponding decoder part, so that more detail information can be retained, which is beneficial to improving the segmentation accuracy of the aircraft takeoff and landing scene model.

[0114] The feature map output by the last convolutional block and upsampling serves as the decoded feature map.

[0115] Next, it further explains how to perform the fusion operation (or multi-scale fusion operation):

[0116] The above global feature map is input into the Multi-scale Feature Attention (MFA) module. The number of channels of the global feature map is adjusted through a 1×1 convolutional layer, and then the global feature map is upsampled using bilinear interpolation. Immediately afterwards, its output value (i.e., the result obtained after upsampling) is used as a weight, and is respectively dot-multiplied with the multi-level features (the multi-level features are the feature maps FM5’, FM6’, FM7’ and FM8’ obtained after each convolutional block and upsampling operation in the aforementioned decoding operation, and can also be called multi-level local features, which have a higher resolution and richer local information compared to the global feature map), resulting in four feature maps. These four feature maps are respectively subjected to upsampling processing and 1×1 convolutional operations to obtain four fused feature maps, that is, the output feature maps of the four multi-scale feature attention modules. These four fused feature maps are then subjected to a concat channel merging operation and a convolutional block operation to form the final fused feature map. The convolutional block in Embodiment 1 may include a 3×3 convolutional layer, a ReLU activation layer, a 1×1 convolutional layer, and a Sigmoid activation layer.

[0117] Through the fusion operation, feature extraction of the feature maps obtained by upsampling each layer in the decoding process is realized. Through the dot-multiplication operation with the feature maps of each layer obtained by decoding, high-level semantic information is implicitly embedded into the multi-level local features, realizing the layer-by-layer embedding of global information to better adapt to the semantic segmentation task and providing information for the detail restoration of the subsequent segmentation inference result (i.e., the segmentation feature map).

[0118] In Embodiment 1, after obtaining the decoding feature map and the fused feature map, they are added together to obtain the recognition result. Then, a permute operation is performed on the recognition result to move the number of channels to the last dimension, and then a softmax operation is applied to extract the category with the maximum probability value corresponding to each pixel point, obtaining the final aircraft takeoff and landing scene recognition result (if a gray bar is added to the input image before inputting it into the aircraft takeoff and landing scene recognition model, the added gray bar part is intercepted from the final prediction result). By judging the category of each pixel point, a specific color is assigned to it to form a segmented image. Finally, the generated segmented image is converted into an image format and its size is adjusted to be the same as the original image (i.e., the aforementioned aircraft takeoff and landing scene image), and finally the segmentation inference result of the aircraft takeoff and landing scene image is obtained. Such an operation finally obtains a segmentation result with the same size as the original image. This process ensures the consistency and accuracy of the image in training and inference.

[0119] In addition, in Embodiment 1, verifying the model performance, quantization processing, and deployment testing include demonstrating the advancement and effectiveness of the optimal aircraft takeoff and landing scene recognition model from both quantitative and qualitative perspectives.

[0120] 1. Quantitative test results: We used the optimal aircraft take-off and landing scene recognition model to test a batch of aircraft take-off and landing scene images, and obtained the intersection-over-union ratio and pixel accuracy indicators of each category of the model on the test image set. The relevant results are as follows: Figure 4 As shown, the model has good expressiveness. In particular, it performs particularly well in pixel-level prediction of large areas such as the background, sky, aircraft body, and aircraft runway, all of which have reached more than 90%. In order to prove the advanced nature of this network, its segmentation performance is compared with that of the classic UNet network on the same batch of test data. It can be seen from the results of both indicators that the test results of the segmentation model have better segmentation performance than the UNet model. In addition, the optimal aircraft take-off and landing scene recognition model before and after quantization was tested on this batch of data sets. The average inference speed of each picture before and after quantization was 22.71Hz and 35.87Hz, which greatly improved the speed of field testing of aircraft take-off and landing scenes. Moreover, the intersection-over-union ratio after quantization is reduced by less than 2% compared with that before quantization (see Figure 4 ); Compared with the pixel accuracy before quantization, the decrease range is within 1.5% (see Figure 4 ), which can fully meet the application needs of actual scenarios.

[0121] 2. Qualitative test results: The optimal aircraft take-off and landing scene recognition model was used to test images and videos of aircraft take-off and landing scenes of different scales. The visual segmentation results are as follows: Figure 5 and Figure 6 shown. Figure 5 The UNet network model and the optimal aircraft take-off and landing scene recognition model are shown in Figure 5 By comparing the visualized segmentation results of the three models, it can be seen that the aircraft take-off and landing scene recognition model provided in Example 1 performs excellently in predicting the global information of the aircraft take-off and landing scene. In addition, by comparing the local details circled in orange in the visualized segmentation results of the three models, it can be concluded that the aircraft take-off and landing scene recognition model provided in Example 1 also performs very well in distinguishing the boundary information between different categories of targets and the local details of the image. Figure 6 The effect of video segmentation testing after quantization of the optimal aircraft take-off and landing scene recognition model is demonstrated, which can fully meet the needs of actual applications.

[0122] Embodiment 1 can achieve the following beneficial effects:

[0123] Compared with the traditional UNet network structure, the aircraft take-off and landing scene recognition model in Example 1 not only retains the encoder-decoder structure framework and jump connections, but also introduces structures such as residual modules and convolutions (such as cov1*1), which can effectively improve the aircraft take-off and landing scene recognition model's ability to reason about details in flight semantic scenes.

[0124] In the first embodiment, during the process of aircraft takeoff and landing scene recognition, the shallow feature map obtained by the encoding operation has a high resolution and contains rich local detail information. The deep feature map obtained by the decoding operation has a low resolution and contains high-level semantic information. By means of the designed multi-scale feature attention module, the high-level semantic information (from the deep feature map) is incorporated into the local features (from the shallow feature map), and the shallow feature map and the deep feature map are fused layer by layer, that is, multi-scale feature fusion, which can enhance the semantic expression ability of local features and provide important information support for the detail restoration of the segmentation feature map, so as to achieve better aircraft takeoff and landing scene recognition results.

[0125] The aircraft takeoff and landing scene recognition model provided by the first embodiment designs a feature retention attention module, which adaptively extracts multi-scale context information and effectively aggregates the global information of the scene to each pixel, effectively improving the aircraft takeoff and landing scene recognition effect.

[0126] The aircraft takeoff and landing scene recognition model provided by the first embodiment uses an end-to-end deep learning model and performs quantization operations, and the model inference speed is relatively fast and the real-time performance is high.

[0127] The aircraft takeoff and landing scene recognition method provided by the first embodiment can realize pixel-level prediction and recognition of multi-category targets (including but not limited to background, sky, airframe, aircraft runway and land) in the aircraft takeoff and landing scene, and can effectively improve the aircraft takeoff and landing scene recognition effect. Moreover, it can predict and recognize aircraft takeoff and landing scene images of various different sizes, and has stronger versatility.

[0128] The aircraft takeoff and landing scene recognition method provided by the first embodiment designs and introduces attention mechanisms and other means to improve the feature capture ability of the aircraft takeoff and landing scene, and realizes multi-category and pixel-level accurate recognition of aircraft takeoff and landing scene images. Especially in the taxiing stage and the approach and landing stage of flight, the aircraft takeoff and landing scene recognition model provided by the first embodiment shows excellent performance, and shows good stability and robustness in practical applications. Moreover, the aircraft takeoff and landing scene recognition method provided by the first embodiment can be used as a key link of the visual assistance module, improving the automation level and reliability of the automatic flight system, and providing reliable technical support for the visual assistance function of the automatic flight system.

[0129] The above is only for the embodiments of this specification and is not intended to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. An aircraft takeoff and landing scene recognition method, characterized in that, The method includes: establishing an aircraft takeoff and landing scene dataset, constructing a semantic segmentation network architecture, training and validating the model performance, and performing quantization processing and deployment testing; Among them, constructing the semantic segmentation network architecture includes: encoding the image to be recognized to obtain an encoded feature map; using a feature-preserving attention module to determine a global feature map based on the encoded feature map; decoding the global feature map to obtain a decoded feature map; using a multi-scale feature attention module to perform a fusion operation on the feature maps obtained by upsampling each layer during the decoding process and the global feature map, and splicing to obtain a fused feature map; adding the decoded feature map and the fused feature map to finally obtain the recognition result of the aircraft takeoff and landing scene.

2. The method according to claim 1, wherein Establishing the aircraft takeoff and landing scene dataset includes: collecting initial images, annotating the initial images to obtain label images corresponding to the initial images; performing data augmentation operations on the initial images and their corresponding labels to obtain a data-augmented dataset. performing data preprocessing operations on the data-augmented dataset to obtain a final multi-class semantic segmentation dataset.

3. The method according to claim 1, characterized in that, The semantic segmentation network architecture includes an encoder and a decoder; The encoder includes multiple layers of residual modules and downsampling; The decoder includes multiple layers of convolutional blocks and upsampling, and each layer is connected to the encoder through a skip link with a convolution; The semantic segmentation network also includes a feature-preserving attention module; The semantic segmentation network also includes a multi-scale feature attention module.

4. The method according to claim 1, wherein Training and validating the model performance includes: dividing the multi-class semantic segmentation dataset according to a ratio to obtain a training set and a validation set, where the RGB images and the label images correspond one by one. During the training process, using a mixed loss function for training supervision to obtain multiple aircraft takeoff and landing scene segmentation models. During the validation process, by calculating the segmentation metrics of each category of the validation set images, obtaining the optimal aircraft takeoff and landing scene segmentation model.

5. The method according to claim 1, characterized in that, Quantization processing and deployment testing include: performing quantization processing with layer-by-layer calibration on the optimal aircraft takeoff and landing scene segmentation model obtained by validation to obtain a model to be deployed, and applying it to the target device for testing.

6. The method according to any one of claims 1 to 5, characterized in that, Encoding the image to be recognized to obtain an encoded feature map includes: sequentially passing the image to be recognized through multiple layers of residual modules and downsampling to obtain an encoded feature map.

7. The method according to any one of claims 1 to 5, characterized in that Using a feature-preserving attention module to determine a global feature map based on the encoded feature map includes: performing max pooling operations with different pooling kernel sizes on the encoded feature map to obtain context feature maps of different scales; performing an upsampling operation on the context feature maps of different scales to restore them to the same scale as the encoded feature map; performing channel merging on the restored context feature map and the encoded feature map; splitting the feature map after channel merging by channels, and adding the splitting results to the encoded feature map to obtain a global feature map.

8. The method according to any one of claims 1 to 8, characterized in that, Decoding the global feature map to obtain a decoded feature map includes: Decode the global feature map through multiple convolutional blocks and upsampling. Among them, when performing convolutional blocks and upsampling for each layer, introduce the information of the same layer obtained during the encoding process into the feature map obtained by upsampling in a skip connection manner with convolution to obtain a decoded feature map.

9. The method according to any one of claims 1 to 5, characterized in that, Use a multi-scale feature attention module to perform a fusion operation on the feature map obtained by upsampling for each layer during the decoding process and the global feature map, and obtain a fused feature map after splicing, including: Use a multi-scale feature attention module to adjust the number of channels of the global feature map through a convolutional layer and perform upsampling using linear interpolation; Use the result obtained after upsampling as a weight, perform a dot product operation with the feature maps obtained by upsampling for each layer during the decoding process, and then perform an upsampling operation and convolutional layer adjustment to obtain the fused feature map for each layer. Splice the results after fusing each layer with the global feature map to obtain the final fused feature map.

Citation Information

Patent Citations

  • Aircraft take-off and landing runway identification method based on PSPNet network

    CN113052106A

  • Infrared remote sensing image runway detection method based on YOLO v3-tiny

    CN116597315A

  • Remote sensing image airport runway extraction method based on boundary and context joint learning

    CN116778315A