Multi-attention-based airport runway and taxiway segmentation method

By introducing the airport runway and taxiway segmentation method with multi-attention module, the problems of low segmentation accuracy and poor multi-view angle adaptability in the prior art are solved, and more efficient and accurate runway and taxiway segmentation are achieved.

CN120411522APending Publication Date: 2025-08-01CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510557712.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing airport runway and taxiway segmentation methods are susceptible to noise and environmental factors in complex scenarios, resulting in low segmentation accuracy and difficult to adapt to accurate segmentation at multiple perspectives, especially the difficulty in distinguishing between taxiways and runways.

Method used

The multi-attention-based airport runway and taxiway segmentation method is adopted, through multiple downsampling and upsampling operations of the encoder and decoder, combined with the multi-attention module, including the three-branch large-core attention module, the spatial attention module and the channel attention module, the jump connection fusion feature is used to enhance the model's adaptability and segmentation accuracy at different perspectives.

Benefits of technology

It improves the segmentation accuracy and efficiency of airport runways and taxiways, reduces the missed segmentation of taxiways and runways, enhances the model's adaptability at multiple perspectives, and improves the accuracy and robustness of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411522A_ABST
    Figure CN120411522A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-attention-based airport runway and taxiway segmentation method, which belongs to the field of image processing and comprises the following steps of: inputting an airport image into an encoder of an airport segmentation model to carry out maximum pooling operation and down-sampling operation; in the down-sampling process, a multi-attention module is embedded to obtain an information feature map related to an airport runway and a taxiway; and inputting the obtained information feature map into a decoder of the airport segmentation model, reconstructing an image through up-sampling operation of an up-sampling layer, fusing features extracted by the decoder with features of the decoder by utilizing jump connection, and then outputting airport runway and taxiway segmentation pictures. According to the invention, the multi-attention module is introduced, so that the model pays more attention to the boundaries of the runway and the taxiway, wrong segmentation caused by easy confusion of the taxiway and the runway is reduced, and the overall segmentation precision and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method for segmenting airport runways and taxiways based on multi-attention. Background Art

[0002] In the aviation field, an airport is an important hub connecting air traffic and ground transportation; the traffic network of an airport mainly consists of runways and taxiways. A runway is a key area for the safe takeoff and landing of an aircraft, and a taxiway connects the runway with other airport facilities. By accurately segmenting airport runways and taxiways, the normal operation of ground traffic and aircraft takeoff and landing can be ensured, and the working efficiency of the airport can be effectively improved.

[0003] Early remote sensing airport runway segmentation methods can be roughly divided into two types: edge information extraction-based and region segmentation-based. The former algorithm has low complexity and fast calculation speed, and the parallel and straight characteristics of the airport runway are also conducive to the extraction of feature information; the latter focuses more on the significant structural features of the airport; however, these algorithms are difficult to adapt to more complex scenarios and are easily affected by factors such as noise and environment, resulting in low segmentation accuracy. At the same time, a large amount of manual intervention and parameter adjustment are required, and the segmentation efficiency is low.

[0004] The development of deep learning and remote sensing technology has greatly promoted the development of image segmentation technology. High-resolution remote sensing images can provide more abundant information and effectively improve the accuracy of image segmentation; the neural network model of deep learning has good feature learning ability. By training a large amount of data, it can not only adapt to various complex scenarios but also improve the operation efficiency and accuracy of image segmentation. To further improve the segmentation accuracy, some methods introduce an attention mechanism. By mimicking the selective attention ability of humans when processing information, the airport segmentation model can dynamically adjust the attention weight to highlight key information. However, in the segmentation task of airport runways and taxiways, there are still problems: 1. Existing research pays less attention to taxiways, and due to factors such as wear and tear of ground markings and blurred signs at the airport, it is difficult to clearly distinguish taxiways from runways. In addition, the geological environment interference around the airport and the complexity of the airport's own structure will also affect the segmentation results. 2. When an aircraft takes off and lands, there are differences between the observation perspectives of the pilots and the airport tower and the perspective in the remote sensing image, and these differences may cause inaccurate segmentation results of the model in complex scenarios. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for segmenting airport runways and taxiways based on multi-attention, which solves the deficiencies existing in the prior art.

[0006] The object of the present invention is achieved by the following technical solutions: A method for segmenting airport runways and taxiways based on multi-attention, the segmentation method comprising:

[0007] Step 1: Input the airport image into the encoder of the airport segmentation model for max-pooling operation and downsampling operation to expand the receptive field and enrich the extracted image features, and obtain an information feature map related to the airport runway and taxiway by embedding a multi-attention module during the downsampling process;

[0008] Step 2: Input the obtained information feature map into the decoder of the airport segmentation model, reconstruct the image through the upsampling operation of the upsampling layer, and at the same time use skip connections to fuse the features extracted by the decoder with the decoder features and then output the airport runway and taxiway segmentation pictures. During the upsampling process, embed a multi-attention module to locate and strengthen the features related to the runway and taxiway, and restore the detailed information of the airport image.

[0009] The specific content of Step 1 includes:

[0010] Extract the important feature map of the airport image through the max-pooling layer and then input it into the encoder to output the airport image feature map. The encoder includes 4 downsampling operations of the downsampling layer. During the 4 downsampling processes, 4, 8, 32, and 4 multi-attention modules are respectively embedded. Each downsampling is connected through the max-pooling layer, and the feature maps output during each downsampling process are respectively used as the input for the next downsampling operation and the corresponding upsampling operation. The output feature map of the downsampling layer is input into the upsampling layer through skip connections, and the airport feature map output by the last downsampling is directly used as the input for the decoder.

[0011] The specific content of Step 2 includes:

[0012] Input the airport feature map output by the encoder into the decoder. The decoder includes 4 upsampling operations of the upsampling layer. During the 4 upsampling processes, 4, 32, 8, and 4 multi-attention modules are respectively embedded. Map the feature map to the same number of channels as the number of classes through a 1×1 convolutional layer, and each channel corresponds to a class. Finally, output a segmentation map of the airport runway and taxiway, where Class is the number of classes. During each upsampling process, the input of the upsampling is directly spliced by the output feature map of the corresponding downsampling layer through skip connections and the output of the previous downsampling. The airport feature map output by the encoder is the input for the first upsampling of the decoder. Among them, W, H, and C are the width, height, and number of channels of the original image respectively.

[0013] The multi-attention module includes a three-branch large-kernel attention module, a spatial attention module, and a channel attention module;

[0014] The three-branch large kernel attention module is configured to perform feature extraction on the input airport feature map by expanding the receptive field;

[0015] The spatial attention module and the channel attention module are configured to extract key information of the airport runway and taxiway, so as to focus on the boundaries of the runway and taxiway and improve the generalization ability of the model under multiple perspectives.

[0016] The three-branch large kernel attention module simulates the working mechanism of a large convolution kernel through a combination of depth convolution and dilated convolution to achieve the purpose of expanding the receptive field;

[0017] The spatial attention module takes the feature maps extracted by two branches in the three-branch large kernel attention module as input, and realizes the dimensionality reduction operation through the strategy of flattening the features to reduce the computational amount. At the same time, it focuses on the boundaries between the airport taxiway and the runway to achieve precise segmentation;

[0018] The channel attention module dynamically adjusts the weights of the channels according to the correlation with the features of the airport runway and taxiway, and improves the adaptability of the airport segmentation model under different perspectives.

[0019] The encoder performs feature extraction on the input airport image through 4 times of downsampling and max pooling layers. Among them, the multi-attention module is used to complete the downsampling operation and extract the key features of the runway and taxiway areas;

[0020] The decoder reconstructs the feature image of the airport through 4 times of upsampling, and uses the multi-attention module to restore the spatial and detailed information of the runway and taxiway in the airport image;

[0021] The skip connection directly transmits the information to be input to the subsequent layers, enabling the information to jump between different layers. The skip connection is introduced to splice the airport feature maps extracted by each layer of the encoder with the airport feature maps of the corresponding layers of the decoder, retaining the low-level features from the encoder to the decoder, enriching the detailed features, and enhancing the perception ability of the airport segmentation model for detailed information such as the boundary area between the runway and the taxiway.

[0022] The present invention has the following advantages: A method for segmenting airport runways and taxiways based on multi-attention. The introduction of the multi-attention module enables the model to pay more attention to the boundaries of the runway and taxiway, not only reducing the wrong segmentation caused by the easy confusion between the taxiway and the runway, improving the ability to extract key features, but also enhancing the adaptability of the model under different perspectives, and improving the overall segmentation accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0024] Figure 2Schematic structural diagram of the encoder of the present invention;

[0025] Figure 3 Schematic structural diagram of the feature processing unit of the present invention;

[0026] Figure 4 Schematic structural diagram of the decoder of the present invention. Specific embodiments

[0027] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and illustrated herein can generally be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present application provided below with reference to the accompanying drawings is not intended to limit the protection scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application. The present invention will be further described below with reference to the accompanying drawings.

[0028] As Figure 1 shown, the present invention specifically relates to a method for segmenting airport runways and taxiways based on multi-attention, which specifically includes the following contents:

[0029] Step 1: Input the airport image into the encoder for a series of max-pooling operations and downsampling operations to expand the receptive field and extract rich image features. During the downsampling process, further obtain more detailed information features about the airport runway and taxiway by embedding a multi-attention module.

[0030] Step 2: Input the airport feature map obtained by the encoder into the decoder, and reconstruct the image through an upsampling operation symmetric to the downsampling; meanwhile, use skip connections to fuse the features extracted by the decoder with the decoder features, and finally output the segmented image of the airport runway and taxiway. During the upsampling process, locate and strengthen the features related to the runway and taxiway by embedding a multi-attention module, and restore the detailed information of the airport image.

[0031] Further, step 1 specifically includes:

[0032] After extracting the important feature map of the airport image through the max-pooling layer, input it into the encoder, and output the airport image feature map. The encoder includes 4 downsampling operations of the downsampling layer. During the 4 downsampling processes, 4, 8, 32, and 4 multi-attention modules are respectively embedded. Each downsampling is connected through the max-pooling layer, and the corresponding feature map sizes are , , and , and during each downsampling process, the output feature maps are respectively used as the inputs for the next downsampling operation and the corresponding upsampling operation. The output feature maps of the downsampling layer are input into the upsampling layer through skip connections. The airport feature map output by the last downsampling is directly used as the input to the decoder. Here, W, H, and C are the width, height, and number of channels of the original image.

[0033] Furthermore, step two specifically includes:

[0034] Input the airport feature map output by the encoder into the decoder. The decoder includes 4 upsampling operations of the upsampling layer. During the 4 upsampling processes, 4, 32, 8, and 4 multi-attention modules are respectively embedded. The feature map is mapped to the same number of channels as the number of classes through a 1×1 convolutional layer. Each channel corresponds to a class, and finally an segmentation map of the airport runway and taxiway is output. Class is the number of classes. During each upsampling process, the input of the upsampling is the output feature map of the corresponding downsampling layer directly concatenated with the output of the previous downsampling through skip connections. The airport feature map output by the encoder is the input for the first upsampling of the decoder.

[0035] Among them, as Figure 2 shown, the multi-attention module runs through the entire encoder, effectively extracts airport image features under multi-perspective conditions with the help of the spatial attention module and the channel attention module, and pays more attention to the important features for distinguishing the boundaries of the runway and taxiway. The skip connections input the feature information extracted by each downsampling in the encoder into the corresponding upsampling operation in the decoder for feature fusion, which is beneficial to improving the detail perception ability of the model, retaining the low-level features, and effectively improving the segmentation accuracy and model performance. In order to enable the airport segmentation model to better adapt to multi-perspective scenarios and clearly distinguish the boundaries of the runway and taxiway, a multi-attention module is introduced as the main structure of the encoder to obtain richer visual features.

[0036] Among them, the specific method of the encoder is as follows:

[0037] After inputting the remote sensing airport image, it will go through 4 max-pooling layers and downsampling operations to expand the receptive field and extract image features. During these 4 downsampling processes, 3, 3, 12, and 3 multi-attention modules are respectively embedded, and the feature information extracted by each downsampling will also be directly concatenated with the upsampling features of the corresponding decoder.

[0038] As Figure 3As shown in the figure, there is a feature processing unit that extracts rich visual features from the input airport feature image and uses a multi-attention module to enhance the attention to the airport runway and taxiway areas and their demarcation areas, reducing the computational burden and enhancing the feature extraction ability. The feature processing unit consists of two parts. The first part is composed of a 1×1 convolutional layer, an activation function layer Gelu, a multi-attention module, and a 1×1 convolutional layer, mainly aiming to efficiently extract important features by introducing the multi-attention module. Among them, the spatial attention module helps the model better distinguish the runway and taxiway, which is beneficial to improving the accuracy of the segmentation result; the channel attention module dynamically adjusts the channel weights, flexibly processes problems in multi-perspective scenarios, and pays more attention to the important features related to the runway and taxiway. The second part is composed of a 1×1 convolutional layer, an activation function layer Gelu, a multi-layer perceptron MLP, and a 1×1 convolutional layer, further extracting effective information from the airport feature image. Such a design enables the model to pay more attention to the distinction between the runway and taxiway in multi-perspective scenarios, improve the generalization performance of the model, and more flexibly handle various complex scenarios.

[0039] As Figure 4 shown, the decoder is also mainly composed of a multi-attention module. The spatial and channel attention modules can accurately locate and strengthen the features of the runway and taxiway areas, better reconstructing the airport image. At the same time, the skip connection fuses the low-level feature information extracted by downsampling with the feature map output by the decoder's upsampling, improving the robustness of the model. In order to more accurately restore the spatial and detailed information of the airport image and precisely segment the runway and taxiway, the multi-attention module mechanism is introduced to focus on important features and fuse them with the low-level feature information, improving the performance and segmentation accuracy of the model. The specific method of the decoder is as follows:

[0040] After the feature information extracted by the encoder is input into the decoder, it undergoes 4 upsampling operations, embedding 4, 32, 8, and 4 multi-attention modules respectively, and fusing them with the low-level feature information output by the encoder through skip connections; finally, a segmentation map of the airport runway and taxiway with multiple channels is output, and the number of channels is the number of categories.

[0041] Skip connection: The skip connection directly transmits the information to be input to the subsequent layers, allowing information to jump between different layers. The skip connection is introduced to splice the airport feature maps extracted by each layer of the encoder with the corresponding layer of the decoder's airport feature maps, retaining the low-level features from the encoder to the decoder, enriching the detailed features, and enhancing the perception ability of the segmentation model for detailed information such as the demarcation area between the runway and taxiway.

[0042] Furthermore, the multi-attention module consists of a spatial attention module (SAM) and a channel attention module (CAM), which solve the problems of distinguishing runways from taxiways and segmentation in multi-view scenarios respectively. In addition, the module also introduces a three-branch large kernel attention module, and the large kernel attention module (LKA) is its main component. By replacing the role of the large convolution kernel with depth convolution and dilated convolution, it not only expands the range of the receptive field to capture richer information, but also reduces the computational burden. The calculation formula of the large kernel attention module is as follows:

[0043] After the input airport feature image is processed by the large kernel attention module, it captures broader information. The airport feature maps extracted from two of the three-branch large kernel attention modules are transformed from three-dimensional to two-dimensional through a flattening strategy, and then enter the spatial attention module. This can significantly reduce the computational complexity and the number of parameters, and effectively capture the important information in the local features of the airport, solving the problem that runways and taxiways are easily confused. The channel attention module solves the segmentation problem in multi-view scenarios by dynamically adjusting the channel weights. For channels with a high degree of relevance to airport runways and taxiways, larger weights are assigned so that more attention can be given to these channels in subsequent work; for channels with a low degree of relevance to important features, the weights are reduced to minimize their impact on the model segmentation results.

[0044] Among them, the large kernel attention module is LKA(*), the spatial attention module is SAM(*), the channel attention module is CAM(*), and the input image feature is I. It can be expressed by the formula:

[0045] ,

[0046] Furthermore, the three-branch large kernel attention module: It consists of large kernel attention modules, mainly simulating the working mechanism of the large convolution kernel through the combination of depth convolution and dilated convolution, so as to achieve the effect of expanding the receptive field and extracting richer information. The large kernel attention module can be formulated as:

[0047] ,

[0048] Among them, * represents the product of elements, I represents the input feature map, represents point convolution, represents depth convolution, represents depthwise separable convolution.

[0049] The spatial attention module: Taking the feature maps extracted from two branches of the three-branch large kernel attention module as input, it realizes the dimensionality reduction operation through the strategy of flattening the features. While reducing the amount of calculation, it enables the module to more effectively focus on the boundaries between airport taxiways and runways, which is beneficial for accurate segmentation. This module can be formulated as:

[0050] ,

[0051] Among them, , , where W is the width of the original feature map, H is the height of the original feature map, C is the number of channels of the original feature map, is a learnable weight initialized to 0.

[0052] Channel attention module: It can dynamically adjust the weights of channels according to the correlation with the features of airport runways and taxiways, improving the adaptability of the airport segmentation model under different perspectives. This module can be formulated as:

[0053] ,

[0054] Where , W is the width of the original feature map, H is the height of the original feature map, C is the number of channels of the original feature map, is a learnable weight initialized to 0.

[0055] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and improvements, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and alterations made by those skilled in the art that do not depart from the spirit and scope of the present invention shall all be within the protection scope of the appended claims of the present invention.

Claims

1. A multi-attention-based method for airport runway and taxiway segmentation, characterized in that: The segmentation method includes: Step 1: Input the airport image into the encoder of the airport segmentation model for max pooling operation and downsampling operation to expand the receptive field and enrich the extracted image features, and obtain an information feature map related to the airport runway and taxiway through embedding a multi-attention module during the downsampling process; Step 2: Input the obtained information feature map into the decoder of the airport segmentation model, reconstruct the image through the upsampling operation of the upsampling layer, and at the same time use skip connections to fuse the features extracted by the decoder with the decoder features and then output the airport runway and taxiway segmentation pictures. During the upsampling process, the multi-attention module is embedded to locate and strengthen the features related to the runway and taxiway, and restore the detailed information of the airport image.

2. The method for segmenting airport runways and taxiways based on multi-attention according to claim 1, wherein: Specifically, Step 1 includes: After extracting the important feature map of the airport image through the max pooling layer, input it into the encoder, and output the airport image feature map. The encoder includes 4 downsampling operations of the downsampling layer. During the 4 downsampling processes, 4, 8, 32, and 4 multi-attention modules are respectively embedded. Each downsampling is connected through the max pooling layer, and the feature maps output during each downsampling process are respectively used as the input of the next downsampling operation and the corresponding upsampling operation. The output feature map of the downsampling layer is input into the upsampling layer through skip connections, and the airport feature map output by the last downsampling is directly used as the input of the decoder.

3. A method for segmenting airport runways and taxiways based on multi-attention according to claim 1, characterized in that: Specifically, Step 2 includes: The airport feature map output by the encoder is input into the decoder. The decoder includes four upsampling operations of the upsampling layer. Four, 32, eight, and four multi-attention modules are respectively embedded in the four upsampling processes. The feature map is mapped to the same number of channels as the number of classes through a 1×1 convolutional layer. Each channel corresponds to a class. Finally, a segmentation map of the airport runway and taxiway is output. Class is the number of classes. In each upsampling process, the input of the upsampling is directly spliced with the output of the previous downsampling through a skip connection of the output feature map of the corresponding downsampling layer. The airport feature map output by the encoder is the input of the first upsampling of the decoder. Among them, W, H, and C are the width, height, and number of channels of the original image respectively.

4. A method for segmenting airport runways and taxiways based on multi-attention according to claim 1, characterized in that: The multi-attention module includes a three-branch large kernel attention module, a spatial attention module, and a channel attention module; The three-branch large kernel attention module is configured to extract features from the input airport feature map by expanding the receptive field; The spatial attention module and the channel attention module are configured to extract the key information of the airport runway and taxiway, so as to focus on the boundaries of the runway and taxiway and improve the generalization ability of the model under multiple perspectives.

5. A method for segmenting airport runways and taxiways based on multi-attention, as claimed in claim 4, wherein: The three-branch large kernel attention module simulates the working mechanism of a large convolution kernel through a combination of depth convolution and dilated convolution to achieve the purpose of expanding the receptive field; The spatial attention module takes the feature maps extracted from two branches of the three-branch large kernel attention module as input, and realizes the dimensionality reduction operation through the strategy of flattening the features to reduce the computational amount, and at the same time focuses on the boundaries between the airport taxiway and the runway to achieve accurate segmentation; The channel attention module dynamically adjusts the weights of the channels according to the correlation with the features of the airport runway and taxiway, and improves the adaptability of the airport segmentation model under different perspectives.

6. The method for segmenting airport runways and taxiways based on multi-attention according to claim 1, characterized in that: The encoder extracts features from the input airport image through 4 downsamplings and the max pooling layer, and uses the multi-attention module to complete the downsampling operation and extract the key features of the runway and taxiway areas; The decoder reconstructs the feature image of the airport through 4 upsamplings, and uses the multi-attention module to restore the spatial and detailed information of the runway and taxiway in the airport image; The skip connection directly transmits the information to be input to subsequent layers, enabling the information to skip between different layers. Introducing the skip connection concatenates the airport feature maps extracted by each layer of the encoder with the airport feature maps of the corresponding layer of the decoder, retaining the low-level features from the encoder to the decoder, enriching the detailed features, and enhancing the perception ability of the airport segmentation model for detailed information such as the demarcation area between the runway and the taxiway.