Multi-scale cross-space learning general aviation aircraft landing runway detection method

By building an AES-YOLOv8 network model for multi-scale cross-space learning, the high cost and insufficient detection accuracy problems of general aviation aircraft landing navigation systems were solved, efficient and safe runway detection was achieved, and the operational risks of pilots were reduced.

CN120689781APending Publication Date: 2025-09-23SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510775408.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing general aviation aircraft landing navigation system has high deployment costs, and the visual navigation system has deficiencies in runway detection accuracy and safety, making it difficult to effectively assist pilots in achieving automatic landing.

Method used

The AES-YOLOv8 network model with multi-scale cross-space learning is adopted. By constructing a lightweight channel attention convolution module LA-AKConv, introducing the EMA attention mechanism and the C2f_SCConv module, and combining it with the improved loss function S_CIOU, a runway detection model is trained to achieve multi-scale target detection on the runway.

Benefits of technology

It improves the accuracy of runway detection, reduces the risk of pilot landing operations, reduces system deployment costs, and enhances flight safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689781A_ABST
    Figure CN120689781A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scale cross-space learning general aviation aircraft landing runway detection method, and relates to the technical field of computer vision. The method comprises the following steps: firstly, collecting a runway image data set when a general aviation aircraft lands, preprocessing the data set, and dividing the data set into a training set and a test set; then, an AES-YOLOv8 network model for detecting a multi-scale runway target is constructed; the model is based on a YOLOv8s network of multi-scale cross-space learning, and a Backbone network comprises a plurality of improved lightweight channel attention convolution modules LA-AKConv; an EMA attention mechanism is introduced after a spatial pyramid pooling module SPPF of the Backbone network; the Neck network comprises a plurality of C2fSCConv modules fused by C2f and a space and channel reconstruction convolution module SCConv; and the AES-YOLOv8 network model is trained based on the preprocessed data set to obtain a multi-scale cross-space learning airport runway detection model, and the runway state of the general aviation aircraft during landing is detected in real time through the multi-scale cross-space learning airport runway detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a general aviation aircraft landing runway detection method based on multi-scale cross-space learning. Background Art

[0002] With the rapid development of machine vision technology, numerous aviation safety agencies, research organizations, and leading companies are actively developing, testing, and certifying aircraft visual landing systems (VLS). An aircraft's flight phases include takeoff, cruising, and landing. Landing is the most challenging phase and carries the greatest risk of accidents. A successful landing requires maintaining an appropriate glide path angle and descent speed, ensuring that the aircraft's flight path aligns with the runway centerline and passes through the designated landing point. Therefore, reducing pilot workload and improving landing safety are key goals for the aviation industry.

[0003] Existing navigation systems for general aviation aircraft during the landing phase include instrument landing systems and ground-based augmentation systems, but the deployment costs of these systems are high. In recent years, visual navigation systems have become an emerging development in this field due to their low cost, and many studies aim to achieve automatic landing with the help of computer-aided technology. General aviation airport runway detection, as an important component of general aviation visual navigation systems, performs pixel-level classification of runway markings. The detection results can indicate whether the general aviation aircraft is aligned with the runway centerline and whether the current glide path angle is reasonable, thereby helping general aviation pilots better perceive the runway position, achieve automatic landing, and improve landing safety. Therefore, for general aviation aircraft landing, visual image-based detection technology is the most suitable real-time detection method. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a general aviation aircraft landing runway detection method based on multi-scale cross-space learning to realize the detection of general aviation aircraft landing runways.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] The present invention provides a general aviation aircraft landing runway detection method based on multi-scale cross-space learning, comprising:

[0007] S1. Collect a dataset of runway images of general aviation aircraft landing and preprocess the dataset.

[0008] S2. Construct an AES-YOLOv8 network model for detecting multi-scale runway targets. The AES-YOLOv8 network model is based on the YOLOv8s network for multi-scale cross-space learning. The Backbone network includes several improved lightweight channel attention convolution modules LA-AKConv. The EMA attention mechanism is introduced after the spatial pyramid pooling module SPPF of the Backbone network to ensure that spatial semantic features are evenly distributed in each feature group. To reduce spatial and channel redundancy in standard convolution and enhance feature representation, the Neck network includes several C2f_SC Conv modules that fuse C2f with spatial and channel reconstruction convolution modules SCConv.

[0009] S3. Establish the loss function S_CIOU of the AES-YOLOv8 network model for detecting multi-scale runway targets;

[0010] S4. The AES-YOLOv8 network model is trained based on the preprocessed dataset to obtain a multi-scale cross-space learning airport runway detection model. The multi-scale cross-space learning airport runway detection model is used to detect the runway status of general aviation aircraft during landing in real time.

[0011] Furthermore, in step S1, the process of obtaining the runway image dataset includes: autonomously collecting general aviation airport runway images, collecting open-source airport runway detection datasets, and integrating the autonomously collected general aviation airport runway images with the open-source airport runway datasets to obtain a final runway image dataset;

[0012] The preprocessing process of the dataset includes: labeling the airport runway images in the dataset, that is, labeling the overall airport runway labels; dividing the labeled dataset into a training set and a test set in a ratio of 9:1;

[0013] Furthermore, the AES-YOLOv8 network model constructed in step S2 is based on a multi-scale cross-space learning YOLOv8s network, including a backbone network Backbone, a neck network Neck, and a head network Head;

[0014] The backbone network Backbone includes a conventional convolution layer Conv, several lightweight channel attention convolution modules LA-AKConv, and an EMA attention mechanism module with cross-space learning introduced after the spatial pyramid pooling module SPPF of the backbone network. In the backbone network, the image first passes through several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules to gradually extract feature maps of different scales. Each group of lightweight channel attention convolution modules LA-AKConv and C2f modules inputs the extracted features to the next layer for higher-level feature extraction, until the output is to the EMA attention mechanism module introduced after the spatial pyramid pooling module SPPF of the backbone network Backbone.

[0015] The neck network Neck includes two upsampling modules Upsample, several concatenation modules Concat, at least one C2f module, and several C2f_SCConv modules that fuse C2f with SCConv space and channel reconstruction convolution modules. In the neck network, the high-level feature map output by the EMA attention mechanism module of the backbone network is first upsampled, and then fused with the low-level feature maps gradually extracted by several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules of the backbone network through the concatenation module Concat. The fused feature map is subjected to feature extraction by the convolution module C2f and the C2f_SCConv module to reduce spatial and channel redundancy and obtain a multi-scale feature map.

[0016] The head network Head obtains the multi-scale feature map from the neck network Neck and performs the final target detection task through the detection module Detect, including bounding box regression and category prediction.

[0017] Furthermore, the lightweight channel attention convolution module LA-AKConv in step S2 introduces a depth-wise separable convolution architecture for spatial feature extraction based on the AKConv convolution, introduces a channel attention mechanism to parallelly calculate channel attention weights to enhance key channels, enhance the channel perception ability of offset prediction, improve the initial sampling template generation strategy, and enhance geometric adaptability; the features input to the lightweight channel attention convolution module LA-AKConv are first subjected to spatial feature extraction through 3×3 depth-wise separable convolution, and the channel attention weights are calculated in parallel to enhance key channels; then the dynamic sampling offset is predicted using the attention-modulated features, and adaptive sampling coordinates are generated in combination with the preset basic template; the feature map is geometrically resampled through bilinear interpolation, and the multi-sampling point features are reorganized into strip features; and finally, the channel information is fused and output through 1×1 point convolution.

[0018] Furthermore, the EMA attention mechanism module with cross-space learning takes the feature map output by the spatial pyramid pooling module SPPF as input, and the feature map is divided into g groups according to the channel dimension for feature extraction, and the size of each group is c / / g×h×w, where / / is a separator, c is the channel dimension, g is the number of image groups, and h and w are the height and width of the image respectively; and two parallel branches are used to extract global information and local spatial information of the grouped input images to generate an attention map; the first branch uses a 1D global pooling layer to extract the global information of the input image, and then performs feature fusion through a 1*1 convolution layer, and finally generates an attention map through a Sigmoid function; the second branch uses a 3*3 convolution kernel to capture the local spatial information of the input image, and then uses a Softmax function to standardize and generate an attention map; the attention maps generated by the two branches are subjected to matrix multiplication operations to generate a global attention map, which captures the pairing relationship at the pixel level; the final attention map is combined with the input feature map through addition and multiplication operations to generate the output features of the backbone network.

[0019] Furthermore, in the neck network Neck, three C2f_SCConv modules are used to replace the second C2f module, the third C2f module, and the fourth C2f module respectively; the C2f_SCConv module achieves module-level functional enhancement by replacing the original bottleneck structure Bottleneck in the C2f module with a customized Bottleneck_SCConv structure. This inheritance mechanism maintains the multi-branch characteristics of the original C2f while introducing a new convolution operation.

[0020] Furthermore, the Bottleneck_SCConv performs two-stage improvements on the original Bottleneck in the C2f module:

[0021] The first stage maintains standard convolution for channel expansion;

[0022] In the second stage, the SCConv module is used to replace the traditional convolution. The module includes the spatial reconstruction unit SRU and the channel reconstruction unit CRU:

[0023] The spatial reconstruction unit (SRU) uses group normalization + adaptive gating mechanism to achieve dynamic selection and reorganization of feature channels;

[0024] The channel reconstruction unit (CRU) uses a channel segmentation-transformation-fusion strategy to enhance the multi-scale feature expression capability through GWC / PWC hybrid convolution;

[0025] Within the C2f framework, multiple improved Bottleneck_SCConvs form a cyclic chain structure to achieve dense connection enhancement. Each bottleneck SCConv module uses space-sensitive feature selection and intelligent fusion of channel dimensions to enable continuous adaptive optimization of features when they are transferred between layers. Ultimately, more efficient gradient flow and feature reuse are achieved through cross-layer connections, forming a C2f_SCConv module.

[0026] Furthermore, the AES-YOLOv8 model considers the alignment and multi-scale of bounding boxes, and uses an improved loss function S_CIoU in bounding box prediction, that is, the SIoU loss function is introduced on the basis of the CIoU loss function.

[0027] Furthermore, the SIoU loss function is composed of angle loss Δ, distance loss ▽, shape loss Ω and IoU. The SIoU formula is as follows:

[0028]

[0029] The distance loss formula is as follows:

[0030]

[0031] in, c w , c h The difference between the horizontal and vertical coordinates of the center point of the predicted frame and the center point of the real frame, respectively. α is the angle between the line connecting the center point of the predicted frame and the center point of the real frame and the horizontal line. (b cx ,b cy ) is the center coordinate of the prediction box, is the center coordinate of the real frame;

[0032] The angle loss formula Δ is as follows:

[0033]

[0034] in, x is the sine value of the angle α, which is normalized by get;

[0035] The shape loss Ω formula is as follows:

[0036]

[0037] in, σ is the distance between the center point of the real box and the predicted box, θ is the parameter range used to control the degree of attention to shape loss; ω w and ω his the normalized difference between width and height. The normalization logic is: the denominator takes the maximum value between the predicted value and the true value, ensuring that the difference value is between [0,1]. w and h are the width and height of the predicted box, w gt and h gt is the width and height of the real frame;

[0038] The improved S_CIoU loss function formula is as follows:

[0039]

[0040] in, ρ 2 (b,b gt ) represents the Euclidean distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the real box. and Represent the aspect ratios of the true box and the predicted box respectively.

[0041] The beneficial effects of adopting the above technical solution are: the multi-scale cross-space learning general aviation aircraft landing runway detection method provided by the present invention includes: collecting a general aviation aircraft landing video image dataset; training an AES-YOLOv8 network model based on the video dataset, the AES-YOLOv8 network model including the eighth-generation YOLO (You Only Look Once, YOLO) model, wherein the initial convolution structure of the 0th layer of the Backbone part of the YOLOv8s model is not replaced, the original convolution structure is maintained, and the standard convolution positions 1, 3, 5, and 7 are replaced with LA-AKConv with the same kernel_size=3; the EMA attention mechanism is introduced after the spatial pyramid pooling module SPPF; C2f_SCConv is introduced into its neck network to replace the 2nd, 3rd, and 4th C2f modules; a loss function S_CIOU of the YOLOv8 network structure for detecting multi-scale runway targets is established; based on the trained AES-YOLOv8 model, the general aviation aircraft landing video is input for detection, and the detection effect is observed. This application is based on the AES-YOLOv8 model and can detect airport runways for general aviation aircraft, ultimately improving accuracy and reducing the operational risks of general aviation aircraft pilots during landing. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A model structure diagram of a general aviation aircraft landing runway detection method using multi-scale cross-space learning provided by an embodiment of the present invention;

[0043] Figure 2 Provides a schematic diagram of the AKConv convolution module structure for an embodiment of the present invention;

[0044] Figure 3 Schematic diagram of the improved lightweight channel attention AKConv convolution process provided by an embodiment of the present invention;

[0045] Figure 4 EMA attention mechanism module structure diagram provided by an embodiment of the present invention;

[0046] Figure 5 A network structure diagram of the improved C2f_SCConv module provided in an embodiment of the present invention;

[0047] Figure 6 A network structure diagram of the SCConv module provided in an embodiment of the present invention;

[0048] Figure 7 Schematic diagram of the training results of the AES-YOLOv8 model provided in an embodiment of the present invention, wherein (a) represents the training set used to supervise the regression of the detection box, and the error loss graph between the predicted box and the calibration box, (b) represents the training set used to supervise the category classification, and the loss graph for calculating whether the anchor box and the corresponding calibration classification are correct, (c) represents the training set used to optimize the Distribution Focal Loss of the bbox, (d) represents the precision, (e) represents the recall rate, (f) represents the validation set used to supervise the regression of the detection box, and the error loss graph between the predicted box and the calibration box, (g) represents the validation set used to supervise the category classification, and the loss graph for calculating whether the anchor box and the corresponding calibration classification are correct, (h) represents the validation set used to optimize the Distribution Focal Loss of the bbox, (i) represents the mAP value change curve when the intersection-over-union (IoU) threshold is 0.5, and (j) represents the average mAP at different IoU thresholds (from 0.5 to 0.95, with a step size of 0.05);

[0049] Figure 8 This is a diagram of some runway detection results provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0051] In this embodiment, the multi-scale cross-space learning method for general aviation aircraft landing runway detection includes the following steps:

[0052] S1. Collect a dataset of runway images of general aviation aircraft landing and preprocess the dataset.

[0053] S2. Build an AES-YOLOv8 network model for detecting multi-scale runway targets; Figure 1As shown in the figure, the AES-YOLOv8 network model is based on the YOLOv8s network for multi-scale cross-space learning. The Backbone network includes several improved lightweight channel attention convolution modules LA-AKConv. The EMA attention mechanism is introduced after the spatial pyramid pooling module SPPF of the Backbone network, so that the spatial semantic features are evenly distributed in each feature group. To reduce the spatial and channel redundancy in the standard convolution and enhance the feature representation, the Neck network includes several C2f_SCConv modules that are fused with the spatial and channel reconstruction convolution modules SCConv.

[0054] S3. Establish the loss function S_CIOU of the YOLOv8 network model for detecting multi-scale runway targets;

[0055] S4. Train the AES-YOLOv8 network model based on the preprocessed dataset to obtain a multi-scale cross-space learning airport runway detection model. This multi-scale cross-space learning airport runway detection model is used to detect the runway status of general aviation aircraft during landing in real time. This application, based on the recognition of the improved YOLOv8 network model, can detect general aviation aircraft runways and reduce the operational risks for general aviation aircraft pilots during landing.

[0056] In this embodiment, step S1 mainly searches and downloads image datasets from the perspective of general aviation aircraft landing from domestic and foreign websites, and at the same time takes runway images from the perspective of general aviation aircraft landing at the airport, and then screens and integrates them to obtain a self-made detection dataset. Among them, the self-made detection dataset includes web page images, open source datasets, and images collected on-site at the General Aviation Laboratory in Shenyang, Liaoning Province. The open source datasets include two airport runway detection datasets, BARS and RLD. The BARS dataset is simulated based on the FAA-certified X-Plane simulation platform and collects 10,256 airport runway images, including runway images of various airports, aircraft views, weather conditions, and time intervals. Among them, very similar images and images that cannot be annotated due to long distances are removed. The RLD dataset, also simulated using the FAA-certified X-Plane simulation platform, contains 12,239 1280×720 pixel images of general aviation runways in various terrains (cities, villages, grasslands, mountains, oceans, etc.) and in various weather conditions (sunny, cloudy, foggy, and night). Some of these images were collected from YouTube videos of civil aircraft landings captured by front-facing cameras. The field images include images of general aviation runways on both sunny and cloudy days.

[0057] The integrated data is labeled, that is, the entire airport runway part from the aircraft landing perspective is labeled, and divided into training set and test set.

[0058] The AES-YOLOv8 network model constructed in step S2 is based on a YOLOv8s network for multi-scale cross-space learning, including a backbone network, a neck network, and a head network.

[0059] The backbone network Backbone includes a conventional convolution layer Conv, several lightweight channel attention convolution modules LA-AKConv, and an EMA attention mechanism module with cross-space learning introduced after the spatial pyramid pooling module SPPF of the backbone network. In the backbone network, the image first passes through several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules to gradually extract feature maps of different scales. Each group of lightweight channel attention convolution modules LA-AKConv and C2f modules inputs the extracted features to the next layer for higher-level feature extraction, until the output is to the EMA attention mechanism module introduced after the spatial pyramid pooling module SPPF of the backbone network Backbone.

[0060] The neck network Neck includes two upsampling modules Upsample, several concatenation modules Concat, at least one C2f module, and several C2f_SCConv modules that fuse C2f with SCConv space and channel reconstruction convolution modules. In the neck network, the high-level feature map output by the EMA attention mechanism module of the backbone network is first upsampled, and then fused with the low-level feature maps gradually extracted by several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules of the backbone network through the concatenation module Concat. The fused feature map is subjected to feature extraction by the convolution module C2f and the C2f_SCConv module to reduce spatial and channel redundancy and obtain a multi-scale feature map.

[0061] The head network Head obtains the multi-scale feature map from the neck network Neck and performs the final target detection task through the detection module Detect, including bounding box regression and category prediction.

[0062] The lightweight channel attention convolution module LA-AKConv introduces a depth-wise separable convolution architecture for spatial feature extraction based on the AKConv convolution, introduces a channel attention mechanism to parallelly calculate channel attention weights to enhance key channels, enhance the channel perception ability of offset prediction, improve the initial sampling template generation strategy, and enhance geometric adaptability; the features input to the lightweight channel attention convolution module LA-AKConv are first subjected to 3×3 depth-wise separable convolution for spatial feature extraction, and the channel attention weights are calculated in parallel to enhance the key channels; then the attention-modulated features are used to predict dynamic sampling offsets, and adaptive sampling coordinates are generated in combination with the preset basic template; the feature map is geometrically resampled through bilinear interpolation, and the multi-sampling point features are reorganized into strip features; finally, the channel information is fused and output through 1×1 point convolution.

[0063] In this embodiment, in step S2, the specific process of LA-AKConv replacing the standard Conv structure in the YOLOv8 network model includes: not replacing the initial convolution structure of the 0th layer of the Backbone network of the YOLOv8s model, maintaining the original convolution structure, and replacing the standard convolution positions 1, 3, 5, and 7 with LA-AKConv with the same kernel_size=3;

[0064] like Figure 2 As shown in the figure, AKConv (Convolutional Kernel with Arbitrary Sampled Shapes and Arbitrary Number of Parameters) is an innovative convolution layer in CNN. It achieves efficient feature extraction by breaking through the fixed sampling mode of traditional convolution kernels. Compared with standard convolution (fixed k×k grid sampling) and deformable convolution (requiring additional parameter prediction offset), its parameter quantity only grows linearly with the convolution kernel size (standard / deformable convolution grows quadratically); it supports sampling point distribution and parameter quantity configuration of arbitrary shapes. Figure 3 As shown in the figure, the improved LA-AKConv module based on AKConv further reduces the number of parameters by introducing depthwise separable convolution and proposes an offset prediction mechanism guided by the channel attention module (ChannelAttention) to enhance the semantic perception ability of dynamic sampling, including:

[0065] Input features X∈R C×H×W First, spatial features are extracted through depth-wise separable convolution:

[0066]

[0067] in, is the channel-by-channel spatial convolution, W dw ∈RC×3×3 is a learnable parameter; each input channel performs a 3×3 convolution independently to extract local spatial features, and the output dimension remains C×H×W;

[0068] The channel attention module runs in parallel with the depthwise separable convolution to semantically enhance the original input features; the channel attention module compresses the spatial information of each channel through global average pooling (GAP) to generate a channel description vector z c ∈R C :

[0069]

[0070] Learn the channel weights α∈R through two fully connected layers (compression rate r=8) C :

[0071] α C =σ(W2δ(W1z c ))

[0072] Where W1∈R C / r×C , W2∈R C×C / r It forms a bottleneck structure, where σ is the Sigmoid activation function; the original feature X is multiplied by α channel by channel to enhance the response of important channels and suppress noise;

[0073] The attention-modulated features are input to the offset prediction network (3×3 convolutional layer) to generate the dynamic offset ΔP∈R of N sampling points 2N×H×W :

[0074] ΔP=F offset (α c ⊙X)

[0075] Among them, F offset Represents a 3×3 convolutional layer, where the output channels 2N correspond to the two-dimensional offset of N sampling points, where the first N channels are the horizontal offset Δx and the last N channels are the vertical offset Δy. The initial weight of the offset convolution kernel is set to 0 to ensure that the sampling points are gradually adjusted from a regular grid at the beginning of training.

[0076] Based on a three-level coordinate synthesis method, the image sampling position is determined by using a regular grid, an adaptive template, and dynamic offset to improve the model's robustness to geometric deformation and local features:

[0077] The reference grid generates regular grid coordinates P0 according to the feature map size and step size (for example, when the step size is 2, the coordinate interval is 2 pixels); the initial template generates the initial sampling point P based on the adaptive strategy. n , supports any number N, if N = m 2, arranged in an m×m grid; if N is not a square number, a row and column filling strategy is used (for example, N=5 generates a 2×3 grid and cuts off the extra points); finally, the dynamic offset is superimposed to obtain the dynamic coordinates of the final image sampling position:

[0078] P=P0+P n +ΔP

[0079] Based on the dynamic coordinate P, features are extracted from the depthwise convolution output through bilinear interpolation. For each sampling point p = (x, y), its four nearest grid corner points are determined. The interpolation weight is calculated based on the normalized distance, and the weighted sum of the four corner point features is used to obtain a smooth sampling value that preserves spatial continuity.

[0080] Subsequently, the features of N sampling points are concatenated along the height dimension to obtain a strip feature map with a dimension of C×(H·N)×W. The information is integrated across channels through 1×1 convolution. Finally, the LA-AKConv module outputs a feature map with a dimension of C out ×H×W, compatible with standard convolution, where C out is the number of output feature map channels;

[0081] An efficient multi-scale attention mechanism (EMA) with cross-spatial learning is introduced after the spatial pyramid pooling module (SPPF) of the backbone network BackBone to perform multi-scale cross-spatial feature extraction.

[0082] The EMA mechanism is used for feature extraction in neural networks. It groups input features by channel dimension and uses two parallel branches to capture global and local features. The output features of the two branches are fused through matrix multiplication to generate an attention map, which is combined with the input features. This ultimately improves the model's ability to capture global and local information while reducing computational complexity.

[0083] The EMA attention mechanism module with cross-space learning takes the feature map output by the spatial pyramid pooling module SPPF as input, and the feature map is divided into g groups according to the channel dimension for feature extraction. The size of each group is c / / g×h×w, where / / is a separator, c is the channel dimension, g is the number of image groups, and h and w are the height and width of the image respectively; and two parallel branches are used to extract global information and local spatial information from the grouped input images to generate an attention map; the first branch uses a 1D global pooling layer to extract the global information of the input image, and then performs feature fusion through a 1*1 convolution layer, and finally generates an attention map through a Sigmoid function; the second branch uses a 3*3 convolution kernel to capture the local spatial information of the input image, and then uses a Softmax function to standardize and generate an attention map; the attention maps generated by the two branches are subjected to matrix multiplication operations to generate a global attention map, which captures the pairing relationship at the pixel level; the final attention map is combined with the input feature map through addition and multiplication operations to generate the output features of the backbone network.

[0084] Specifically, if Figure 4 As shown, EMA selects only the shared component of the 1x1 convolution from the CA module, which is named the 1x1 branch in EMA. To aggregate multi-scale spatial structure information, a 3x3 kernel is placed in parallel with the 1x1 branch for fast response, namely the 3x3 branch. Considering the feature grouping and multi-scale structure, it is advantageous to effectively establish short-range and long-range dependencies for better performance.

[0085] The feature map output by the SPPF module is used as the input of EMA X∈R C×H×W , where C, H, and W are the number of channels, height, and width of the feature map X respectively; EMA will divide the feature map X into G sub-features in the channel dimension direction to learn different semantics. Its group-style group style can be represented by X=[X0,X i ,...,X G-1 ,],X i ∈R C / / G×H×WGiven. Without loss of generality, let G << C and assume that the attention weight descriptors learned by EMA will be used to enhance the feature representation of the regions of interest in each sub - feature. The large local receptive field of neurons enables neurons to collect multi - scale spatial information. Therefore, EMA uses three parallel paths to extract the attention weight descriptors of the grouped feature maps. Two parallel paths are 1x1 branches, and the third path is a 3x3 branch. To capture the dependencies between all channels and reduce the computational budget, cross - channel information interaction is modeled in the channel direction. More specifically, two 1D global average pooling operations are used in the 1x1 branches to encode the channels along two spatial directions respectively, and only a single 3x3 kernel is stacked in the 3x3 branch to capture multi - scale feature representations.

[0086] The parameter dimension of a normal 2D convolutional kernel in the deep learning framework Pytorch is [oup, inp, k, k], which does not involve the batch dimension, where oup represents the output plane of the input, inp represents the input plane of the input feature, and k represents the kernel size respectively. Accordingly, g groups are reshaped and replaced into the batch dimension, and the input tensor is re - defined with the shape of C / / G * H * W.

[0087] On the one hand, through a process similar to CA, the two encoded features are concatenated with respect to the image height direction and share the same 1x1 convolution in the 1x1 branch without dimensionality reduction. After factoring the output of the 1 * 1 convolution into two vectors, two non - linear Sigmoid functions are used to fit the 2D binomial distribution on the linear convolution. To achieve different cross - channel interaction features between the two parallel routes in the 1x1 branch, the two channel - wise attention maps within each group are aggregated together through simple multiplication. On the other hand, the 3x3 branch captures local cross - channel interaction through 3x3 convolution to expand the feature space. Specifically, not only is the cross - channel information encoded to adjust the importance of different channels, but the precise spatial structure information is also preserved in the channels.

[0088] The cross - spatial information aggregation in different spatial dimension directions introduces two tensors, one of which is the output of the 1x1 branch and the other is the output of the 3x3 branch. Then, two - dimensional global average pooling is used to encode the global spatial information of the output of the 1x1 branch. Before the channel feature joint activation mechanism, the output of the smallest branch is directly transformed into the corresponding dimensional shape, that is The formula for the two - dimensional global pooling operation is

[0089]

[0090] where, z c is the two - dimensional global pooling output under channel c; x c(i, j) represents the input characteristics under channel c;

[0091] Two-dimensional global pooling is used to encode global information and model long-range dependencies. Softmax, a natural nonlinear function of the two-dimensional Gaussian map, is applied to the output of the two-dimensional global average pooling to fit the above linear transformation. The first spatial attention map is obtained by multiplying the output of the above parallel processing with a matrix dot product operation. To observe this, EMA collects spatial information at different scales in the same processing stage.

[0092] In addition, the 2D global average pooling is also used to encode the global spatial information in the 3x3 branch, and the 1x1 branch is directly converted to the corresponding dimensional shape before the channel feature joint activation mechanism, that is,

[0093] Building on this, the second branch derives a second spatial attention map that preserves the entire precise spatial location information. Finally, the output feature map within each group is calculated as a set of the two generated spatial attention weights. A sigmoid function is then applied to capture pixel-level pairwise relationships and highlight the global context of all pixels. The final output of the EMA is the same size as the general airport runway feature map X, making it efficient for stacking into the model.

[0094] As mentioned above, the attention factor is only guided by the similarity between the global and local feature descriptors in each group.

[0095] Taking into account the cross-spatial information aggregation method, precise location information is embedded in the EMA while modeling long-range dependencies. Fusion of contextual information at different scales enables the CNN to generate better pixel-level attention on high-level feature maps. Subsequently, by using a cross-spatial learning method, parallelization of convolutional kernels is a more powerful architecture for handling both short-term and long-term dependencies. In contrast to the progressive behavior of limited receptive field formation, the parallel use of 3x3 and 1x1 convolutions in intermediate feature maps captures more contextual information.

[0096] In the neck network Neck, three C2f_SCConv modules are used to replace the second C2f module, the third C2f module and the fourth C2f module respectively; Figure 5 As shown, the C2f_SCConv module inherits from the base C2f module, retaining its core structure while implementing key modifications to the bottleneck network Neck. By replacing the original Bottleneck with the customized Bottleneck_SCConv, module-level functionality is enhanced. This inheritance mechanism maintains the multi-branch nature of the original C2f while introducing new convolutional operations. Bottleneck_SCConv undergoes a two-stage transformation:

[0097] The first stage maintains the standard convolution (Conv) for channel expansion;

[0098] In the second stage, the SCConv module is innovatively used to replace the traditional convolution. The SCConv module includes the spatial reconstruction unit SRU and the channel reconstruction unit CRU.

[0099] The spatial reconstruction unit (SRU) uses group normalization + adaptive gating mechanism to achieve dynamic selection and reorganization of feature channels;

[0100] The channel reconstruction unit (CRU) uses a channel segmentation-transformation-fusion strategy to enhance the multi-scale feature expression capability through GWC / PWC hybrid convolution;

[0101] Within the C2f framework, multiple improved Bottleneck_SCConvs form a recurrent chain structure to achieve dense connectivity enhancement. The SCConv module of each bottleneck network uses spatially sensitive feature selection (SRU's gated threshold mechanism) and intelligent fusion of channel dimensions (CRU's soft attention weighting) to continuously and adaptively optimize features as they are transferred between layers. Ultimately, cross-layer connections achieve more efficient gradient flow and feature reuse, forming the C2f_SCConv module.

[0102] The SCConv (spatial and channel reconstruction convolution) module is as follows Figure 6 As shown in Figure 1, it is an efficient convolution module designed to optimize the performance of convolutional neural networks (CNNs) and reduce computing resource consumption by reducing spatial and channel redundancy. The module consists of two core units:

[0103] (1) The spatial reconstruction unit SRU uses separation and reconstruction operations: the purpose of the separation operation is to separate the feature map with rich information from the feature map with less information corresponding to the spatial content. The reconstruction operation adds the feature with rich information to the feature with less information to generate a feature with richer information, thereby saving space. The cross reconstruction operation is used to fully combine the two different information features after weighting and strengthen the information flow between them. The cross-reconstructed features are then spliced ​​to obtain the spatial fine feature map X w .

[0104] (2) The channel reconstruction unit CRU adopts a split transformation and fusion strategy to reduce the redundancy of the channel dimension as well as the computational cost and storage. This unit adopts the split-transform-fusion method. The split operation divides the input spatial refinement features into two parts, one with a channel number of αC and the other with a channel number of (1-α)C, where C is the total number of channels of the image feature and α is the scale factor; then the channel number of the two sets of features is compressed using a 1*1 convolution kernel to obtain the compressed features X and X respectively. up and compression feature X low The conversion operation converts the input compressed features X up As the input of "rich feature extraction", GWC and PWC are performed respectively, and then the output feature Y1 is obtained by adding the input compressed feature X low As a supplement to "rich feature extraction", PWC is performed, and the obtained output is combined with the original input to obtain feature Y2. The fusion operation uses a simplified SKNet method to adaptively merge feature Y1 and feature Y2. Specifically, global average pooling is first used to combine global spatial information and channel statistical information to obtain pooled features S1 and S2. Then, Softmax is performed on features S1 and S2 to obtain feature weight vectors β1 and β2. Finally, the feature weight vector is used to obtain the output feature Y = β1Y1 + β2Y2, where Y is the feature extracted by the SCConv module channel;

[0105] The spatial reconstruction unit SRU and the channel reconstruction unit CRU are placed sequentially. Specifically, for the intermediate input feature X in the bottleneck residual block of the backbone network, the spatial refinement feature X is first obtained through the SRU operation. w , and then the CRU operation is used to obtain the channel-refined feature Y. The SCConv module exploits the spatial and channel redundancy between features to reduce the redundancy between intermediate feature maps and enhance the feature representation of the network.

[0106] Among them, the spatial reconstruction unit SRU (Spatial Reconstruction Unit) uses separation and reconstruction operations. The purpose of the separation operation is to separate those feature maps with rich information from those with less information corresponding to spatial content. The scaling factor in the group normalization layer is used to evaluate the information content of different feature maps. Specifically, given an intermediate feature map X∈R N×C×H×W , where N is the batch dimension, C is the channel dimension, and H and W are the spatial height and width dimensions. First, the input features X are normalized by subtracting the mean μ and dividing by the standard deviation σ, as follows:

[0107]

[0108] Where μ and σ are the mean and standard deviation of X, ε is a small positive constant added for division stability, and γ and β are trainable affine transformations. Using the trainable parameters γ∈R in the GN layer C As a method to measure the spatial pixel variance of each batch and channel. Richer spatial information reflects more changes in spatial pixels, resulting in a larger γ. The normalized correlation weight W is obtained by the following formula γ ∈R C , which shows the importance of different feature maps.

[0109]

[0110] Then, through W γ The weight values ​​of the reweighted feature maps are mapped to the range (0, 1) by the Sigmoid function and gated by the threshold. The weights above the threshold are set to 1 to obtain the informative weight W1, and they are set to 0 to obtain the non-informative weight W2 (the threshold is set to 0.5 in the experiment). The whole process of obtaining the informative weight W can be expressed as:

[0111] W=Gate(Sigmoid(W γ (GN(X))))

[0112] Finally, the input feature X is multiplied by W1 and W2 respectively to produce two weighted features: feature information-rich and less feature information and With little or no information, it is considered redundant. In order to reduce spatial redundancy, the information-rich features and the less information-rich features are cross-reconstructed to fully combine the weighted two different information features and strengthen the information flow between them. Then, the cross-reconstructed features are concatenated. and To obtain the spatially refined feature map X w To generate more informative features and save space.

[0113] Among them, the channel reconstruction unit CRU (Channel Reconstruction Unit) adopts the segmentation transformation and fusion strategy. k ∈R c×k×k represents the k×k convolution kernel, X, Y∈R c×h×w Denote the input and convolution output features respectively. Standard convolution can be defined as Y = M k X. Specifically, CRU replaces the standard convolution through three operators: Split, Transform, and Fuse.

[0114] Among them, Split: For a given spatial refinement feature X w∈R c×h×w , first change X w The channel of is divided into two parts, namely αC channel and (1-α)C channel, where 0≤α≤1 is the split ratio. Subsequently, 1×1 convolution is used to compress the channel of the feature map to improve computational efficiency. A compression ratio r is introduced to control the feature channel to balance the computational cost of CRU. After the segmentation and squeezing operations, the spatially refined feature X is w Divided into upper X up and lower X low .

[0115] Transform: X up It is fed into the upper transformation stage and acts as a "rich feature extractor". Efficient convolution operations (i.e., GWC and PWC) are used to replace the expensive standard k×k convolution to extract high-level representative information and reduce computational cost. Due to the sparse convolutional connections, GWC reduces the number of parameters and computation, but cuts off the information flow between channel groups. PWC compensates for the information loss and helps information flow across feature channels. Therefore, in the same X up k×kGWC and 1×1PWC operations are performed on the up-conversion layer. Then, the outputs are summed to form the combined representative feature map Y1. The up-conversion stage can be expressed as:

[0116] Y1=M G X up +M P1 X up

[0117] in is the learnable weight matrix of GWC and PWC, and Y1∈R c×h×w are the upper input and output feature maps respectively. In short, the upper transformation stage is on the same feature map X up The combination of GWC and PWC is used to extract rich representative features Y1 with less computational cost. low is fed into the lower transformation stages, where a 1×1 PWC operation is applied to generate feature maps with shallow hidden details as a complement to the rich feature extractor. In addition, we reuse X low Features are used to obtain more feature maps without additional cost. Finally, the generated and reused features are concatenated to form the output of the next level Y2 as shown below:

[0118]

[0119] in is the learnable weight matrix of PWC, ∪ is the cascade operation, and Y 2 ∈R c×h×ware the lower input and output feature maps respectively. Finally, the lower transformation stage reuses the previous feature X low And 1×1 PWC is used to obtain feature Y2 with supplementary detailed information.

[0120] Fuse: After executing Transformer, the simplified SKNet method is used to adaptively merge the output features Y1 and Y2 of the upper and lower transformation stages. First, global average pooling is applied to collect global spatial information S with channel statistics. m ∈R c×1×1 , the calculation formula is:

[0121]

[0122] Next, the upper and lower global channel descriptors S1, S2 are stacked together and the channel soft attention operation is used to generate the feature importance vectors β1, β2∈R c , as shown below:

[0123]

[0124] Finally, under the guidance of the feature importance vectors β1 and β2, the channel-refined feature Y can be obtained by merging the upper feature Y1 and the lower feature Y2 in a channel-wise manner, as shown below:

[0125] Y=β1Y1+β2Y2

[0126] The AES-YOLOv8 model considers the alignment and multi-scale of bounding boxes. The improved loss function S_CIoU (Complete Intersection over Union) Loss is used in bounding box prediction. That is, the SIoU loss function is introduced on the basis of the CIoU loss function:

[0127] SIoU consists of angle loss Δ, distance loss ▽, shape loss Ω and IoU. The SIoU formula is as follows:

[0128]

[0129] The distance loss formula is as follows:

[0130]

[0131] in,

[0132] The formula for angle loss Δ is as follows:

[0133]

[0134] in,

[0135] The shape loss Ω formula is as follows:

[0136]

[0137] in, Where c w , c h The difference between the horizontal coordinate and the vertical coordinate of the center point of the predicted frame and the center point of the real frame, respectively, and α is the angle between the line connecting the two points and the horizontal line; (b cx ,b cy ) is the center coordinate of the prediction box, is the center coordinate of the real box, σ is the distance between the center of the real box and the predicted box, and θ controls the degree of attention to shape loss. The parameter range is between 2 and 6. w and h are the width and height of the predicted box, w gt and h gt is the width and height of the real box. The improved S_CIoU loss function formula is as follows:

[0138]

[0139] in, ρ 2 (b,b gt ) represents the Euclidean distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the real box. and Represent the aspect ratios of the true box and the predicted box respectively.

[0140] In this embodiment, the AES-YOLOv8 model is trained using a server. The training environment is Python 3.9, the deep learning framework is PyTorch 1.13.1, CUDA is 12.0, the operating system is Ubuntu 20.04, the processor is an Intel 7-13700KF with 16 cores and 24 threads, the maximum turbo frequency is 5.40GHz, the memory is 32GB, the graphics card is NVIDIA RTX 4090, and the video memory is 24GB; the number of training rounds is set to 100, the batch_size is set to 16, the imgsz is set to 640, and the workers is set to 4. Finally, the trained model results are as follows Figure 7 As shown in the figure, the accuracy is 98.622%, the recall rate is 98.292%, the average detection precision mAP_0.5 is 83.5%, and the average detection precision mAP_0.5-0.95 is 86.7%. Then, in the actual detection process, the real-time airport runway image is imported into the AES-YOLOv8 model to detect the airport runway area of ​​the image to be tested, and the following is obtained: Figure 8The result diagram of the airport runway detection area is shown.

[0141] During the training process of this embodiment, the input image is first sent to the backbone network BackBone for cross-spatial multi-scale feature extraction, the extracted feature map is sent to the Neck part for feature fusion, and finally sent to the Head part for prediction.

[0142] In the backbone network BackBone, the 0th layer processes an input image of size (640*640), the convolution kernel size is 3, the stride is 2, and the padding is 1. The input image size is reduced to a feature map of size (320*320); the first layer LA-AKConv performs the same operation as the previous layer; the second layer is the C2f module. This layer is repeated 3 times to extract the basic features of the input image through the initial convolution block, and further refine and enhance the features through multiple Bottleneck blocks. These Bottleneck blocks can capture more complex patterns and detailed feature fusion. The directly transmitted feature map and the processed feature map are fused through the Concat block, so that the model can comprehensively utilize multi-scale and multi-level information output generation, and the final feature map is generated through the last convolution block, providing rich feature representation for subsequent detection and classification tasks; the third layer performs LA-AKConv convolution operation with an output channel number of 256, a convolution kernel size of 3*3, and a stride of 2. The output feature map size is 80*80*256, and the length and width of the feature map have become 1 / 8 of the input image; the fourth layer C2f module, this layer is repeated 6 times, the number of output channels is 256, and the residual structure is used to splice the features output by the previous layers. After this layer, the feature map size is still 80*80*256; the fifth layer, the LA-AKConv convolution operation with an output channel number of 512, a convolution kernel size of 3*3, and a stride of 2 is performed, and the output feature map size is 40*40 *512, the feature map size has become 1 / 16 of the input image. The sixth layer is a C2f module, repeated six times, with 512 output channels and residual connections. After this layer, the feature map size remains at 40*40*512. The seventh layer performs a LA-AKConv convolution operation with 1024 output channels, a 3*3 kernel size, and a stride of 2. The output feature map size is 20*20*1024, and the feature map size has become 1 / 32 of the input image. The eighth layer is a C2f module, repeated three times, with 1024 output channels and residual connections. After this layer, the feature map size remains at 20*20*1024. The ninth layer is a fast spatial pyramid pooling layer (SPPF) with 1024 output channels and a kernel size of 5. The tenth layer introduces the EMA attention mechanism, which introduces two branches: one for the output of the 1*1 branch and the other for the output of the 3*3 branch. Then, 2D global average pooling is used to encode the global spatial information in the output of the 1*1 branch. The output of the smallest branch will be converted to the corresponding dimensional shape before the joint activation mechanism of the channel features. The first spatial attention map is obtained by multiplying the output of the above parallel processing with the matrix dot product operation.In addition, 2D global average pooling is also used to encode global spatial information in the 3*3 branch. The 1*1 branch will be directly converted to the corresponding dimensional shape before the joint activation mechanism of the channel features. After that, a second spatial attention map is derived, which retains the entire precise spatial position information. Finally, the output feature map within each group is calculated as the aggregation of the two generated spatial attention weight values, followed by a Sigmoid function, which captures pixel-level pairwise relationships and highlights the global context of all pixels. Since the final output of EMA is the same size as the input provided by the previous layer, the feature map size obtained by concat is 20*20*(512*4), and the feature map size obtained after one convolution is 20*20*1024. After EMA, the feature map size is still 20*20*1024.

[0143] In the neck network Neck, feature maps are upsampled and fused. The feature map (20*20) is upsampled by a factor of 2 to obtain a feature map of 40*40. The upsampled feature map is then fused (element-wise addition) with the (40*40) feature map to obtain the feature map F1 (40*40).

[0144] After processing the C2f_SCConv module, feature map F1 is upsampled by a factor of 2 to produce a feature map of size 80*80. This upsampled feature map is fused (element-wise addition) with (80*80) to produce feature map T1 (80*80). Next, the bottom-up fusion process involves convolution and fusion of T1 and F1. After convolution, T1 produces a feature map of size 80*80. This convolution is fused with feature map F1 (40*40) to produce feature map F2 (40*40). F2 is processed by the C2f_SCConv module to produce feature map T2 (40*40). This convolution is also performed on T2 to produce a feature map of size 40*40. This convolution is fused with (20*20) to produce feature map F3 (20*20). Finally, F3 is processed by the C2f_SCConv module to produce feature map T3 (20*20).

[0145] The final output is T1 (80*80), T2 (40*40), and T3 (20*20). These feature maps are sent to the Head for detection.

[0146] Based on the trained AES-YOLOv8 model, an image of an airport runway from the perspective of a general aviation aircraft landing was input for detection testing to determine the detection effect. The input image was LetterBox preprocessed to convert the original image size (1920, 1080) of the general aviation airport runway to the network input size (640, 640).

[0147] Scaling uses proportional scaling, which means finding the long side and scaling it to 640. Then, using the long side's scaling ratio (1920 / 640=3), the short side is scaled to 1080 / 3=360, and the short side is padded with grayscale to 640. However, after LetterBox processing, the input image does not reach 640*640, but 384*640. Since the model is downsampled five times and the EMA input and output sizes are the same, to speed up inference, it is sufficient to ensure that every pixel is valid and divisible by 32. As can be seen from the previous section, a vector of (1, 64, 8400) will be output after passing through the backbone network. 64 is calculated by 4*reg_max (reg_max=16). 4 refers to the distance from the predicted center point to the left (l), top (t), right (r), and bottom (b) of the predicted border. Reg_max refers to the range of the predicted border. When reg_max=16, under each predicted feature map (20*20, 40*40, 80*80), the maximum predicted box size that can be predicted is 30*30. 4*16 can be a matrix with 4 rows and 16 columns. The values ​​of l, t, r, and b follow ∑value=1 after softmax, and the final predicted result is the product of Index and the corresponding value.

[0148] At each feature map scale (20*20, 40*40, 80*80), the maximum bounding box size that can be predicted is 30*30. For example, in a 20*20 feature map, a large-scale feature map is used to predict the runway of a near-point general aviation aircraft. However, 30*30 exceeds the size of the 20*20 feature map, so the detection of close-range runways will not be missed. In a 40*40 feature map, 30*30 can predict most medium- and long-range runway targets (mapped back to 640*640, the target size is approximately 480*480). In an 80*80 feature map, 30*30 is also mainly used to predict small targets in the long-range general aviation airport runway.

[0149] Then, according to the scaling ratio, the bounding box size of 8400 grid_cell predictions is mapped back to the 640*640 scale, that is, the size of the input to the network, and the predicted LTRB representation is changed to XYWH mode, that is, center point / width and height mode.

[0150] Finally, the Cls classification branch in the detection head performs a Sigmoid() operation on all elements. Each element will pass through the Sigmoid() function independently to obtain a value in the (0, 1) interval.

[0151] After the above process, non-maximum suppression (NMS) and Scale_boxes modules are performed.

[0152] The NMS process is divided into three parts. The first part is mainly to filter out a part through the confidence threshold (each grid_cel will have nc predicted category values, and after Sigmoid, they are all between (0, 1), and the maximum value of nc is compared with the threshold), and convert the XYWH format to XYXY format. As a result, only 29 of the 8400 grid_cels are left after filtering.

[0153] The second part mainly uses the Cls tensor to pick out the category confidence and label subscripts of these 29 grid_cels.

[0154] The third part is to add an offset to the box and complete the label box filtering through the NMS that comes with torchvision. Adding an offset to different categories is to distinguish different categories. Finally, a 3-row 6-column matrix will be obtained, representing the three predicted targets and their corresponding XYXY format boxes, the category confidence, and the category subscript. The scale_boxes module maps the prediction results back to the original input image size. First, the predicted box is subtracted from the offset caused by the LetterBox, and the XYXY coordinates of each box are restored to the proportional scaling (360, 640). The XYXY coordinates are then proportionally enlarged to the coordinates of the original image (1920, 1080). Finally, the obtained XYXY coordinate information is cropped to the specified image size range to ensure that the bounding box does not exceed the actual size of the image. With continuous detection iterations, the detection of general aviation airport runways is achieved.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A multi-scale cross-space learning method for general aviation aircraft landing runway detection, characterized by: include: S1. Collect a dataset of runway images of general aviation aircraft landing and preprocess the dataset. S2. Construct an AES-YOLOv8 network model for detecting multi-scale runway targets. The AES-YOLOv8 network model is based on the YOLOv8s network for multi-scale cross-space learning. The Backbone network includes several improved lightweight channel attention convolution modules LA-AKConv. The EMA attention mechanism is introduced after the spatial pyramid pooling module SPPF of the Backbone network to ensure that spatial semantic features are evenly distributed in each feature group. To reduce spatial and channel redundancy in standard convolution and enhance feature representation, the Neck network includes several C2f_SC Conv modules that fuse C2f with spatial and channel reconstruction convolution modules SCConv. S3. Establish the loss function S_CIOU of the AES-YOLOv8 network model for detecting multi-scale runway targets; S4. The AES-YOLOv8 network model is trained based on the preprocessed dataset to obtain a multi-scale cross-space learning airport runway detection model. The multi-scale cross-space learning airport runway detection model is used to detect the runway status of general aviation aircraft during landing in real time.

2. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 1 is characterized by: In step S1, the process of obtaining the runway image dataset includes: autonomously collecting general aviation airport runway images, collecting open-source airport runway detection datasets, and integrating the autonomously collected general aviation airport runway images with the open-source airport runway datasets to obtain a final runway image dataset; The preprocessing process of the dataset includes: labeling the airport runway images in the dataset, that is, labeling the overall airport runway labels; dividing the labeled dataset into a training set and a test set in a ratio of 9:

1.

3. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 1, characterized in that: The AES-YOLOv8 network model constructed in step S2 is based on a YOLOv8s network for multi-scale cross-space learning, including a backbone network, a neck network, and a head network. The backbone network Backbone includes a conventional convolution layer Conv, several lightweight channel attention convolution modules LA-AKConv, and an EMA attention mechanism module with cross-space learning introduced after the spatial pyramid pooling module SPPF of the backbone network. In the backbone network, the image first passes through several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules to gradually extract feature maps of different scales. Each group of lightweight channel attention convolution modules LA-AKConv and C2f modules inputs the extracted features to the next layer for higher-level feature extraction, until the output is to the EMA attention mechanism module introduced after the spatial pyramid pooling module SPPF of the backbone network Backbone. The neck network Neck includes two upsampling modules Upsample, several concatenation modules Concat, at least one C2f module, and several C2f_SCConv modules that fuse C2f with SCConv space and channel reconstruction convolution modules. In the neck network, the high-level feature map output by the EMA attention mechanism module of the backbone network is first upsampled, and then fused with the low-level feature maps gradually extracted by several groups of lightweight channel attention convolution modules LA-AKConv and C2f modules of the backbone network through the concatenation module Concat. The fused feature map is subjected to feature extraction by the convolution module C2f and the C2f_SCConv module to reduce spatial and channel redundancy and obtain a multi-scale feature map. The head network Head obtains the multi-scale feature map from the neck network Neck and performs the final target detection task through the detection module Detect, including bounding box regression and category prediction.

4. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 3 is characterized by: The lightweight channel attention convolution module LA-AKConv in step S2 introduces a depth-wise separable convolution architecture for spatial feature extraction based on the AKConv convolution, introduces a channel attention mechanism to parallelly calculate channel attention weights to enhance key channels, enhances the channel perception ability of offset prediction, improves the initial sampling template generation strategy, and enhances geometric adaptability; the features input to the lightweight channel attention convolution module LA-AKConv are first subjected to spatial feature extraction through 3×3 depth-wise separable convolution, and the channel attention weights are calculated in parallel to enhance the key channels; then the dynamic sampling offset is predicted using the attention-modulated features, and adaptive sampling coordinates are generated in combination with the preset basic template; the feature map is geometrically resampled through bilinear interpolation, and the multi-sampling point features are reorganized into strip features; and finally, the channel information is fused and output through 1×1 point convolution.

5. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 4, characterized in that: The EMA attention mechanism module with cross-space learning takes the feature map output by the spatial pyramid pooling module SPPF as input, and the feature map is divided into g groups according to the channel dimension for feature extraction. The size of each group is c / / g×h×w, where / / is a separator, c is the channel dimension, g is the number of image groups, and h and w are the height and width of the image respectively; and two parallel branches are used to extract global information and local spatial information from the grouped input images to generate an attention map; the first branch uses a 1D global pooling layer to extract the global information of the input image, and then performs feature fusion through a 1*1 convolution layer, and finally generates an attention map through a Sigmoid function; the second branch uses a 3*3 convolution kernel to capture the local spatial information of the input image, and then uses a Softmax function to standardize and generate an attention map; the attention maps generated by the two branches are subjected to matrix multiplication operations to generate a global attention map, which captures the pairing relationship at the pixel level; the final attention map is combined with the input feature map through addition and multiplication operations to generate the output features of the backbone network.

6. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 5, characterized in that: In the neck network Neck, three C2f_SCConv modules are used to replace the second C2f module, the third C2f module, and the fourth C2f module respectively. The C2f_SCConv module achieves module-level functional enhancement by replacing the original bottleneck structure Bottleneck in the C2f module with a customized Bottleneck_SCConv structure. This inheritance mechanism maintains the multi-branch characteristics of the original C2f while introducing a new convolution operation.

7. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 6, characterized in that: The Bottleneck_SCConv performs two-stage improvements on the original Bottleneck in the C2f module: The first stage maintains standard convolution for channel expansion; In the second stage, the SCConv module is used to replace the traditional convolution. The module includes the spatial reconstruction unit SRU and the channel reconstruction unit CRU: The spatial reconstruction unit (SRU) uses group normalization + adaptive gating mechanism to achieve dynamic selection and reorganization of feature channels; The channel reconstruction unit (CRU) uses a channel segmentation-transformation-fusion strategy to enhance the multi-scale feature expression capability through GWC / PWC hybrid convolution; Within the C2f framework, multiple improved Bottleneck_SCConvs form a cyclic chain structure to achieve dense connection enhancement. Each bottleneck SCConv module uses space-sensitive feature selection and intelligent fusion of channel dimensions to enable continuous adaptive optimization of features when they are transferred between layers. Ultimately, more efficient gradient flow and feature reuse are achieved through cross-layer connections, forming a C2f_SCConv module.

8. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 7, characterized in that: The AES-YOLOv8 model considers the alignment and multi-scale of bounding boxes, and uses an improved loss function S_CIoU in bounding box prediction, that is, the SIoU loss function is introduced on the basis of the CIoU loss function.

9. The multi-scale cross-space learning method for general aviation aircraft landing runway detection according to claim 8, characterized in that: The SIoU loss function is composed of angle loss △, distance loss The shape loss Ω and IoU are composed, and the SIoU formula is as follows: The distance loss The formula is as follows: in, c w , c h The difference between the horizontal and vertical coordinates of the center point of the predicted frame and the center point of the real frame, respectively. α is the angle between the line connecting the center point of the predicted frame and the center point of the real frame and the horizontal line. (b cx ,b cy ) is the center coordinate of the prediction box, is the center coordinate of the real frame; The angle loss formula △ is as follows: in, x is the sine value of the angle α, which is normalized by get; The shape loss Ω formula is as follows: in, σ is the distance between the center point of the real box and the predicted box, θ is the parameter range used to control the degree of attention to shape loss; ω w and ω h is the normalized difference between width and height. The normalization logic is: the denominator takes the maximum value between the predicted value and the true value, ensuring that the difference value is between [0,1]. w and h are the width and height of the predicted box, w gt and h gt is the width and height of the real frame; The improved S_CIoU loss function formula is as follows: in, ρ 2 (b,b gt ) represents the Euclidean distance between the center point of the predicted box and the real box, c represents the diagonal distance of the minimum closure area that can contain both the predicted box and the real box. and Represent the aspect ratios of the true box and the predicted box respectively.

Citation Information

Cited By

  • Video key frame extraction method based on FD-SPnet network

    CN116310981A

  • A video key frame extraction method based on an FD-SPnet network

    CN116310981B

  • Attention guidance-based airport runway line detection method

    CN121170742A

  • Road crack segmentation method and device fusing double cross attention and variable kernel convolution

    CN121458983A