Airport bird detection method based on improved YOLOv8 algorithm

By improving the bird detection method of the YOLOv8 algorithm, the existing airport bird detection equipment has solved the problem of poor real-time and low accuracy, and efficient and accurate detection of birds in low-altitude areas of the airport has been achieved, and the airport safety management capabilities have been improved.

CN119964203APending Publication Date: 2025-05-09CIVIL AVIATION UNIV OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137249.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing airport bird detection equipment has problems such as poor real-time and low accuracy, especially the number and specific shape of birds in low-altitude areas are difficult to accurately detect.

Method used

The bird detection method based on the improved YOLOv8 algorithm is adopted. By obtaining the image to be detected at the airport and inputting the bird detection model of the improved YOLOv8 algorithm for detection, the bird detection results are output. The model includes backbone network, intermediate layer and detection head. Through technical means such as multi-layer convolution and attention mechanism, image features are extracted and fused to improve detection accuracy.

Benefits of technology

It improves the real-time and accuracy of airport bird detection, can more effectively monitor the number and shape of birds in low-altitude areas, and enhances the airport's safety management capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964203A_ABST
    Figure CN119964203A_ABST
Patent Text Reader

Abstract

The invention provides an airport bird detection method based on an improved YOLOv8 algorithm. The airport bird detection method comprises the steps of obtaining a to-be-detected image of an airport; inputting the to-be-detected image into a flying bird detection model based on an improved YOLOv8 algorithm, and carrying out flying bird detection; and outputting a bird detection result. According to the airport bird detection method based on the improved YOLOv8 algorithm, based on the improvement of the improved YOLOv8 algorithm, the accuracy and the detection speed in the airport bird detection process are improved, and bird detection is efficiently completed in the airport environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an airport bird detection method based on an improved YOLOv8 algorithm. Background Art

[0002] At present, bird strikes have always been a difficult problem for the civil aviation industry. With the vigorous development of the civil aviation industry and the surge in passenger traffic, the frequency of bird strikes has continued to increase worldwide, and the safety of the aircraft operating environment has received great attention worldwide. In particular, effective supervision of aircraft runways and areas below 60 meters vertically is an important problem that needs to be solved urgently to ensure civil aviation safety. It is of great significance to effectively monitor the flying birds in the low-altitude areas of airports to avoid birds threatening civil aviation safety.

[0003] Among the methods for detecting birds at airports, manual observation has defects such as poor real-time performance and low accuracy. Acoustic detection equipment is easily disturbed by environmental noise, and the effective detection distance is too short. Airports can use radar monitoring equipment for target detection, but due to the limitation of the resolution ability of radar equipment, it is difficult for airport radar to distinguish the number and specific forms of birds in low-altitude areas. With the development of deep learning, convolutional neural networks have made great achievements in image recognition. The use of deep learning convolutional neural networks for airport bird detection is expected to become a new type of airport bird detection technology due to its rapidity and accuracy. Summary of the invention

[0004] One of the purposes of the present invention is to provide an airport bird detection method based on an improved YOLOv8 algorithm to solve the problems of poor real-time performance and low accuracy of existing airport equipment in bird detection.

[0005] An embodiment of the present invention provides an airport bird detection method based on an improved YOLOv8 algorithm, comprising:

[0006] Obtain the image to be detected of the airport;

[0007] Input the image to be detected into the bird detection model based on the improved YOLOv8 algorithm to perform bird detection;

[0008] Output the bird detection results.

[0009] Optionally, the bird detection model based on the improved YOLOv8 algorithm includes: a backbone network, an intermediate layer and a detection head.

[0010] Optionally, the backbone network includes: a P1 layer, a P2 layer, a P3 layer, a P4 layer, a P5 layer, a P6 layer, a P7 layer, a P8 layer, a P9 layer, a P10 layer, a P11 layer, a P12 layer, and a P13 layer connected in sequence;

[0011] Among them, the P1 layer includes: the first Conv module;

[0012] The P2 layer includes: the second Conv module;

[0013] The P3 layer includes: a first MF module;

[0014] The P4 layer includes: the third Conv module;

[0015] The P5 layer includes: the first EMA module;

[0016] The P6 layer includes: a second MF module;

[0017] The P7 layer includes: the fourth Conv module;

[0018] The P8 layer includes: the second EMA module;

[0019] The P9 layer includes: the third MF module;

[0020] The P10 layer includes: the fifth Conv module;

[0021] The P11 layer includes: the fourth MF module;

[0022] The P12 layer includes: FocalModulation module;

[0023] The P13 layer includes: the third EMA module;

[0024] Among them, the convolution kernel size of the first Conv module, the second Conv module, the third Conv module, the fourth Conv module, and the fifth Conv module are all 3×3 and the step size is all 2;

[0025] The convolution kernel size of the first MF module, the second MF module, the third MF module, and the fourth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

[0026] Optionally, the intermediate layer includes: a first Upsample module, a first Concat module, a fifth MF module, a second Upsample module, a second Concat module, a sixth MF module, a sixth Conv module, a third Concat module, a seventh MF module, a seventh Conv module, a fourth Concat module, and an eighth MF module connected in sequence;

[0027] Among them, the convolution kernel size of the sixth Conv module and the seventh Conv module is the same as 3×3, and the step size is the same as 2;

[0028] The convolution kernel size of the fifth MF module, the sixth MF module, the seventh MF module, and the eighth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

[0029] Optionally, the detection head includes: a first Detect module, a second Detect module, and a third Detect module.

[0030] Optionally, the pre-training steps for the bird detection model based on the improved YOLOv8 algorithm are as follows:

[0031] Obtain a sample set; wherein the sample set includes a large number of groups of one-to-one corresponding historical airport images and flying bird labels;

[0032] The sample set is randomly divided into training set, validation set and test set in a ratio of 7:2:1;

[0033] Based on the improved YOLOv8 algorithm, training is performed according to the training set to obtain a training model;

[0034] Verify the training model based on the validation set;

[0035] Based on the verification results, the training model is tuned;

[0036] Test the tuned training model based on the test set;

[0037] The trained model that passes the test is used as a flying bird detection model based on the improved YOLOv8 algorithm.

[0038] Optionally, after the image to be detected is input into the bird detection model based on the improved YOLOv8 algorithm, the first Conv module of the P1 layer and the second Conv module of the P2 layer sequentially perform preliminary global feature extraction operations on the image to be detected, the first MF module of the P3 layer extracts the characteristics of the bird target and the surrounding environment information from the extracted preliminary features to form an initial feature map, the third Conv module of the P4 layer convolves the initial feature map, the first EMA module in the P5 layer emphasizes the spatial feature information of the bird in the initial feature map after the convolution, and the second MF module in the P6 layer extracts global features from the initial feature map after the spatial feature information of the bird is emphasized to obtain the feature map C1;

[0039] The fourth Conv module in the P7 layer convolves the feature map C1, the second EMA module in the P8 layer emphasizes the small target position information ignored by the shallow layer in the convolved feature map C1, the third MF module in the P9 layer extracts global feature information from the convolved feature map C1 after emphasizing the small target position information ignored by the shallow layer, the fifth Conv module in the P10 layer convolves the extracted global feature information, and the fourth MF module in the P11 layer performs global information feature equalization on the convolved global feature information to obtain the feature map C2;

[0040] The FocalModulation module in the P12 layer performs background denoising on the feature map C2, and the third EMA module in the P13 layer generates a feature map C3 with the minimum granularity based on the feature map C2 after background denoising;

[0041] The first Upsample module in the middle layer upsamples the feature map C3. The first Concat module fuses the upsampled feature map C3 with the feature map C2 to obtain the first fused feature map. The fifth MF module extracts global features from the first fused feature map to obtain feature maps C6 and N1. The first Detect module in the detection head performs large target flying bird detection based on feature map C6.

[0042] Optionally, the second Upsample module upsamples the N1 feature map, the second Concat module fuses the upsampled N1 feature map with the feature map C1 to obtain a second fused feature map, the sixth MF module performs global feature extraction on the second fused feature map to obtain feature map C5 and N2 feature map, and the second Detect module performs target flying bird detection based on the feature map C5.

[0043] Optionally, the sixth Conv module performs a convolution operation on the N2 feature map, the third Concat module fuses the N2 feature map after the convolution operation with the feature map C6 to obtain the N3 feature map, the seventh MF module performs global feature extraction on the N3 feature map to obtain the feature map to be convolved, the seventh Conv module performs a convolution operation on the feature map to be convolved, the fourth Concat module fuses the feature map to be convolved after the convolution operation with the feature map C3 to obtain a third fused feature map, the eighth MF module performs global feature extraction on the third fused feature map to obtain the feature map C4, and the third Detect module performs small target flying bird detection based on the feature map C4.

[0044] Optionally, based on the improved YOLOv8 algorithm, when training according to the training set, the momentum stochastic gradient descent method is used, the optimizer uses SGD, and the polynomial learning decay strategy is used. The loss function L used is:

[0045] L=L DFL +L CIoU +L CLS

[0046] Where L CLS is the classification loss, L CIoU is the bounding box IoU loss, L DFL is the bounding box regression loss.

[0047] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0048] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0050] Figure 1 Schematic diagram of an airport bird detection method based on an improved YOLOv8 algorithm in an embodiment of the present invention;

[0051] Figure 2 This is a schematic diagram of the algorithm network structure in an embodiment of the present invention;

[0052] Figure 3 This is a structural diagram of the EMA attention mechanism in an embodiment of the present invention;

[0053] Figure 4 is a structural diagram of FocalModulation in an embodiment of the present invention;

[0054] Figure 5 is a structural diagram of the MF module in an embodiment of the present invention;

[0055] Figure 6 Schematic diagram of the bird detection effect in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0057] The embodiment of the present invention provides an airport bird detection method based on an improved YOLOv8 algorithm. Figure 1 As shown, including:

[0058] S1: Obtain the image to be detected of the airport;

[0059] S2: Input the image to be detected into the bird detection model based on the improved YOLOv8 algorithm to perform bird detection;

[0060] S3: Output the bird detection result.

[0061] The flying bird detection model based on the improved YOLOv8 algorithm includes: a backbone network, an intermediate layer and a detection head.

[0062] The backbone network includes: a P1 layer, a P2 layer, a P3 layer, a P4 layer, a P5 layer, a P6 layer, a P7 layer, a P8 layer, a P9 layer, a P10 layer, a P11 layer, a P12 layer, and a P13 layer connected in sequence;

[0063] Among them, the P1 layer includes: the first Conv module;

[0064] The P2 layer includes: the second Conv module;

[0065] The P3 layer includes: a first MF module;

[0066] The P4 layer includes: the third Conv module;

[0067] The P5 layer includes: the first EMA module;

[0068] The P6 layer includes: a second MF module;

[0069] The P7 layer includes: the fourth Conv module;

[0070] The P8 layer includes: the second EMA module;

[0071] The P9 layer includes: the third MF module;

[0072] The P10 layer includes: the fifth Conv module;

[0073] The P11 layer includes: the fourth MF module;

[0074] The P12 layer includes: FocalModulation module;

[0075] The P13 layer includes: the third EMA module;

[0076] Among them, the convolution kernel size of the first Conv module, the second Conv module, the third Conv module, the fourth Conv module, and the fifth Conv module are all 3×3 and the step size is all 2;

[0077] The convolution kernel size of the first MF module, the second MF module, the third MF module, and the fourth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

[0078] The middle layer includes: a first Upsample module, a first Concat module, a fifth MF module, a second Upsample module, a second Concat module, a sixth MF module, a sixth Conv module, a third Concat module, a seventh MF module, a seventh Conv module, a fourth Concat module, and an eighth MF module connected in sequence;

[0079] Among them, the convolution kernel size of the sixth Conv module and the seventh Conv module is the same as 3×3, and the step size is the same as 2;

[0080] The convolution kernel size of the fifth MF module, the sixth MF module, the seventh MF module, and the eighth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

[0081] The detection head includes: a first Detect module, a second Detect module, and a third Detect module.

[0082] The parameters of the Conv layer are the same, the convolution kernel is 3 and the stride is 2.

[0083] The parameters of the MF module are the same, with a dilation rate of 1, a convolution kernel of 3, and a stride of 2.

[0084] Upsample is an upsampling operation that can enlarge the width and height of the feature map.

[0085] Contact is a concatenation operation that concatenates two feature maps with the same width and height, and is used in conjunction with Upsample.

[0086] The pre-training steps of the bird detection model based on the improved YOLOv8 algorithm are as follows:

[0087] Obtain a sample set; wherein the sample set includes a large number of groups of one-to-one corresponding historical airport images and flying bird labels;

[0088] The sample set is randomly divided into training set, validation set and test set in a ratio of 7:2:1;

[0089] Based on the improved YOLOv8 algorithm, training is performed according to the training set to obtain a training model;

[0090] Verify the training model based on the validation set;

[0091] Based on the verification results, the training model is tuned;

[0092] Test the tuned training model based on the test set;

[0093] The trained model that passes the test is used as a flying bird detection model based on the improved YOLOv8 algorithm.

[0094] During training, the sample set is randomly divided into a training set, a validation set and a test set in a ratio of 7:2:1; based on the improved YOLOv8 algorithm, the optimal weight of the model proposed by the present invention is obtained according to the training set, the validation set and the test set; based on the training set, the weight, the bias value, the regularization parameter and the like in the model training process are adjusted by forward feedback through the size of the loss function value; based on the verification result, the generalization ability of the model for unknown parameters is determined, and the weight trained in this round is fine-tuned; based on the test set, the weight of the model finally trained is tested; the trained model and the generated weight are used as a flying bird detection model based on the improved YOLOv8 algorithm.

[0095] After the image to be detected is input into the bird detection model based on the improved YOLOv8 algorithm, the first Conv module of the P1 layer and the second Conv module of the P2 layer perform preliminary global feature extraction operations on the image to be detected, such as Figure 5 As shown in the red box in (c), the first MF module of the P3 layer extracts the features of the flying bird target and the surrounding environment information from the extracted preliminary features to form an initial feature map, as shown in Figure 5 As shown in the purple and yellow boxes in (b), the third Conv module in the P4 layer convolves the initial feature map, the first EMA module in the P5 layer emphasizes the spatial feature information of the bird in the initial feature map after convolution, improves the network's attention to the target and reduces the attention to the background, and the second MF module in the P6 layer extracts the global features of the initial feature map after emphasizing the spatial feature information of the bird, and obtains the feature map C1;

[0096] The fourth Conv module in the P7 layer convolves the feature map C1, but the feature map of C1 has a large attention scale, and the partially occluded small target bird is easily missed. The second EMA module in the P8 layer emphasizes the position information of the small target ignored in the shallow layer in the convolutional feature map C1 (that is, the second EMA attention mechanism module in the P8 layer performs deep mining on the convolutional feature map C1 to improve the network's attention to the small targets missed in the upper layer step). The third MF module in the P9 layer extracts global feature information from the convolutional feature map C1 after emphasizing the position information of the small targets ignored in the shallow layer. The fifth Conv module in the P10 layer convolves the extracted global feature information. The fourth MF module in the P11 layer performs global information feature balance on the convolutional global feature information to obtain the feature map C2 (that is, the fourth MF module in the P11 layer performs feature fusion of convolution and dilation convolution on the convolutional global feature information again. At this time, the network obtains the minimum scale feature map C2 of the image, which captures the occluded target features in the image to the greatest extent).

[0097] The FocalModulation module in the P12 layer performs background denoising on the feature map C2, and the third EMA module in the P13 layer generates a feature map C3 with the minimum granularity based on the feature map C2 after background denoising;

[0098] The first Upsample module in the middle layer upsamples the feature map C3. The first Concat module fuses the upsampled feature map C3 with the feature map C2 to obtain the first fused feature map. The fifth MF module extracts global features from the first fused feature map to obtain feature maps C6 and N1. The first Detect module in the detection head performs large target flying bird detection based on feature map C6.

[0099] The second Upsample module upsamples the N1 feature map, the second Concat module fuses the upsampled N1 feature map with the feature map C1 to obtain a second fused feature map, the sixth MF module performs global feature extraction on the second fused feature map to obtain feature map C5 and N2 feature map, and the second Detect module performs target flying bird detection based on feature map C5.

[0100] The sixth Conv module performs a convolution operation on the N2 feature map. The third Concat module fuses the N2 feature map after the convolution operation with the feature map C6 to obtain the N3 feature map. The seventh MF module performs a global feature extraction on the N3 feature map to obtain the feature map to be convolved. The seventh Conv module performs a convolution operation on the feature map to be convolved. The fourth Concat module fuses the feature map to be convolved after the convolution operation with the feature map C3 to obtain the third fused feature map. The eighth MF module performs a global feature extraction on the third fused feature map to obtain the feature map C4. The third Detect module performs small target flying bird detection based on the feature map C4.

[0101] Based on the improved YOLOv8 algorithm, when training according to the training set, the momentum stochastic gradient descent method is used, the optimizer uses SGD, and the polynomial learning attenuation strategy is used. The loss function used is:

[0102] L=L DFL +L CIoU +L CLS

[0103] Where L CLS is the classification loss, L CIoU is the bounding box IoU loss, L DFL is the bounding box regression loss.

[0104] In the above technical solution, firstly, images of the airport and its surroundings are collected, the location information of the flying birds in each image is annotated and a label file is generated (the label file of each image contains the location of the flying bird in the image, usually the coordinates of the bounding box), and the collected images are used as sample images and the label files to form a sample set;

[0105] In the sample set, sample images and corresponding label files are randomly divided into training set, validation set and test set in a ratio of 7:2:1. The model uses the training set as the main data source for model training, including target features and label data. The main purpose is to adjust internal parameters through the training set and make reasonable predictions for new input samples; the validation set is used to evaluate the performance of the model during the training process. Its main function is to help identify whether the model is overfitting and provide a basis for hyperparameter tuning; the test set is independent of the training set and validation set, and does not participate in the training and tuning process. Its main purpose is to evaluate the final performance of the model and measure the performance of the model on unknown images.

[0106] The sample image is input into the network model and first scaled to a 640*640 standard image and two Conv modules and one MF module. The Conv module uses a 3*3 convolution kernel with a step size of 2. After each Conv module, the feature map size is reduced by half. These two Conv blocks are responsible for the preliminary extraction of P1 and P2 layer features; the MF module uses a 3*3 convolution kernel and a 3*3 convolution kernel with an expansion rate of 1 to extract the features of the flying bird target and the surrounding environment information respectively, and finally the extracted feature maps are fused through the Concat layer.

[0107] After the feature map is reduced again by the 3*3 convolution of the P4 layer, the EMA attention mechanism of the P5 layer begins to emphasize the spatial feature information of the bird for the first time. The EMA attention mechanism does not change the size of the feature map. The subsequent feature map is sent to the MF module of the P6 layer for global feature extraction. At this time, the maximum feature map C1 of 80*80 is obtained.

[0108] To further detect small targets in the image, the 3*3 Conv of the P7 layer convolves the feature map again to generate a 40*40 feature image. The EMA attention mechanism of the P8 layer emphasizes the position information of small targets ignored by the shallow layer, and sends the feature map to the MF module of the P9 layer to extract global feature information. At this time, the backbone network performs the next 3*3 Conv operation, and the generated feature map is balanced by the MF module for global information features to obtain the second 20*20 feature map C2. C2 is input to the FocalModulation of the backbone network for background denoising, and the output feature map is then passed through the EMA attention mechanism to generate the minimum granularity feature map C3.

[0109] After upsampling, the C3 feature map is enlarged to 40*40, and features are fused with the C2 feature map. After passing through the MF module, it enters the Detect stage. The feature map C6 generated by it is used to detect larger flying birds. In the process, the N1 feature map generated by the MF module is enlarged to the same size as C1 through upsampling, and features are extracted by the MF module and enter the Detect stage. The feature map C5 generated by the process is used to detect medium-sized flying bird targets. The N3 feature map is fused with the C3 feature map through the MF module and Conv convolution operation, and the C4 feature map is generated through the MF block. After passing through the Detect stage, it detects small flying bird targets.

[0110] The network is trained using the momentum stochastic gradient descent method, the optimizer uses SGD, and the polynomial learning decay strategy is used. The loss function L used is:

[0111] L=L DFL +L CIoU +L CLS

[0112] Among them, L CLS is the classification loss, L CIoU is the bounding box IoU loss, L DFL is the bounding box regression loss;

[0113] The trained network is used for the actual bird detection task at the airport. The loss function consists of three parts: classification loss, bounding box loss, and bounding box regression loss. The goal is to optimize these loss functions so that the model can effectively predict the category and position of the bird.

[0114] Each EMA module has a built-in EMA attention mechanism, such as Figure 3 As shown, this is the EMA attention mechanism structure diagram, which for any given input feature map X∈R C×H×W , take G<<C, EMA divides the feature map X into G sub-features X=[X0,X1,...,X G-1 ], X i ∈R C×H×W , where R represents the input set, C represents the number of input channels, H and W represent the height and width of the input feature space dimension respectively, and X i is [X0,X1,...,X G-1] The cross-space learning part is the attention weights extracted by two 1×1 convolution paths and one 3×3 convolution path. In order to achieve richer feature aggregation, the cross-channel interaction modeling of information in two directions is carried out. The output of two 1×1 branches and the output of 3×3 branches are introduced, and then the global spatial information output by the 1×1 branch is encoded using 2D global average pooling. The channel features of the output of the smallest branch are converted into the dimensional shape matching the 3×3 branch, that is, R 1×C / / G ×R C / / G×HW . C / / G represents the number of channels occupied by each sub-feature.

[0115] The FocalModulation module structure is as follows Figure 4 As shown in , the process of this module consists of two steps: hierarchical context extraction from local to global scope at different granularity levels, and gated aggregation of the obtained context features compressed to the modulator at different granularity, as shown in Figure 5 , 6 As shown in the figure. Given an input feature X, it is projected into the new feature space of the linear layer to obtain the context level representation of L+1 depth convolution directions, and finally sent to the modulator. This part obtains the spatial and level-aware weights through the linear layer, and then performs element-by-element multiplication to perform the weighted sum, and finally obtains a single feature map of the same size as the input X.

[0116] The structure of each MF module is as follows Figure 5 As shown in Figure 1, the MF block contains two main steps: feature learning step and feature improvement step. In the feature learning part, it contains a local feature extractor Gloc(*) and a context feature extractor Gsur(*), which are used to learn local features and surrounding background respectively. Gloc(*) is instantiated as a 3×3 standard convolutional layer, which learns local features from 8 adjacent feature vectors, corresponding to Figure 5 The purple area in (a). Gsur (*) is instantiated as a 3×3 dilated convolution with a dilation rate of 1. The larger receptive field effectively learns contextual environmental features. Figure 5 The yellow area in (b) corresponds to this. To reduce the amount of computation, this paper uses normalization (BN) and activation function (SILU) as the connection layer to output joint features for the context environment and local features of Gsur(*) and Gloc(*). In the feature improvement part, the global context extractor Gglo(*) extracts global environment information to improve the joint features. Gglo(*) is instantiated as a global average pooling layer and a multi-layer perceptron. The global average pooling layer can efficiently aggregate scene images, corresponding to Figure 5The red area in (c) of the image. The multi-layer perceptron consists of two fully connected layers to further extract the global contextual relationship. Finally, a scaling layer is used to re-weight and fuse the features using the extracted global contextual relationship. The row-by-row feature extraction adopted by the MF block can extract more multi-scale contextual information from the joint features, thereby improving the detection accuracy of the target.

[0117] like Figure 6 As shown, the present invention reduces missed detection and repeated detection, indicating that the algorithm of the present application has good detection performance.

[0118] This application has certain effectiveness in the detection of flying birds in airport bird defense. For the problem of insufficient feature extraction ability in the detection process, the concept of multi-feature fusion block is proposed. Its feature extraction fusion of the target, the surrounding of the target, and the global can obtain the flying bird features to the greatest extent; by introducing the attention mechanism, the network's attention to flying birds is improved, and the FocalModulation module is used to enhance the network's processing ability for different bird scale feature information. Based on the improvement of the improved YOLOv8 algorithm, the present invention improves the accuracy and detection speed of the airport flying bird detection process, and realizes the efficient completion of flying bird detection in the airport environment.

[0119] Table 1 compares the simulation results of the method of the present invention (MF-YOLO) with YOLOv5, YOLOv6, YOLOv7, and YOLOv8:

[0120] Table 1 Simulation experiment results data

[0121]

[0122] From Table 1, it can be seen that the method proposed in the present invention has higher accuracy and smaller computational complexity compared with other algorithms. Compared with other algorithms, the accuracy is improved by 2.3%, 6.2%, 5%, and 2.3%, the recall rate is improved by 1.6%, 12.1%, 1.7%, and 2%, respectively, and the computational complexity is reduced by 10.4G, 5.24G, 6.9G, and 2G, respectively, indicating that the method of the present invention has better accuracy while improving the computational efficiency.

[0123] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for detecting birds at an airport based on an improved YOLOv8 algorithm, characterized in that: include: Obtain the image to be detected of the airport; Input the image to be detected into the bird detection model based on the improved YOLOv8 algorithm to perform bird detection; Output the bird detection results.

2. The airport bird detection method based on the improved YOLOv8 algorithm as claimed in claim 1, characterized in that: The flying bird detection model based on the improved YOLOv8 algorithm includes: a backbone network, an intermediate layer and a detection head.

3. The airport bird detection method based on the improved YOLOv8 algorithm as claimed in claim 2, characterized in that: The backbone network includes: a P1 layer, a P2 layer, a P3 layer, a P4 layer, a P5 layer, a P6 layer, a P7 layer, a P8 layer, a P9 layer, a P10 layer, a P11 layer, a P12 layer, and a P13 layer connected in sequence; Among them, the P1 layer includes: the first Conv module; The P2 layer includes: the second Conv module; The P3 layer includes: a first MF module; The P4 layer includes: the third Conv module; The P5 layer includes: the first EMA module; The P6 layer includes: a second MF module; The P7 layer includes: the fourth Conv module; The P8 layer includes: the second EMA module; The P9 layer includes: the third MF module; The P10 layer includes: the fifth Conv module; The P11 layer includes: the fourth MF module; The P12 layer includes: FocalModulation module; The P13 layer includes: the third EMA module; Among them, the convolution kernel size of the first Conv module, the second Conv module, the third Conv module, the fourth Conv module, and the fifth Conv module are all 3×3 and the step size is all 2; The convolution kernel size of the first MF module, the second MF module, the third MF module, and the fourth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

4. The airport bird detection method based on the improved YOLOv8 algorithm as claimed in claim 3, characterized in that: The middle layer includes: a first Upsample module, a first Concat module, a fifth MF module, a second Upsample module, a second Concat module, a sixth MF module, a sixth Conv module, a third Concat module, a seventh MF module, a seventh Conv module, a fourth Concat module, and an eighth MF module connected in sequence; Among them, the convolution kernel size of the sixth Conv module and the seventh Conv module is the same as 3×3, and the step size is the same as 2; The convolution kernel size of the fifth MF module, the sixth MF module, the seventh MF module, and the eighth MF module are all 3×3, the dilation rate is all 1, and the step size is all 2.

5. The airport bird detection method based on the improved YOLOv8 algorithm as claimed in claim 4, characterized in that: The detection head includes: a first Detect module, a second Detect module, and a third Detect module.

6. The airport bird detection method based on the improved YOLOv8 algorithm as claimed in claim 1, characterized in that: The pre-training steps of the bird detection model based on the improved YOLOv8 algorithm are as follows: Obtain a sample set; wherein the sample set includes a large number of groups of one-to-one corresponding historical airport images and flying bird labels; The sample set is randomly divided into training set, validation set and test set in a ratio of 7:2:1; Based on the improved YOLOv8 algorithm, training is performed according to the training set to obtain a training model; Verify the training model based on the validation set; Based on the verification results, the training model is tuned; Test the tuned training model based on the test set; The trained model that passes the test is used as a flying bird detection model based on the improved YOLOv8 algorithm.

7. The method for detecting birds at an airport based on the improved YOLOv8 algorithm according to claim 5, characterized in that: After the image to be detected is input into the bird detection model based on the improved YOLOv8 algorithm, the first Conv module of the P1 layer and the second Conv module of the P2 layer perform preliminary global feature extraction operations on the image to be detected in turn, the first MF module of the P3 layer extracts the characteristics of the bird target and the surrounding environment information from the extracted preliminary features to form an initial feature map, the third Conv module of the P4 layer performs convolution on the initial feature map, the first EMA module in the P5 layer emphasizes the spatial feature information of the bird in the initial feature map after convolution, and the second MF module in the P6 layer extracts global features from the initial feature map after emphasizing the spatial feature information of the bird, and obtains the feature map C1; The fourth Conv module in the P7 layer convolves the feature map C1, the second EMA module in the P8 layer emphasizes the small target position information ignored by the shallow layer in the convolved feature map C1, the third MF module in the P9 layer extracts global feature information from the convolved feature map C1 after emphasizing the small target position information ignored by the shallow layer, the fifth Conv module in the P10 layer convolves the extracted global feature information, and the fourth MF module in the P11 layer performs global information feature equalization on the convolved global feature information to obtain the feature map C2; The FocalModulation module in the P12 layer performs background denoising on the feature map C2, and the third EMA module in the P13 layer generates a feature map C3 with the minimum granularity based on the feature map C2 after background denoising; The first Upsample module in the middle layer upsamples the feature map C3, the first Concat module fuses the upsampled feature map C3 with the feature map C2 to obtain the first fused feature map, the fifth MF module performs global feature extraction on the first fused feature map to obtain feature maps C6 and N1 feature maps, and the first Detect module in the detection head performs large target flying bird detection based on the feature map C6.

8. The method for detecting birds at an airport based on the improved YOLOv8 algorithm according to claim 7, characterized in that: The second Upsample module upsamples the N1 feature map, the second Concat module fuses the upsampled N1 feature map with the feature map C1 to obtain a second fused feature map, the sixth MF module performs global feature extraction on the second fused feature map to obtain feature map C5 and N2 feature map, and the second Detect module performs target flying bird detection based on feature map C5.

9. The method for detecting birds at an airport based on the improved YOLOv8 algorithm according to claim 8, characterized in that: The sixth Conv module performs a convolution operation on the N2 feature map. The third Concat module fuses the N2 feature map after the convolution operation with the feature map C6 to obtain the N3 feature map. The seventh MF module performs a global feature extraction on the N3 feature map to obtain the feature map to be convolved. The seventh Conv module performs a convolution operation on the feature map to be convolved. The fourth Concat module fuses the feature map to be convolved after the convolution operation with the feature map C3 to obtain the third fused feature map. The eighth MF module performs a global feature extraction on the third fused feature map to obtain the feature map C4. The third Detect module performs small target flying bird detection based on the feature map C4.

10. The method for detecting birds at an airport based on an improved YOLOv8 algorithm as claimed in claim 6, characterized in that: Based on the improved YOLOv8 algorithm, when training according to the training set, the momentum stochastic gradient descent method is used, the optimizer uses SGD, and the polynomial learning attenuation strategy is used. The loss function L used is: L=L DFL +L CIoU +L CLS Where L CLS is the classification loss, L CIoU is the bounding box IoU loss, L DFL is the bounding box regression loss.