Detection method of airport runway foreign object detection system (FOD)
By improving the YOLOv7 model and combining super-resolution and target tracking technology, the accuracy and stability of small and medium-sized object detection of external objects at airport runways is solved, and high-precision and real-time external objects detection are achieved to ensure the safety of the aircraft.
Patent Information
- Application Number
- CN202510441414.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing airport runway detection methods rely on manual patrols to have the risk of missed inspection, and the automation system has low accuracy and stability when detecting small targets, which is particularly prone to missed inspections or missed inspections, affecting aircraft safety.
The improved YOLOv7 object detection model is adopted to combine the super-resolution branch structure and the object tracking algorithm. By building a dedicated data set, optimizing the loss function and introducing a Kalman filter, the accuracy of small object detection is improved and repeated reporting is avoided.
It significantly improves the detection accuracy of small targets on the runway, reduces the false detection rate, achieves high-precision and real-time detection of foreign objects, and improves the level of airport safety management.
Smart Images

Figure CN120388330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a detection method of an airport runway foreign object detection (FOD) system. Background Art
[0002] Foreign Object Debris (FOD) detection on airport runways is crucial for ensuring safe aircraft takeoff and landing. FOD on the runway, such as birds, debris, garbage, and other obstacles, can pose a significant threat to aircraft safety if not promptly detected and removed. Especially during the critical moments of takeoff and landing, the impact of FOD on the aircraft cannot be ignored and can even lead to damage, major accidents, or catastrophic consequences. Therefore, effective, timely, and accurate FOD detection on airport runways has become a critical issue that needs to be addressed in airport safety management.
[0003] Traditional methods for detecting foreign objects on airport runways rely primarily on manual inspections. While this method can ensure runway safety to a certain extent, it has significant shortcomings. First, manual inspections are limited by inspection frequency and personnel oversight, resulting in a high risk of missed inspections. Second, each inspection requires temporary runway closure, which seriously impacts airport operational efficiency. This is especially true during peak flight times, where prolonged runway closures can pose a significant risk to aviation safety. Therefore, improving the automation and real-time performance of foreign object detection has become a critical issue that needs to be addressed urgently.
[0004] Currently, existing automated FOD detection systems fall into three main categories: radar detection alone, radar combined with video data, and video image recognition alone. Radar systems offer strong penetration in FOD detection, detecting targets based on reflected signals and achieving high accuracy, but they are also prone to numerous false detections. Therefore, regardless of the technology employed, relying on video image data for correction and auxiliary verification remains an essential step.
[0005] Although existing FOD detection systems have incorporated image recognition technology, image data processing still faces numerous challenges. The locations of FOD are highly random, and numerous small targets exist. These targets are typically less than 20 pixels wide and high in a standard 1920×1080 resolution image. Due to the small size, irregular shape, and complex backgrounds of these targets, traditional video image recognition algorithms have low accuracy and stability when detecting such small targets, and are particularly prone to false detections or missed detections. Therefore, improving the detection accuracy and reducing the false detection rate of image-based target detection algorithms has become a core technical challenge facing current FOD detection systems. Summary of the Invention
[0006] The object of the present invention is to provide an artificial intelligence-based foreign object detection method and system for airport runways. By improving the YOLOv7 model and combining super-resolution and target tracking technologies, high-precision and real-time detection of foreign objects on the runway is achieved, thereby improving the airport safety management level.
[0007] The technical solution of the present invention is to provide a detection method for a foreign object detection system (FOD) for airport runways, and the method includes:
[0008] S1. Construct a dedicated dataset suitable for the foreign object detection task on the airport runway. First, use a high-resolution device to capture image data of the constructed simulated runway scene and various foreign objects. After annotating the foreign objects in the images, perform preprocessing to segment the training set and the test set;
[0009] S2. Build a YOLOv7 object detection model, and introduce a super-resolution branch structure to enhance the detection accuracy of the model for small targets. Then optimize the loss function of the model. The optimized loss function consists of two parts: the object detection loss and the super-resolution reconstruction loss;
[0010] S3. Introduce a target tracking algorithm to avoid duplicate reporting of the same foreign object. Connect the output of the YOLOv7 object detection model to a target detector, and perform subsequent processing on the detection results through the target tracking algorithm;
[0011] S4. Input the training set into the YOLOv7 network for training. After training is completed, output the new image to be detected to the YOLOv7 network model, and the target detector outputs the detection results to confirm whether there are foreign objects on the runway.
[0012] In any of the above technical solutions, further, the structure of the YOLOv7 object detection model built in step S2 includes: an input layer, a backbone network, a super-resolution branch, and a detection head;
[0013] The input layer receives the input image and transfers it to the backbone network. The backbone network uses a deep convolutional neural network CNN to extract the features of the input image; the purpose of the backbone network is to extract low-level features and high-level features from the image, and gradually abstract the content of the image through multiple convolutional operations; the low-level features come from the output of the second ELAN module in the backbone network, and the high-level features come from the output of the last ELAN module in the backbone network; the extracted low-level and high-level features respectively contain the detailed information and overall semantic information of the target, and have different spatial and semantic levels;
[0014] The super-resolution branch processes the low-resolution feature map extracted from the backbone network, restores the image details, and improves the detection accuracy of small targets;
[0015] The detection head receives the feature maps output from the second to the fourth ELAN modules of the backbone network, and after processing, outputs three key results: the class prediction of the target, the bounding box regression of the target, and the confidence prediction.
[0016] In any of the above technical solutions, further, in the super-resolution branch, the low-level features and the high-level features are first fused through an encoder. The encoder uses a convolutional-ReLU module to process the low-level features and simultaneously upsamples the high-level features. After these processes, the low-level features and the high-level features are combined together through a concatenation operation, and then further processed by convolution and the ReLU activation function to enhance the expressive ability of the features.
[0017] In the decoder part, the decoder receives the fused features output by the encoder, performs a convolution operation on the input features to extract more high-level features, then uses an upsampling operation to increase the spatial resolution of the feature map to the resolution of the low-level features. Finally, the upsampled feature map is processed through multiple convolutional layers to generate a high-resolution output map.
[0018] In any of the above technical solutions, further, in the training stage of the YOLOv7 object detection model, the features from the training set are enhanced through the super-resolution branch structure, and the model learns more detailed image details to improve the detection accuracy of the model. In the inference stage, the super-resolution branch structure is removed to accelerate the inference speed. At this time, the trained network can quickly generate high-precision detection results.
[0019] In any of the above technical solutions, further, in the loss function of the YOLOv7 object detection model, the object detection loss is used to evaluate the position, class, and existence of the object, while the super-resolution reconstruction loss measures the difference between the input image and the super-resolution output.
[0020] The improved loss function consists of the object detection loss L o and the super-resolution reconstruction loss L s The total loss L total The expression is:
[0021] L total = c1L o + c2L s ;
[0022] Where c1 and c2 are coefficients for balancing the object detection loss and the SR reconstruction loss.
[0023] The super-resolution reconstruction loss L s is used to measure the difference between the input image X and the super-resolution result S, and is specifically calculated through the L1 loss. The L s The expression is:
[0024] L s = ||S - X||1;
[0025] Where S is the output after super - resolution reconstruction and X is the input image; The L1 loss can effectively measure the pixel - level difference, thus ensuring that the details of the reconstructed image are preserved;
[0026] The object detection loss L o includes the object existence loss L obj , the position loss L loc and the classification loss L cls , and its overall expression is:
[0027]
[0028] Where l represents the number of layers of the output layer; a l , b l and c l are the weight coefficients of the object position, object existence, and object classification losses at different layers respectively; λ loc , λ obj and λ cls are hyperparameters used to adjust the relative importance of each part of the loss.
[0029] In any of the above - mentioned technical solutions, further, the object tracking algorithm in step S3 specifically includes:
[0030] First, the real - time video stream is passed as input to the object detector, which analyzes each frame of the image and outputs the detection results including detection boxes and corresponding confidence scores; According to two pre - set confidence thresholds τ high and τ low , the detection boxes are divided into a high - confidence detection box set D high and a low - confidence detection box set D low ; The high - confidence detection box set D higt includes detection boxes with confidence scores greater than τ high , and the low - confidence detection box set D low includes detection boxes with confidence scores between τ low and τ high ;
[0031] Subsequently, for the trajectory set T in the existing object trajectory set in the current frame, the Kalman filter is used to predict each trajectory to generate the predicted position of each trajectory in the next frame;
[0032] Next, the first association is performed between the high - confidence detection box set D high and the trajectory set T: Calculate the cost matrix between each detection box and the trajectory, and use the Hungarian algorithm for optimal matching;
[0033] For the successfully matched trajectories, update the state parameters of their Kalman filters and retain these trajectories in the trajectory set of the current frame; for the unmatched trajectories, store them in the unmatched trajectory set T remain ; meanwhile, store the unmatched detection boxes in the unmatched detection box set D remain ;
[0034] After the first association, perform a second association on the low-confidence detection box set and the unmatched trajectory set, and repeat the above optimal matching operation;
[0035] For the trajectories successfully matched through the second association, update the Kalman filter in the same way and add them to the current frame trajectory set; for the unmatched trajectories, classify them into the lost trajectory set T lost ; for the unmatched low-confidence detection boxes, directly delete them to avoid adverse effects of low-confidence detection results on subsequent processing;
[0036] For the lost trajectory set T lost if a certain trajectory in it has not been updated for more than 30 frames, delete it from the trajectory set to avoid excessive consumption of computing resources by the system;
[0037] Subsequently, for the unmatched detection box set D in the first association process remain , the system compares these detection boxes with the preset tracking threshold ε; if the confidence of a detection box is greater than this threshold and satisfies the condition in two consecutive frames, this detection box is initialized as a new target trajectory and added to the trajectory set, and at the same time, the Kalman filter is used to predict the position of the new trajectory in the next frame;
[0038] Finally, the system outputs the current frame trajectory set after the above multi-round matching and processing, and this set will be used for target detection and tracking in the next frame of image. The Kalman filter continuously predicts each target trajectory to ensure the tracking stability of the target in consecutive frames, so as to achieve the purpose of reporting a single foreign object only once.
[0039] In any of the above technical solutions, further, the content of labeling foreign objects in the image in step S1 includes: the bounding box of each foreign object, the category label of the foreign object, and its position information;
[0040] After the labeling is completed, save the data set in the PASCAL VOC format, and store the labeling information as an XML format file, and save the image file in the JPEG format; then convert the data labeling file in the PASCAL VOC format to the TXT file format required by YOLO.
[0041] The beneficial effects of the present invention are:
[0042] At long distances or in complex backgrounds, small targets are difficult to be accurately captured by traditional detection algorithms due to low resolution and unclear texture. The technical solution of the present invention improves the YOLOv7 model by introducing a super-resolution branch structure, effectively restoring the detailed information in low-resolution images, thereby significantly improving the detection accuracy of small targets (such as gravel, tools, screws, nuts, etc.) on the runway; compared with traditional object detection methods, the present invention improves the detection accuracy of small targets by more than 10%.
[0043] By combining object detection with the Bytetrack object tracking algorithm, the present invention can continuously track the same foreign object in consecutive frames, avoiding repeated reporting, thereby reducing the burden of redundant information processing in the system and improving the accuracy and real-time performance of the detection results; the Kalman filter is used to predict and update the target trajectory, making the target tracking more stable, ensuring the accurate positioning of the foreign object in consecutive image frames, and maintaining a high tracking accuracy even during movement, further enhancing the overall robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The above and additional aspects of the present invention will become obvious and easy to understand in the description of the embodiments in conjunction with the following drawings, where:
[0045] Figure 1 is a schematic diagram of an improved model structure of the detection method of a foreign object detection system (FOD) for airport runways according to an embodiment of the present invention;
[0046] Figure 2 is a flow chart of the combination of the object detection model and the tracking model of the detection method of a foreign object detection system (FOD) for airport runways according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0048] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0049] This embodiment provides a detection method for a foreign object detection system (FOD) for airport runways, and the method includes:
[0050] S1. Build a dataset: First, build a dedicated dataset suitable for the task of detecting foreign objects on airport runways. The construction of this dataset includes data collection, annotation, and preprocessing:
[0051] Data collection: Set up a simulated airport runway scene and use high-resolution imaging equipment to collect image data of various types of foreign objects; the categories of foreign objects include but are not limited to gravel, maintenance tools, screws, nuts, etc.; to ensure that the dataset has broad adaptability, the collected image data covers a variety of typical environmental conditions, such as runway scenes under different lighting conditions during the day and at night.
[0052] Data annotation: Use annotation tools (such as LabelImg) to perform object annotation on the collected image data; the annotation content includes the bounding box of each foreign object, the category label of the foreign object, and its location information; after annotation, save the dataset in the PASCAL VOC format and store the annotation information as an XML format file, and save the image files in the JPEG format.
[0053] Data preprocessing: To adapt to the YOLOv7 training framework, it is necessary to convert the data annotation files in the PASCAL VOC format to the TXT file format required by YOLO to ensure data format compatibility; in addition, data augmentation processing is also performed on the data, including operations such as image normalization, random cropping, horizontal flipping, and color adjustment, to enhance the robustness and adaptability of the model.
[0054] Divide the annotated dataset into an 80% training set and a 20% test set. During the division process, ensure that the training data and the test data are balanced and consistent in terms of category distribution.
[0055] S2. Improve the YOLOv7 object detection model: To improve the accuracy of the YOLOv7 model in small object detection tasks, the present invention improves the YOLOv7 detection algorithm, and the improvement content includes introducing a super-resolution branch structure and optimizing the loss function.
[0056] In the existing YOLOv7 object detection model, the feature extraction network of the image usually processes the input image through multiple convolutional layers, gradually reducing the size of the feature map. This processing method can capture the overall semantic information of the image, but when extracting the features of small objects, detailed information may be lost due to the decrease in image resolution; to improve the detection accuracy of small objects, the present invention introduces a super-resolution (SR for short) branch structure into the backbone network of the YOLOv7 model to restore the high-resolution features of the image and ensure that detailed information can be effectively retained.
[0057] As Figure 1As shown in the figure, the YOLOv7 object detection model structure provided by the present invention includes: an input layer, a backbone network, a super-resolution branch, and a detection head ( Figure 1 The content in the lower right white background box represents the specific composition of some modules).
[0058] Input layer: The input image received by the YOLOv7 model is first processed by the input layer, usually an image of a fixed size (such as 640×640) to meet the input requirements of the network. The main task of the input layer is to convert the original image data into a form that can be processed by the neural network.
[0059] Backbone network: In the present invention, the backbone network of YOLOv7 adopts a deep convolutional neural network (CNN) to extract the features of the input image. Its overall structure combines multiple convolutional modules with feature aggregation modules to achieve the gradual abstraction and refinement of low-level features (texture information) and high-level features (semantic information) in the image. Specifically, the backbone network first sends the input image into four initial stages composed of CBS (Convolution-BatchNorm-SiLU, convolution, batch normalization, activation) modules, which initially extract the basic features of the image. Subsequently, after the fourth CBS module, four ELAN (Efficient Layer Aggregation Network) modules are connected in series in sequence, and an MP-1 downsampling module is connected between every two ELAN modules; the ELAN module is used to efficiently aggregate and fuse features from different levels, and strengthen the feature expression through a multi-path design, enabling the network to capture both the detailed information and the overall semantic information of the image simultaneously.
[0060] Through the above design, the high-level features come from the output of the last ELAN module of the backbone network, and the low-level features come from the output of the second ELAN module in the backbone network; the low-level and high-level features extracted in this way respectively contain the detailed information and the overall semantic information of the target, and have different spatial and semantic levels.
[0061] In the traditional YOLOv7 model, the backbone network can already extract relatively rich features, but its ability to extract small targets and detailed information is weak. Therefore, the present invention further enhances it by introducing a super-resolution branch structure.
[0062] Super-Resolution Branch: The super-resolution branch enhances the detection accuracy of small targets by restoring the detailed information in low-resolution images and improving the image resolution. In the super-resolution branch structure, low-level features and high-level features are first fused through an encoder. The encoder (Encoder) uses a convolutional-ReLU (Conv-ReLU, abbreviated as CR) module to process the low-level features and upsample the high-level features simultaneously. After these processes, the low-level features and high-level features are combined through a concatenation operation and then further processed by convolution and the ReLU activation function, enhancing the feature representation ability. In the decoder part, an EDSR (Enhanced Deep Super-Resolution Network) structure is used to further process the fused features, thereby restoring a higher-resolution feature map. The decoder receives the fused features output by the encoder and restores the feature map to a higher resolution through a series of convolutional layers and upsampling operations. First, a convolutional operation is performed on the input features to extract more high-level features. Then, the spatial resolution of the feature map is increased to the resolution of the low-level features through an upsampling operation. Finally, the upsampled feature map is processed by multiple convolutional layers to generate a high-resolution output map. The purpose of the decoder is to restore more image details through super-resolution technology, thereby improving the detection accuracy.
[0063] The detection head receives the feature maps output from the second to the fourth ELAN modules of the backbone network. After multi-module processing, it finally outputs three key results: the classification prediction (Classification) of the target, the bounding box regression (Regression) of the target, and the confidence prediction (Confidence). The classification prediction is used to determine which type of foreign object the target belongs to. The regression result determines the specific position and size of the target in the image. The confidence measures the reliability of the detection result, facilitating the subsequent screening or further processing of low-confidence targets. In Figure 1 , the ELAN-W in the detection head part is a widened version of the ELAN module. ELAN-W further strengthens the feature aggregation effect and improves the fusion ability of semantic information and detailed information by widening the channels or expanding the network structure. MP-2 is a downsampling module with different parameters from the MP-1 downsampling module. REP (Re-parameterization) represents the re-parameterization operation. CBM (Convolution-BatchNorm-Module, convolution, batch normalization, activation) transforms, normalizes, and refines the features, playing a role in stabilizing and enhancing the feature representation in the network structure.
[0064] During the training phase of the YOLOv7 object detection model, the features from the training set are enhanced through the super-resolution branch structure. The model learns more detailed image details, especially in the representation of texture and semantic information, improving the detection accuracy of the model. During the inference phase, the super-resolution branch structure is removed to accelerate the inference speed. At this time, the trained network can quickly generate high-precision detection results.
[0065] To effectively combine object detection and super-resolution tasks, the present invention designs an optimized loss function. The optimized loss function consists of two parts: object detection loss (Lo) and super-resolution reconstruction loss (Ls). The object detection loss is used to evaluate the position, category, and existence of the object, while the super-resolution reconstruction loss measures the difference between the input image and the super-resolution output. By optimizing these two loss functions, the model can improve the ability to restore image details while performing object detection.
[0066] The improved loss function consists of two parts: object detection loss Lo and super-resolution reconstruction loss Ls. The expression of its total loss Ltotal is:
[0067] L total = c1Lo + c2Ls;
[0068] Among them, c1 and c2 are coefficients for balancing the object detection loss and the SR reconstruction loss.
[0069] The super-resolution reconstruction loss Ls is used to measure the difference between the input image X and the super-resolution result S. Specifically, it is calculated through the L1 loss (i.e., the absolute value difference). The expression of Ls is:
[0070] Ls = ||S - X||1;
[0071] Among them, S is the output after super-resolution reconstruction, and X is the input image. The L1 loss can effectively measure the pixel-level difference, ensuring that the details of the reconstructed image are retained.
[0072] The object detection loss Lo includes object existence loss L obj , position loss L loc and classification loss L cls , and its overall expression is:
[0073]
[0074] Among them, l represents the number of layers of the output layer; al, bl, and cl are the weight coefficients of the object position, object existence, and object classification losses in different layers; λ loc , λ obj and λ cls are hyperparameters used to adjust the relative importance of each part of the loss.
[0075] S3. As Figure 2 shown, a target tracking algorithm is introduced to avoid duplicate reporting of the same foreign object. In this embodiment, a method combining a target detection model and the Bytetrack target tracking algorithm is adopted to perform subsequent processing on the detection results, so as to achieve that a single foreign object is reported only once during the entire runway scanning process.
[0076] First, the real-time video stream is passed as input to the target detector, which analyzes each frame of the image and outputs the detection results including detection boxes (Bounding Boxes) and corresponding confidence scores; according to two preset confidence thresholds τ high and τ low , the detection boxes are divided into a set D high of high-confidence detection boxes (detection boxes with confidence scores greater than τ high ) and a set D low of low-confidence detection boxes (detection boxes with confidence scores between τ low and τ high ).
[0077] Subsequently, for the trajectory set T in the set of target trajectories existing in the current frame, the Kalman filter is used to predict each trajectory to generate the predicted position of each trajectory in the next frame.
[0078] Next, a first association is performed between the set D high of high-confidence detection boxes and the trajectory set T: calculate the cost matrix between each detection box and the trajectory, and use the Hungarian algorithm for optimal matching.
[0079] For the successfully matched trajectories, update the state parameters of their Kalman filters and retain the trajectories in the trajectory set of the current frame; for the unmatched trajectories, store them in the unmatched trajectory set T remain ; at the same time, store the unmatched detection boxes in the unmatched detection box set D remain .
[0080] After the first association, a second association is performed for the set of low-confidence detection boxes and the set of unmatched trajectories, repeating the above optimal matching operation.
[0081] For the trajectories successfully matched through the second association, similarly update the Kalman filter and add them to the trajectory set of the current frame; for the unmatched trajectories, classify them into the lost trajectory set T lost ; and for the unmatched low-confidence detection boxes, directly delete them to avoid adverse effects on subsequent processing caused by low-confidence detection results.
[0082] For the lost trajectory set Tlost If a certain trajectory has not been updated for more than 30 frames, it will be deleted from the trajectory set to avoid excessive consumption of computing resources by the system.
[0083] Subsequently, for the set of detection boxes D that were not matched during the first association process remain , the system compares these detection boxes with a preset tracking threshold ∈; if the confidence of a detection box is greater than this threshold and the condition is met in two consecutive frames, then this detection box is initialized as a new target trajectory and added to the trajectory set, and at the same time, the Kalman filter is used to predict the position of the new trajectory in the next frame.
[0084] Finally, the system outputs the current frame trajectory set after the above-mentioned multiple rounds of matching and processing. This set will be used for object detection and tracking in the next frame image. The Kalman filter continuously predicts each target trajectory to ensure the tracking stability of the target in consecutive frames, thereby achieving the purpose of reporting a single foreign object only once.
[0085] S4. Connect the object detector provided in step S3 after the output of the constructed YOLOv7 network model, input the training set into the YOLOv7 network for training, and after training is completed, output the new image to be detected to the YOLOv7 network model. The object detector outputs the detection result to confirm whether there is a foreign object on the runway.
[0086] In summary, the present invention proposes a detection method for an airport runway foreign object detection system (FOD), including:
[0087] S1. Construct a dedicated data set suitable for the airport runway foreign object detection task. First, use a high-resolution device to capture image data of a simulated runway scene and various foreign objects, annotate the foreign objects in the images, and then perform preprocessing to divide the training set and the test set.
[0088] S2. Build a YOLOv7 object detection model, introduce a super-resolution branch structure to enhance the detection accuracy of the model for small objects, and then optimize the loss function of the model. The optimized loss function consists of two parts: the object detection loss and the super-resolution reconstruction loss.
[0089] S3. Introduce an object tracking algorithm to avoid repeated reporting of the same foreign object. Connect the output of the YOLOv7 object detection model to the object detector, and perform subsequent processing on the detection results through the object tracking algorithm.
[0090] S4. Input the training set into the YOLOv7 network for training, and after training is completed, output the new image to be detected to the YOLOv7 network model. The object detector outputs the detection result to confirm whether there is a foreign object on the runway.
[0091] The steps in the present invention can be adjusted in sequence, combined, and deleted according to actual needs.
[0092] The units in the device of the present invention can be combined, divided, and deleted according to actual needs.
[0093] Although the present invention has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present invention. The protection scope of the present invention is defined by the appended claims and may include various variations, modifications, and equivalent solutions made to the invention without departing from the protection scope and spirit of the present invention.
Claims
1. A detection method for a foreign object detection system (FOD) of an airport runway, characterized in that, The method includes: S1. Construct a dedicated dataset suitable for the task of detecting foreign objects on airport runways. First, use a high-resolution device to capture image data of a simulated runway scene and various foreign objects. After annotating the foreign objects in the images, perform preprocessing and segment the training set and the test set. S2. Build a YOLOv7 object detection model, and introduce a super-resolution branch structure to enhance the detection accuracy of the model for small objects. Then optimize the loss function of the model. The optimized loss function consists of two parts: object detection loss and super-resolution reconstruction loss. S3. Introduce an object tracking algorithm to avoid duplicate reporting of the same foreign object. Connect the output of the YOLOv7 object detection model to an object detector, and perform subsequent processing on the detection results through the object tracking algorithm. S4. Input the training set into the YOLOv7 network for training. After training is completed, output the new images to be detected into the YOLOv7 network model, and the object detector outputs the detection results to confirm whether there are foreign objects on the runway.
2. The detection method of the foreign object debris (FOD) detection system for airport runway according to claim 1, characterized in that, The structure of the YOLOv7 object detection model built in step S2 includes: an input layer, a backbone network, a super-resolution branch, and a detection head. The input layer receives the input image and passes it to the backbone network. The backbone network uses a deep convolutional neural network CNN to extract the features of the input image. The purpose of the backbone network is to extract low-level features and high-level features from the image, and gradually abstract the content of the image through multiple convolutional operations. The low-level features come from the output of the second ELAN module in the backbone network, and the high-level features come from the output of the last ELAN module in the backbone network. The extracted low-level and high-level features respectively contain the detailed information and overall semantic information of the target, and have different spatial and semantic levels. The super-resolution branch processes the low-resolution feature map extracted from the backbone network, restores the image details, and improves the detection accuracy of small objects. The detection head receives the feature maps output from the second to fourth ELAN modules of the backbone network, and outputs three key results after processing: class prediction of the target, bounding box regression of the target, and confidence prediction.
3. The detection method of the foreign object debris (FOD) detection system for airport runway according to claim 2, characterized in that, In the super-resolution branch, the low-level features and high-level features are first fused through an encoder. The encoder uses a convolutional-ReLU module to process the low-level features and upsample the high-level features at the same time. After these processes, the low-level features and high-level features are combined through a concatenation operation, and then further processed by convolution and ReLU activation functions to enhance the expression ability of the features. In the decoder part, the decoder receives the fused features output by the encoder, performs convolutional operations on the input features, and extracts more high-level features. Then, the spatial resolution of the feature map is increased to the resolution of the low-level features through an upsampling operation. Finally, the upsampled feature map is processed through multiple convolutional layers to generate a high-resolution output map.
4. The detection method of the foreign object debris (FOD) detection system for airport runways according to claim 1, characterized in that, In the training stage of the YOLOv7 object detection model, the features from the training set are enhanced through the super-resolution branch structure, and the model learns more detailed image details, improving the detection accuracy of the model. In the inference stage, the super-resolution branch structure is removed to accelerate the inference speed. At this time, the trained network can quickly generate high-precision detection results.
5. The detection method of the foreign object debris (FOD) detection system for airport runway as claimed in claim 1, wherein, In the loss function of the YOLOv7 object detection model, the object detection loss is used to evaluate the position, category, and existence of the object, while the super-resolution reconstruction loss measures the difference between the input image and the super-resolution output. The improved loss function consists of the object detection loss L o and the super-resolution reconstruction loss L s The total loss L total is expressed as: L total = c1L o + c2L s ; Among them, c1 and c2 are coefficients for balancing the object detection loss and the SR reconstruction loss. Super-resolution reconstruction loss L s It is used to measure the difference between the input image X and the super-resolution result S, and is specifically calculated through the L1 loss, L s The expression is: L s = ||S - X||1; Among them, S is the output after super-resolution reconstruction, and X is the input image. The L1 loss can effectively measure the pixel-level difference, thus ensuring that the details of the reconstructed image are retained. Target detection loss L o including target existence loss L obj , position loss L loc and classification loss L cls , and its overall expression is: Among them, l represents the number of layers of the output layer; a l , b l and c l are the weight coefficients of the target position, target existence, and target classification losses at different layers respectively; λ loc , λ obj and λ cls are hyperparameters used to adjust the relative importance of each part of the loss.
6. The detection method of the foreign object debris (FOD) detection system for airport runway according to claim 1, characterized in that, The object tracking algorithm in step S3 specifically includes: First, the real-time video stream is passed as input to the object detector, which analyzes each frame of the image and outputs the detection results containing the detection boxes and the corresponding confidence scores; according to two preset confidence thresholds τ high and τ low , the detection boxes are divided into a set D high of high-confidence detection boxes and a set D low of low-confidence detection boxes; the set D high of high-confidence detection boxes includes the detection boxes with confidence scores greater than τ higt , and the set D low of low-confidence detection boxes includes the detection boxes with confidence scores between τ low and τ high ; Subsequently, for the trajectory set T in the existing object trajectory set in the current frame, the Kalman filter is used to predict each trajectory to generate the predicted position of each trajectory in the next frame. Next, perform the first association on the set D of high-confidence detection boxes high and the set T of trajectories: calculate the cost matrix between each detection box and the trajectories, and use the Hungarian algorithm for optimal matching; For the successfully matched trajectories, update the state parameters of their Kalman filters and retain these trajectories in the trajectory set of the current frame; for the trajectories that fail to match, store them in the unmatched trajectory set T remain ; Meanwhile, store the unmatched detection bounding boxes in the unmatched detection bounding box set D remain ; After the first association, a second association is performed on the set of low-confidence detection boxes and the set of unmatched trajectories, and the above optimal matching operation is repeated. For the trajectories successfully matched through the second association, update the Kalman filter and add them to the current frame trajectory set as well; for the unmatched trajectories, classify them into the lost trajectory set T lost ; for the unmatched detection boxes with low confidence, directly delete them to avoid adverse effects on subsequent processing caused by low-confidence detection results; For the lost trajectory set T lost If a certain trajectory in it has not been updated for more than 30 frames, it will be deleted from the trajectory set to avoid excessive consumption of computing resources by the system; Subsequently, for the set D of detection boxes that were not matched during the first association process remain , the system compares these detection boxes with a preset tracking threshold ∈; if the confidence of a detection box is greater than this threshold and the condition is satisfied in two consecutive frames, then this detection box is initialized as a new target trajectory and added to the trajectory set, and at the same time, the Kalman filter is used to predict the position of the new trajectory in the next frame; Finally, the system outputs the trajectory set of the current frame after the above multi-round matching and processing. This set will be used for object detection and tracking in the next frame image. The Kalman filter continuously predicts each object trajectory to ensure the tracking stability of the object in consecutive frames, thus achieving the purpose of reporting a single foreign object only once.
7. The detection method of the foreign object detection system (FOD) for airport runway according to claim 1, characterized in that, The content of labeling foreign objects in the image in step S1 includes: the bounding box of each foreign object, the category label of the foreign object, and its position information. After labeling, the dataset is saved in the PASCAL VOC format, and the labeling information is stored as an XML format file. The image file is saved in the JPEG format. Then, the data labeling file in the PASCAL VOC format is converted into the TXT file format required by YOLO.
Citation Information
Cited By
Foreign matter identification and detection method for unmanned motor sweeper in airport
CN122073045A