FOD recognition system and method based on machine vision

Through the FOD recognition system based on machine vision, multimodal data fusion and multi-level verification decision-making, the problems of inefficiency of traditional FOD detection methods and poor detection results in harsh environments are solved, and the FOD recognition effect with high accuracy and rapid processing is achieved.

CN120088583AActive Publication Date: 2025-06-03SHANDONG EAGLE INFORMATION ENG CO LTD

Patent Information

Application Number
CN202510562713.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional FOD detection methods are inefficient, making it difficult to comprehensively and promptly discover tiny or hidden foreign objects in complex airport environments, and the detection effect is not good under severe weather conditions.

Method used

The FOD recognition system based on machine vision is adopted to synchronize multi-modal image data through visible light cameras, infrared cameras and millimeter wave radars, and dynamic environment adaptive preprocessing, multi-attention feature extraction, multi-modal data fusion and multi-level verification decisions, ultimately real-time alarm and closed-loop control.

Benefits of technology

It improves the accuracy of identification of foreign objects, reduces the rate of missed detection and false detection, enhances the system's adaptability in harsh environments, realizes rapid processing from detection to alarm to clearing, and reduces security risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088583A_ABST
    Figure CN120088583A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of airport safety guarantee, and discloses an FOD recognition system and method based on machine vision. The system comprises a multispectral image acquisition module, a dynamic environment adaptive preprocessing module, a multi-attention feature extraction module, a multi-modal data fusion module, a multi-stage verification decision module and a real-time alarm and closed-loop control module. Pavement data are acquired through multi-modal image acquisition, after preprocessing, feature extraction and fusion, multi-stage verification decision is carried out to detect foreign objects, and finally real-time alarm and elimination are realized. The system can effectively overcome complex environmental interference, accurately identify foreign objects, quickly process and improve the safety and operation efficiency of the airfield pavement, and has the advantages of high identification precision, strong adaptability, good reliability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airport security, and specifically to a FOD recognition system and method based on machine vision. Background Art

[0002] In the aviation field, the presence of foreign objects on the airport pavement poses a serious threat to the safe takeoff and landing of aircraft. FOD (Foreign Object Debris) can be various objects such as metal fragments, stones, luggage items, etc. Once inhaled by an aircraft engine or colliding with components such as aircraft tires and landing gears, it is extremely likely to cause serious accidents such as engine failures and tire bursts, resulting in huge economic losses and casualties. Traditional FOD detection methods have many limitations. The manual inspection method relies on human labor, which is not only inefficient, but also difficult for humans to comprehensively and timely detect small or hidden FOD in a complex airport environment, and the detection accuracy cannot be guaranteed. At the same time, manual inspection is greatly affected by environmental factors such as weather and light. Under bad weather conditions, the detection effect will be greatly reduced.

[0003] Some sensor-based detection technologies have improved the detection efficiency to a certain extent, but there are also deficiencies. For example, a single sensor (such as only using a visible light camera) can only obtain limited information. In low visibility environments such as at night, in rain or fog, the detection ability is greatly limited and it is impossible to accurately identify FOD. And the simple combination of multiple sensors, due to the lack of an effective data fusion and processing mechanism, the data between different sensors cannot work together, and it is difficult to achieve accurate positioning and recognition of FOD.

[0004] In terms of image data processing, airport pavement images are easily affected by various noises, such as uneven illumination, rain and fog occlusion, motion blur, etc. These problems will reduce the image quality and affect subsequent feature extraction and target recognition. Existing image preprocessing technologies often cannot effectively solve multiple interference problems at the same time when dealing with complex environments, resulting in details being lost and insufficient clarity in the processed images, thereby affecting the accuracy of FOD recognition.

[0005] In the FOD recognition algorithm, existing object detection algorithms are difficult to adapt to the complex and changeable scenarios of airport pavements. Some algorithms have low detection accuracy for small targets and are prone to missing detections; some algorithms have low classification accuracy when dealing with multi-category FODs and cannot accurately distinguish different types of foreign objects. Moreover, most algorithms do not fully utilize the advantages of multi-modal data and fail to achieve deep fusion of multi-modal data, which limits the performance improvement of the FOD recognition system. After detecting foreign objects, existing FOD recognition systems lack an efficient alarm and processing mechanism. They cannot timely and accurately transmit the position information of foreign objects to relevant staff, nor can they quickly drive the cleaning device for processing, resulting in foreign objects staying on the pavement for a long time and increasing safety risks. Therefore, it is urgent to develop a machine vision-based FOD recognition system and method that can overcome the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a machine vision-based FOD recognition system and method to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solutions: A machine vision-based FOD recognition system, the system includes the following modules: A multi-spectral image acquisition module, used to synchronously acquire multi-modal image data of the airport pavement through a visible light camera, an infrared camera, and a millimeter wave radar, including visible light images, infrared thermal images, and radar point cloud data; A dynamic environment adaptive preprocessing module, used to perform super-resolution reconstruction and denoising processing on the multi-modal image data, including a rain and fog removal sub-module based on a generative adversarial network, a light intensity equalization sub-module, and a blur correction sub-module, and output a high-definition pavement image; A multi-attention feature extraction module, used to perform multi-scale feature fusion on the preprocessed image, including a spatial attention mechanism, a channel attention mechanism, and a time series attention mechanism, and construct a multi-dimensional feature map of the pavement scene; A multi-modal data fusion module, used to perform cross-modal alignment and fusion of visible light image features, infrared thermal image features, and radar point cloud features, and generate a joint feature vector through an adaptive weight allocation strategy; A multi-level verification and decision-making module, used to detect foreign objects based on the joint feature vector, including a preliminary positioning sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module, and generate foreign object coordinates and confidence information; A real-time alarm and closed-loop control module, used to trigger an alarm signal according to the foreign object coordinates, and map the foreign object position to the airport pavement geographic information system through a coordinate conversion algorithm, and drive the mechanical cleaning device to perform a cleaning task.

[0008] Preferably, in the dynamic environment adaptive preprocessing module, the generator of the generative adversarial network adopts a U-Net structure, including an encoder and a decoder. The encoder extracts multi-scale features through multiple layers of convolution, and the decoder restores high-resolution images through transposed convolution and skip connections; the discriminator adopts a PatchGAN structure to distinguish real images from reconstructed images; the loss function includes adversarial loss, perceptual loss, and L1 reconstruction loss.

[0009] Preferably, in the multi-attention feature extraction module, the spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through convolutional kernels, the channel attention mechanism learns the interdependence between channels through global average pooling and fully connected layers, and the time series attention mechanism uses gated recurrent units to model the temporal correlation between consecutive frames.

[0010] Preferably, in the multi-modal data fusion module, the cross-modal alignment realizes the unification of the visible light image coordinate system, the infrared image coordinate system, and the radar point cloud coordinate system through an affine transformation matrix, and the fusion strategy adopts a multi-modal confidence weighting method based on cross-entropy to dynamically adjust the contribution ratio of each modal feature in the joint feature vector.

[0011] Preferably, in the real-time alarm and closed-loop control module, the coordinate transformation algorithm adopts a robust registration method based on RANSAC to map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system. The error function is defined as:

[0012] where is the total registration error of the coordinate transformation, is the affine transformation matrix, is the image coordinate, is the geographic coordinate, is the Huber loss function, is the noise threshold.

[0013] Preferably, the preliminary positioning sub-module adopts an improved YOLOv8 network architecture, including: The backbone network adopts a CSPDarknet53 structure to reduce the computational complexity through cross-stage partial connections; The neck network introduces a cascaded structure of a spatial attention module and a channel attention module to dynamically enhance the feature response of the target area; The output layer of the prediction head adopts a composite loss function, including CIoU localization loss, Focal classification loss, and target existence confidence loss, to generate initial candidate boxes and their confidence scores; its expression is:

[0014] where is the balance factor, is the improved intersection over union loss, is the focal classification loss, is the confidence loss.

[0015] Preferably, the semantic segmentation verification sub-module is constructed based on the DeepLabv3+ network and includes: The segmentation head output layer adopts a hybrid attention mechanism, combines a spatial gating unit and a channel recalibration module to optimize the boundary accuracy of the pavement area and foreign objects; The post-processing unit eliminates segmentation noise through morphological closing operation filtering and uses a connected component analysis algorithm to filter out interference regions with an area smaller than a preset threshold.

[0016] Preferably, the radar reflection intensity verification sub-module includes: The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on statistical analysis of historical radar data, and the threshold calculation is:

[0017] where, is the dynamic reflection intensity threshold, is the time window is the mean value of the reflection intensity within the time window, is the standard deviation, is the confidence coefficient; The moving target detection unit calculates the radial velocity of foreign objects through Doppler frequency shift, combines Kalman filtering to predict the trajectory, and the state equation is defined as:

[0018] where, is the state vector, is the state vector at the next moment, is the state transition matrix, is the process noise; The verification logic unit verifies the geometric consistency of the visual positioning result by constructing an affine transformation equation between the image coordinates and the radar coordinates, and eliminates false alarm targets with spatial offset exceeding the tolerance.

[0019] Preferably, in the multi-spectral image acquisition module, the millimeter-wave radar adopts a frequency-modulated continuous wave mode, detects moving foreign objects through Doppler frequency shift, and separates them from static targets.

[0020] Preferably, the present invention also includes a FOD recognition method based on machine vision, and the method includes the following steps: S1: Synchronously collect multi-modal image data of the airport pavement using a visible light camera, an infrared camera, and a millimeter-wave radar, including visible light images, infrared thermal images, and radar point cloud data; S2: Perform super-resolution reconstruction and denoising processing on the collected multi-modal image data. Specifically, operate through a rain and fog removal sub-module, a lighting equalization sub-module, and a blur correction sub-module based on a generative adversarial network to output a high-definition pavement image; S3: Apply spatial attention mechanism, channel attention mechanism, and time series attention mechanism to the preprocessed images for multi-scale feature fusion to construct a multi-dimensional feature map of the pavement scene; S4: Align and fuse visible light image features, infrared thermal image features, and radar point cloud features across modalities, and adopt an adaptive weight allocation strategy to generate a joint feature vector; S5: Perform foreign object detection based on the joint feature vector, and sequentially pass through a preliminary localization sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module to generate foreign object coordinates and confidence information; S6: Trigger an alarm signal according to the foreign object coordinates, and map the foreign object position to the airport pavement geographic information system through a coordinate conversion algorithm to drive the mechanical cleaning device to perform the cleaning task.

[0021] Compared with the prior art, the beneficial effects of the present invention are: The FOD recognition system of the present invention synchronously collects multi-modal image data, including visible light images, infrared thermal images, and radar point cloud data, using a visible light camera, an infrared camera, and a millimeter-wave radar. The multi-modal data reflects the airport pavement situation from different angles. For example, visible light images provide intuitive visual information, infrared thermal images can detect targets in low light or special object situations, and radar point cloud data is used to detect moving targets. The multi-modal data fusion module aligns and fuses these different modal features across modalities, and adopts an adaptive weight allocation strategy to generate a joint feature vector, giving full play to the advantages of each modal data. Compared with single-modal data recognition, it greatly improves the accuracy of foreign object recognition and effectively reduces the missed detection and false detection rates.

[0022] The dynamic environment adaptive preprocessing module addresses the image interference problem in complex environments and adopts a rain and fog removal sub-module, a lighting equalization sub-module, and a blur correction sub-module based on a generative adversarial network. The generator of the generative adversarial network adopts a U-Net structure, which can effectively remove rain and fog, equalize lighting, and correct blur, and output a high-definition pavement image. This enables the system to still work stably under harsh conditions such as rain and fog, uneven lighting, and blur, ensuring the accuracy of subsequent feature extraction and recognition, greatly enhancing the adaptability of the system to complex environments, and broadening the application scenarios of the system.

[0023] The multi-attention feature extraction module applies spatial attention mechanism, channel attention mechanism and time series attention mechanism. The spatial attention mechanism adjusts local weights according to the importance of image regions to highlight potential FOD regions; the channel attention mechanism learns the dependence between channels to enhance the features of key channels; the time series attention mechanism uses gated recurrent units to model the temporal correlation of consecutive frames to capture the dynamic changes of foreign objects. The three mechanisms work together to construct a multi-dimensional feature map, comprehensively and accurately reflecting the pavement features, improving the feature extraction ability for small, hidden and dynamically changing foreign objects, and further enhancing the recognition accuracy.

[0024] The multi-level verification and decision-making module performs foreign object detection based on the joint feature vector, including a preliminary localization sub-module, a semantic segmentation verification sub-module and a radar reflection intensity verification sub-module. The preliminary localization sub-module uses an improved YOLOv8 network architecture to quickly locate potential foreign objects; the semantic segmentation verification sub-module optimizes the boundary accuracy between the pavement and foreign objects based on the DeepLabv3+ network; the radar reflection intensity verification sub-module eliminates false alarm targets through dynamic threshold setting, detection of moving targets and geometric consistency verification. The multi-level verification cooperates with each other to verify the detection results from different angles, effectively improving the reliability of foreign object detection and ensuring that the system outputs accurate foreign object coordinates and confidence information.

[0025] The real-time alarm and closed-loop control module triggers an alarm signal according to the foreign object coordinates, and maps the position to the airport pavement geographic information system through a coordinate conversion algorithm to drive the mechanical cleaning device to perform the cleaning task. The coordinate conversion algorithm uses a robust registration method based on RANSAC to ensure the accuracy of coordinate conversion. This closed-loop control process realizes fast processing from detection to alarm and then to cleaning, greatly shortening the residence time of foreign objects on the pavement, effectively reducing safety risks, and improving the safety and operation efficiency of the airport pavement. Brief Description of the Drawings

[0026] Figure 1 It is the working principle diagram of the FOD recognition system based on machine vision described in the present invention; Figure 2 It is the working flow chart of the multi-modal data fusion module; Figure 3 It is the working flow chart of the coordinate conversion of the real-time alarm and closed-loop control module; Figure 4 It is the step diagram of the radar reflection intensity verification sub-module. Detailed Embodiments

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] Please refer to Figures 1-4 , the present invention relates to a FOD recognition system based on machine vision, which aims to improve the accuracy and efficiency of foreign object recognition on airport pavements and ensure the safety of airport operations. The specific implementation manners of the present invention will be elaborated in detail below. The system includes: Multi-spectral image acquisition module: This module synchronously acquires multi-modal image data of the airport pavement using a visible light camera, an infrared camera, and a millimeter-wave radar. Among them, the visible light camera is responsible for acquiring the visible light image of the airport pavement, which can clearly present details such as the texture and color of the pavement; the infrared camera obtains an infrared thermal image, and through the thermal radiation differences of different objects, it can more clearly identify some objects that are not easily detected under visible light, such as foreign objects with a temperature difference from the pavement; the millimeter-wave radar collects radar point cloud data, which uses the principle of electromagnetic wave reflection to accurately measure information such as the position and speed of the target object. Through the synchronous acquisition of these three devices, comprehensive and rich pavement information is obtained, providing an original data basis for subsequent analysis and processing.

[0029] Dynamic environment adaptive preprocessing module: This module performs super-resolution reconstruction and denoising processing on the acquired multi-modal image data. In a complex airport environment, the images may be affected by problems such as rain, fog, uneven illumination, and blurring. Therefore, multiple sub-modules are required to optimize the image quality. The rain and fog removal sub-module based on the generative adversarial network uses the characteristics of the generative adversarial network, where the generator and the discriminator learn against each other to effectively remove rain and fog interference in the image; the illumination equalization sub-module adjusts the illumination of the image through a specific algorithm to make the brightness of each part of the image more uniform and enhance the readability of the image; the blur correction sub-module corrects the blurred image caused by device jitter or other reasons, and finally outputs a high-definition pavement image, providing high-quality data for subsequent feature extraction and analysis.

[0030] Multi - attention Feature Extraction Module: It performs multi - scale feature fusion on the pre - processed image and constructs a multi - dimensional feature map of the pavement scene through spatial attention mechanism, channel attention mechanism and time - series attention mechanism. The spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through convolutional kernels, enabling the model to pay more attention to the spatial location where the target object is located; the channel attention mechanism learns the channel - to - channel dependence relationship through global average pooling and fully - connected layers, automatically adjusting the importance of different channels; the time - series attention mechanism uses gated recurrent units to model the temporal correlation between consecutive frames, capturing the dynamic change information of objects, so as to more accurately identify foreign objects in motion.

[0031] Multi - modal Data Fusion Module: It performs cross - modal alignment and fusion of visible - light image features, infrared thermal - imaging features and radar point - cloud features. First, it realizes the unification of the visible - light image coordinate system, infrared image coordinate system and radar point - cloud coordinate system through an affine transformation matrix, enabling data of different modalities to be fused in the same coordinate system. The fusion strategy adopts a multi - modal confidence - weighted method based on cross - entropy. According to the reliability of different - modality data in identifying foreign objects, it dynamically adjusts the contribution ratio of each - modality feature in the joint feature vector, generating a more representative joint feature vector and improving the accuracy of recognition.

[0032] Multi - level Verification and Decision - making Module: It performs foreign - object detection based on the joint feature vector, and this module contains multiple sub - modules. The preliminary localization sub - module uses a specific network architecture to generate initial candidate boxes and their confidence scores; the semantic segmentation verification sub - module performs precise segmentation and verification of the pavement area and foreign objects based on a specific network; the radar reflection intensity verification sub - module further improves the detection accuracy through the analysis and verification of radar reflection intensity, and finally generates foreign - object coordinates and confidence information.

[0033] Real - time Alarm and Closed - loop Control Module: It triggers an alarm signal according to the foreign - object coordinates, and at the same time maps the foreign - object position to the airport pavement geographic information system through a coordinate conversion algorithm. This coordinate conversion algorithm uses a specific robust registration method to ensure the accuracy of coordinate conversion. Then it drives the mechanical cleaning device to perform the cleaning task, realizing the timely handling of foreign objects and ensuring the safety of the airport pavement.

[0034] The technical solutions of the present invention will be further described in detail below in conjunction with specific embodiments.

[0035] Embodiment 1

[0036] In this embodiment, the specific implementation of the generative adversarial network in the dynamic - environment adaptive pre - processing module is elaborated in detail. The generator of the generative adversarial network adopts a U - Net structure, which consists of an encoder and a decoder. The role of the encoder is to extract multi - scale features through multiple layers of convolution. Assuming the input image is , the The layer convolution operation can be expressed as . After multiple layers of convolution, the size of the image gradually shrinks, and the feature information is concentrated. The decoder restores the high-resolution image through deconvolution and skip connections. The deconvolution operation can be expressed as , where represents the number of deconvolution layers. The skip connection connects the feature maps of the corresponding layers in the encoder with the corresponding layers in the decoder, enabling the decoder to utilize the feature information of different scales in the encoder and better restore the image details.

[0037] The discriminator adopts the PatchGAN structure, whose purpose is to distinguish between real images and reconstructed images. The discriminator makes patch-level judgments on the input images, determining whether each patch comes from a real image or a reconstructed image generated by the generator. The loss function includes adversarial loss, perceptual loss, and L1 reconstruction loss. The adversarial loss is used to measure the difference between the image generated by the generator and the real image, prompting the generator to generate results closer to the real image. Its formula is: , where represents the real image, is the data distribution of the real image, is the random noise, is the noise distribution, is the discriminator, is the generator. The perceptual loss measures the difference from the semantic and feature levels of the image. It is based on a pre-trained neural network and optimizes the generator by comparing the differences between the generated image and the real image on the network feature layer. The formula is: , where represents the feature extraction function of the th layer of the pre-trained network, is the number of elements of the feature map of the th layer. The L1 reconstruction loss is used to ensure the pixel-level similarity between the generated image and the real image. The formula is: . Through the combined action of these three loss functions, the generative adversarial network can effectively perform operations such as rain and fog removal and super-resolution reconstruction on multi-modal image data.

[0038] Embodiment 2

[0039] This embodiment is used to describe the specific working methods of the spatial attention mechanism, channel attention mechanism, and time series attention mechanism in the multi-attention feature extraction module.

[0040] The spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through convolutional kernels. Let the input feature map be , the spatial attention mechanism first performs average pooling and max pooling operations on the feature map in the channel dimension to obtain the average pooling feature map and the maximum pooling feature map . Then these two feature maps are concatenated, and a convolutional layer is passed through to generate a spatial attention weight map , and its formula is: , where is the Sigmoid activation function, represents the convolutional operation, represents concatenating the two feature maps along the channel dimension. Finally, the generated spatial attention weight map is multiplied by the original feature map to obtain the enhanced feature map , enabling the model to pay more attention to the spatial location where the target object is located.

[0041] The channel attention mechanism learns the interdependencies between channels through global average pooling and fully connected layers. For the input feature map , global average pooling is performed to obtain the mean vector on the channel dimension, and its calculation formula is: , where and are the height and width of the feature map respectively. Then the mean vector passes through two fully connected layers. The first fully connected layer reduces the dimension, and the second fully connected layer increases the dimension. Finally, the channel attention weight vector is obtained through the Sigmoid activation function, and the formula is: , where and represent the two fully connected layers respectively, is the ReLU activation function. The channel attention weight vector is weighted with the original feature map along the channel dimension to obtain the enhanced feature map , realizing the adjustment of the importance of different channels.

[0042] The time series attention mechanism uses a gated recurrent unit (GRU) to model the temporal correlation between consecutive frames. Assuming the input sequence of consecutive frame feature maps is , the input of the GRU is the current frame feature map and the hidden state at the previous moment. Inside the GRU, the information flow is controlled through the update gate and the reset gate . The calculation formula of the update gate is: , and the calculation formula of the reset gate is: , where and are weight matrices, represents concatenating the current frame feature map and the hidden state at the previous moment. Then through the candidate hidden state to calculate the hidden state at the current moment , and the formula is: , , where is the weight matrix, represents element-wise multiplication. Through the processing of GRU, the temporal correlation between consecutive frames can be effectively captured, improving the recognition ability of moving objects.

[0043] Example 3

[0044] This example details the specific implementation of the cross-modal alignment and fusion strategy in the multi-modal data fusion module.

[0045] The cross-modal alignment realizes the unification of the visible light image coordinate system, the infrared image coordinate system, and the radar point cloud coordinate system through an affine transformation matrix. For a point in the visible light image, the corresponding point in the infrared image, and the corresponding point in the radar point cloud coordinate system, the affine transformation matrix can transform them to the same coordinate system. Taking the transformation from the visible light image coordinate system to the unified coordinate system as an example, the transformation formula is: , where is the coordinate after transformation to the unified coordinate system. The affine transformation matrix is obtained by calculating the coordinates of known corresponding points, and methods such as the least squares method can be used for solving.

[0046] The fusion strategy adopts a multi-modal confidence weighted method based on cross-entropy. Let the visible light image feature be , the infrared thermal imaging feature be , and the radar point cloud feature be . Their corresponding confidences are respectively , , . First, calculate the weighted coefficient of each modal feature. Taking the visible light image feature as an example, the calculation formula of its weighted coefficient is: . Then, generate the joint feature vector by weighted summation, and the formula is: . In this way, according to the reliability of different modal data in identifying foreign objects, the contribution ratio of each modal feature in the joint feature vector is dynamically adjusted to improve the recognition accuracy.

[0047] Example 4

[0048] The coordinate transformation algorithm in the real-time alarm and closed-loop control module uses a robust registration method based on RANSAC (Random Sample Consensus algorithm) to map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system. Assume there is a series of foreign object coordinate points in the image coordinate system , and there are corresponding coordinate points in the airport geographic coordinate system . The basic idea of the RANSAC algorithm is to randomly sample and select some point pairs to calculate the affine transformation matrix , and then use this matrix to transform all points and calculate the error between the transformed points and the true geographic coordinate points. The error function is defined as: , where is the total registration error of coordinate transformation, is the affine transformation matrix, is the image coordinate, is the geographic coordinate, is the Huber loss function, is the noise threshold. The expression of the Huber loss function is: , where is a threshold used to balance the calculation method of the error. When the error is small, the squared error is used, and when the error is large, the linear error is used, which can improve the robustness of the algorithm to noise and outliers.

[0049] In practical applications, the RANSAC algorithm continuously repeats the processes of random sampling, calculating the affine transformation matrix, and error, and selects the affine transformation matrix with the smallest error as the final coordinate transformation matrix. In this way, the coordinates of foreign objects in the image coordinate system can be accurately mapped to the airport geographic coordinate system, providing accurate position information for the subsequent mechanical cleaning device.

[0050] Example 5

[0051] This example focuses on describing the specific details of the improved YOLOv8 network architecture adopted by the preliminary positioning sub-module. The backbone network uses the CSPDarknet53 structure, which reduces the computational amount through cross-stage partial connection. The CSPDarknet53 structure divides the feature map of the backbone network into two parts. One part is directly passed to the next stage, and the other part is merged with the directly passed part after a series of convolutional operations. Assume the input feature map is , and after being processed by a certain stage in the CSPDarknet53 structure, the calculation process of the output feature map can be expressed as: , , , where and respectively represent different convolution operations. Through this cross-stage partial connection method, without losing too much accuracy, the computational complexity is effectively reduced, and the running efficiency of the network is improved.

[0052] The neck network introduces a cascaded structure of a spatial attention module and a channel attention module to dynamically enhance the feature response of the target region. The working mode of the spatial attention module is as described in Embodiment 2. By generating a spatial attention weight map to weight the feature map. The channel attention module also generates a channel attention weight vector in the same way as in Embodiment 2 . For the input feature map , after being processed by the neck network, the output feature map is calculated by the formula: . Through this cascaded structure, the features of the target region can be more prominent, and the detection accuracy of the target object can be improved.

[0053] The output layer of the prediction head adopts a composite loss function, including CIoU localization loss, Focal classification loss, and object existence confidence loss, to generate initial candidate boxes and their confidence scores. Its expression is: , where , , is a balance factor used to adjust the importance of different loss functions; is the improved intersection over union loss. On the basis of the traditional intersection over union loss, it takes into account factors such as the distance, overlap rate, and scale consistency between the predicted box and the ground truth box. The calculation formula is: , where is the intersection over union, is the center point distance between the predicted box and the ground truth box , is the diagonal length of the smallest closed region containing the two boxes, is the weight coefficient, is a parameter to measure the aspect ratio consistency; is the focal classification loss used to solve the problem of sample imbalance. The formula is: , where is the number of samples, is the class weight, is the predicted probability, is the focal parameter; is the confidence loss used to measure the confidence of whether there is an object in the predicted box. Through this composite loss function, the localization and classification accuracy of the network for target objects can be improved.

[0054] Embodiment 6

[0055] This embodiment elaborates in detail the specific implementation of the semantic segmentation verification sub-module based on the DeepLabv3+ network. The output layer of the segmentation head adopts a hybrid attention mechanism, combining a spatial gating unit and a channel recalibration module to optimize the boundary accuracy between the pavement area and foreign objects. The spatial gating unit analyzes the information of spatial positions to generate spatial gating weights , and weights the feature map in the spatial dimension. Let the input feature map be , and the feature map after being processed by the spatial gating unit is calculated as follows: . The channel recalibration module generates channel recalibration weights by learning the relationships between channels, and weights the feature map in the channel dimension, with the formula: . Through this hybrid attention mechanism, the boundary features between the pavement area and foreign objects can be more accurately highlighted, improving the segmentation accuracy.

[0056] The post-processing unit eliminates segmentation noise through morphological closing filtering and uses a connected component analysis algorithm to filter out interference regions with an area smaller than a preset threshold. The morphological closing operation consists of dilation and erosion operations. Assuming the structuring element is , for the segmented binary image , the dilation operation expands the object boundaries in the image outward, and its calculation formula is: , where is the image area covered by the structuring element when the center of the structuring element is translated to the pixel point in the image . The erosion operation contracts the object boundaries inward, with the formula: . The closing operation that first performs dilation and then erosion can fill small holes inside the object and smooth the object boundaries, with the formula: .

[0057] The connected component analysis algorithm marks the connected regions in the image and calculates the area of each region. Let be a certain connected region, and its area is calculated by counting the number of pixel points in the region . The area is compared with the preset threshold . If , then the region is determined as an interference region and filtered, thereby further improving the accuracy of the segmentation result.

[0058] Embodiment 7

[0059] The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on the statistical analysis of historical radar data. The threshold is calculated as follows: . Among them, is the dynamic reflection intensity threshold, which will be dynamically adjusted as time changes; is the time window The average value of the reflection intensity within reflects the average level of the radar reflection intensity during this time period. The calculation method is to divide the sum of all reflection intensity values within the time window by the number of reflection intensity values , that is ; is the standard deviation, which is used to measure the degree of dispersion of the reflection intensity values around the mean value. The calculation formula is ; is the confidence coefficient, which is set according to the requirements and experience of the actual scene and is used to adjust the strictness of the threshold. By setting the threshold dynamically in this way, different airport environments and weather conditions can be adapted.

[0060] The moving target detection unit calculates the radial velocity of foreign objects through the Doppler frequency shift and predicts the trajectory in combination with the Kalman filter. When a foreign object moves relative to the radar, a Doppler frequency shift will be generated. According to the principle of the Doppler effect, the relationship between it and the radial velocity of the foreign object is , where is the wavelength of the electromagnetic wave emitted by the radar. By measuring the Doppler frequency shift , the radial velocity of the foreign object can be calculated.

[0061] The Kalman filter is used to predict the trajectory of foreign objects. Its state equation is defined as: . Among them, is the state vector, which contains information such as the position and velocity of the foreign object at moment; is the state vector at the next moment; is the state transition matrix, which is determined according to the motion model of the system and is used to describe the transition relationship of the state vector from the current moment to the next moment; is the process noise, which is used to represent the uncertain factors existing in the system and is usually assumed to follow a Gaussian distribution.

[0062] The verification logic unit verifies the geometric consistency of the visual positioning result by constructing an affine transformation equation between the image coordinates and the radar coordinates, and eliminates the false alarm targets with excessive spatial offsets. Assume the image coordinates are , and the radar coordinates are , the affine transformation equation is: , where is the affine transformation matrix. By calculating the difference between the coordinates of the foreign object obtained through visual positioning after affine transformation and the radar measurement coordinates, if the difference exceeds a certain threshold, the target is determined as a false alarm target and eliminated, thereby improving the accuracy of FOD recognition.

[0063] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0064] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A FOD recognition system based on machine vision, characterized in that: Includes the following modules: Multispectral image acquisition module, used to synchronously collect multimodal image data of the airport pavement through visible light camera, infrared camera and millimeter wave radar, including visible light image, infrared thermal imaging and radar point cloud data; A dynamic environment adaptive preprocessing module, used for super-resolution reconstruction and denoising of the multimodal image data, including a rain and fog removal submodule, an illumination equalization submodule and a blur correction submodule based on a generative adversarial network, and outputting a high-definition road surface image; The multi-attention feature extraction module is used to perform multi-scale feature fusion on the pre-processed image, including spatial attention mechanism, channel attention mechanism and time series attention mechanism, to construct a multi-dimensional feature map of the road scene; Multimodal data fusion module, used to align and fuse visible light image features, infrared thermal imaging features and radar point cloud features across modalities, and generate joint feature vectors through adaptive weight allocation strategy; A multi-level verification decision module is used to detect foreign objects based on joint feature vectors, including a preliminary positioning submodule, a semantic segmentation verification submodule, and a radar reflection intensity verification submodule to generate foreign object coordinates and confidence information; The real-time alarm and closed-loop control module is used to trigger an alarm signal based on the coordinates of foreign objects, and map the location of foreign objects to the airport pavement geographic information system through a coordinate conversion algorithm, thereby driving the mechanical removal device to perform the removal task.

2. The system according to claim 1, characterized in that In the dynamic environment adaptive preprocessing module, the generator of the generative adversarial network adopts a U-Net structure, including an encoder and a decoder, the encoder extracts multi-scale features through multi-layer convolution, and the decoder restores high-resolution images through deconvolution and jump connection; The discriminator adopts the PatchGAN structure to distinguish real images from reconstructed images; the loss functions include adversarial loss, perceptual loss and L1 reconstruction loss.

3. The system according to claim 1, characterized in that In the multi-attention feature extraction module, the spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through the convolution kernel, the channel attention mechanism learns the dependency between channels through global average pooling and the fully connected layer, and the time series attention mechanism uses the gated recurrent unit to model the temporal association between consecutive frames.

4. The system according to claim 1, characterized in that In the multimodal data fusion module, the cross-modal alignment realizes the unification of the visible light image coordinate system, the infrared image coordinate system and the radar point cloud coordinate system through an affine transformation matrix. The fusion strategy adopts a multimodal confidence weighted method based on cross entropy to dynamically adjust the contribution ratio of each modal feature in the joint feature vector.

5. The system according to claim 1, characterized in that In the real-time alarm and closed-loop control module, the coordinate conversion algorithm uses a robust registration method based on RANSAC to map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system. The error function is defined as: ; in, is the total registration error of coordinate transformation, is the affine transformation matrix, are image coordinates, is the geographic coordinate, is the Huber loss function, is the noise threshold.

6. The system according to claim 1, characterized in that The preliminary positioning submodule adopts an improved YOLOv8 network architecture, including: The backbone network adopts the CSPDarknet53 structure, which reduces the amount of calculation by partially connecting across stages; The neck network introduces a cascade structure of spatial attention module and channel attention module to dynamically enhance the characteristic response of the target area; The prediction head output layer uses a composite loss function, including CIoU positioning loss, Focal classification loss and target existence confidence loss, to generate the initial candidate box and its confidence score; its expression is: ; in, is the balance factor, For the improved intersection-over-union loss, To focus on classification loss, is the confidence loss.

7. The system according to claim 6, characterized in that The semantic segmentation verification submodule is built based on the DeepLabv3+ network and includes: The segmentation head output layer uses a hybrid attention mechanism, combined with a spatial gating unit and a channel recalibration module to optimize the boundary accuracy of the road surface area and foreign objects; The post-processing unit eliminates segmentation noise through morphological closing operation filtering, and uses the connected domain analysis algorithm to filter out interference areas whose area is smaller than a preset threshold.

8. The system according to claim 7, characterized in that The radar reflection intensity verification submodule includes: The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on the statistical analysis of historical radar data. The threshold is calculated as: ; in, is the dynamic reflection intensity threshold, For time window The mean reflection intensity within is the standard deviation, is the confidence coefficient; The mobile target detection unit calculates the radial velocity of the foreign object through the Doppler frequency shift and predicts the trajectory by combining the Kalman filter. The state equation is defined as: ; in, is the state vector, is the state vector at the next moment, is the state transfer matrix, is the process noise; The verification logic unit verifies the geometric consistency of the visual positioning results by constructing the affine transformation equation between the image coordinates and the radar coordinates, and eliminates false alarm targets with excessive spatial offset.

9. The system according to claim 1, characterized in that In the multispectral image acquisition module, the millimeter wave radar adopts a frequency modulated continuous wave mode to detect moving foreign objects through Doppler frequency shift and separate them from static targets.

10. A FOD identification method based on machine vision, characterized in that: The following steps are involved: S1: Use visible light cameras, infrared cameras and millimeter wave radars to synchronously collect multi-modal image data of the airport pavement, including visible light images, infrared thermal imaging and radar point cloud data; S2: Perform super-resolution reconstruction and denoising on the collected multimodal image data, specifically through the rain and fog removal submodule, illumination equalization submodule and blur correction submodule based on the generative adversarial network to output high-definition road surface images; S3: Use spatial attention mechanism, channel attention mechanism and time series attention mechanism to perform multi-scale feature fusion on the pre-processed image to construct a multi-dimensional feature map of the road scene; S4: Cross-modal alignment and fusion of visible light image features, infrared thermal imaging features and radar point cloud features, and the adaptive weight allocation strategy is used to generate a joint feature vector; S5: Detect foreign objects based on the joint feature vector, and generate foreign object coordinates and confidence information through the preliminary positioning submodule, semantic segmentation verification submodule and radar reflection intensity verification submodule in sequence; S6: Trigger an alarm signal based on the coordinates of the foreign object, and map the location of the foreign object to the airport pavement geographic information system through a coordinate conversion algorithm, driving the mechanical removal device to perform the removal task.

Citation Information

Patent Citations

  • FOD detection method, device and system

    CN112285706A

  • Visible light, infrared and radar fusion target detection method based on deep learning

    CN114254696A

  • Computer visual identification system and method

    CN116958771A

  • Road safety monitoring method based on LiDAR and thermal infrared visible light information fusion

    CN118038226A

  • Rail transit safe operation and early warning system

    CN119796274A

Cited By

  • Vehicle-mounted monitoring system based on deep learning

    CN120385999A

  • A vehicle monitoring system based on deep learning

    CN120385999B

  • Image recognition-based canteen catering compliance real-time monitoring system and method

    CN120673341A

  • Universal autonomous charging method and system for mobile robot and storage medium

    CN120779941A

  • A universal autonomous charging method, system, and storage medium for mobile robots

    CN120779941B