An FOD recognition system and method based on machine vision
Through multi-spectral image acquisition and multi-modal data fusion technology, the detection efficiency and accuracy of the FOD recognition system in complex environments is solved, efficient identification and timely removal of foreign objects are achieved, and the safety and operational efficiency of the airport road surface are improved.
Patent Information
- Application Number
- CN202510562713.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing FOD identification system has low detection efficiency and insufficient accuracy in complex environments, making it difficult to achieve deep integration of multimodal data, and lacks efficient alarm and processing mechanisms, resulting in increased security risks.
The multi-spectral image acquisition module, dynamic environment adaptive preprocessing module, multi-attention feature extraction module, multi-modal data fusion module and multi-level verification decision-making module are used to collect multi-modal image data in combination with visible light cameras, infrared cameras and millimeter wave radars, super-resolution reconstruction and denoising processing, multi-scale feature fusion, and foreign object coordinates and confidence information are generated through multi-level verification, and the mechanical cleaning device is driven to perform the removal task.
It improves the accuracy and efficiency of FOD identification, enhances the system's adaptability to complex environments, shortens the residence time of foreign objects on the road surface, reduces safety risks, and improves the safety and operational efficiency of the airport road surface.
Smart Images

Figure CN120088583B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of airport security, and particularly to a FOD recognition system and method based on machine vision. Background Art
[0002] In the aviation field, the presence of foreign objects on the airport pavement poses a serious threat to the safe takeoff and landing of aircraft. FOD (Foreign Object Debris) can be various objects such as metal fragments, stones, luggage items, etc. Once inhaled by an aircraft engine or collided with components such as aircraft tires and landing gears, it is extremely likely to cause serious accidents such as engine failures and tire bursts, resulting in huge economic losses and casualties. Traditional FOD detection methods have many limitations. The manual inspection method relies on manpower, which is not only inefficient, but also in a complex airport environment, it is difficult for humans to comprehensively and timely detect small or hidden FOD, and the detection accuracy cannot be guaranteed. At the same time, manual inspection is greatly affected by environmental factors such as weather and light. Under bad weather conditions, the detection effect will be greatly reduced.
[0003] Some sensor-based detection technologies have improved the detection efficiency to a certain extent, but there are also deficiencies. For example, a single sensor (such as only using a visible light camera) can only obtain limited information. In low visibility environments such as at night, in rain or fog, the detection ability is greatly limited and it is impossible to accurately identify FOD. And the simple combination of multiple sensors, due to the lack of an effective data fusion and processing mechanism, the data between different sensors cannot work together, and it is difficult to achieve accurate positioning and recognition of FOD.
[0004] In terms of image data processing, airport pavement images are easily interfered by various noises, such as uneven illumination, rain and fog occlusion, motion blur, etc. These problems will reduce the image quality and affect subsequent feature extraction and target recognition. Existing image preprocessing technologies often cannot effectively solve multiple interference problems at the same time when dealing with complex environments, resulting in details loss and insufficient clarity in the processed images, thereby affecting the accuracy of FOD recognition.
[0005] In the FOD recognition algorithm, existing object detection algorithms are difficult to adapt to the complex and changeable scenarios of airport pavements. Some algorithms have low detection accuracy for small targets and are prone to missed detections; some algorithms have low classification accuracy when dealing with multi-category FODs and cannot accurately distinguish different types of foreign objects. Moreover, most algorithms do not fully utilize the advantages of multi-modal data and fail to achieve deep fusion of multi-modal data, restricting the performance improvement of the FOD recognition system. After detecting foreign objects, existing FOD recognition systems lack an efficient alarm and processing mechanism. They cannot timely and accurately transmit the position information of foreign objects to relevant staff, nor can they quickly drive the cleaning device for processing, resulting in foreign objects staying on the pavement for a long time and increasing safety risks. Therefore, it is urgent to develop a machine vision-based FOD recognition system and method that can overcome the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a machine vision-based FOD recognition system and method to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: A machine vision-based FOD recognition system, the system includes the following modules:
[0008] A multi-spectral image acquisition module, used to synchronously acquire multi-modal image data of the airport pavement through a visible light camera, an infrared camera, and a millimeter wave radar, including visible light images, infrared thermal images, and radar point cloud data;
[0009] A dynamic environment adaptive preprocessing module, used to perform super-resolution reconstruction and denoising processing on the multi-modal image data, including a rain and fog removal sub-module, a light intensity equalization sub-module, and a blur correction sub-module based on a generative adversarial network, and output a high-definition pavement image;
[0010] A multi-attention feature extraction module, used to perform multi-scale feature fusion on the preprocessed image, including a spatial attention mechanism, a channel attention mechanism, and a time series attention mechanism, and construct a multi-dimensional feature map of the pavement scene;
[0011] A multi-modal data fusion module, used to perform cross-modal alignment and fusion of visible light image features, infrared thermal image features, and radar point cloud features, and generate a joint feature vector through an adaptive weight allocation strategy;
[0012] A multi-level verification decision module, used to detect foreign objects based on the joint feature vector, including a preliminary positioning sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module, and generate foreign object coordinates and confidence information;
[0013] The real-time alarm and closed-loop control module is used to trigger an alarm signal according to the coordinates of foreign objects, and map the position of foreign objects to the airport pavement geographic information system through a coordinate conversion algorithm, and drive the mechanical cleaning device to perform the cleaning task.
[0014] Preferably, in the dynamic environment adaptive preprocessing module, the generator of the generative adversarial network adopts a U-Net structure, including an encoder and a decoder. The encoder extracts multi-scale features through multi-layer convolution, and the decoder restores high-resolution images through deconvolution and skip connections; the discriminator adopts a PatchGAN structure to distinguish real images from reconstructed images; the loss function includes adversarial loss, perceptual loss, and L1 reconstruction loss.
[0015] Preferably, in the multi-attention feature extraction module, the spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through a convolution kernel, the channel attention mechanism learns the dependence relationship between channels through global average pooling and a fully connected layer, and the time series attention mechanism uses a gated recurrent unit to model the temporal correlation between consecutive frames.
[0016] Preferably, in the multi-modal data fusion module, the cross-modal alignment realizes the unification of the visible light image coordinate system, the infrared image coordinate system, and the radar point cloud coordinate system through an affine transformation matrix, and the fusion strategy adopts a multi-modal confidence weighting method based on cross-entropy to dynamically adjust the contribution ratio of each modal feature in the joint feature vector.
[0017] Preferably, in the real-time alarm and closed-loop control module, the coordinate conversion algorithm adopts a robust registration method based on RANSAC to map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system, and the error function is defined as: where, is the total registration error of the coordinate conversion, is the affine transformation matrix, is the image coordinate, is the geographic coordinate, is the Huber loss function, is the noise threshold.
[0018] Preferably, the preliminary positioning sub-module adopts an improved YOLOv8 network architecture, including:
[0019] The backbone network adopts a CSPDarknet53 structure to reduce the computational amount through cross-stage partial connection;
[0020] The neck network introduces a cascaded structure of a spatial attention module and a channel attention module to dynamically enhance the feature response of the target area;
[0021] The prediction head output layer uses a composite loss function, including CIoU positioning loss, Focal classification loss and target existence confidence loss, to generate the initial candidate box and its confidence score; its expression is: in, is the balance factor, For the improved intersection-over-union loss, Focusing on classification loss, is the confidence loss.
[0022] Preferably, the semantic segmentation verification submodule is constructed based on the DeepLabv3+ network, including:
[0023] The segmentation head output layer uses a hybrid attention mechanism, combined with a spatial gating unit and a channel recalibration module to optimize the boundary accuracy of the road surface area and foreign objects;
[0024] The post-processing unit eliminates segmentation noise through morphological closing operation filtering, and uses the connected domain analysis algorithm to filter out interference areas whose area is smaller than a preset threshold.
[0025] Preferably, the radar reflection intensity verification submodule includes:
[0026] The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on the statistical analysis of historical radar data. The threshold is calculated as: in, is the dynamic reflection intensity threshold, For time window The mean reflection intensity within is the standard deviation, is the confidence coefficient; the mobile target detection unit calculates the radial velocity of the foreign object through the Doppler frequency shift and combines the Kalman filter to predict the trajectory. The state equation is defined as: in, is the state vector, is the state vector at the next moment, is the state transfer matrix, is the process noise;
[0027] The verification logic unit verifies the geometric consistency of the visual positioning results by constructing the affine transformation equation between the image coordinates and the radar coordinates, and eliminates false alarm targets with excessive spatial offset.
[0028] Preferably, in the multispectral image acquisition module, the millimeter wave radar adopts a frequency modulated continuous wave mode to detect moving foreign objects through Doppler frequency shift and separate them from static targets.
[0029] Preferably, the present invention also includes a FOD identification method based on machine vision, the method comprising the following steps:
[0030] S1: Synchronously collect multi-modal image data of the airport pavement using visible light cameras, infrared cameras, and millimeter-wave radars, including visible light images, infrared thermal images, and radar point cloud data;
[0031] S2: Perform super-resolution reconstruction and denoising processing on the collected multi-modal image data, specifically through the rain and fog removal sub-module, illumination equalization sub-module, and blur correction sub-module based on the generative adversarial network, and output high-definition pavement images;
[0032] S3: Apply spatial attention mechanism, channel attention mechanism, and time series attention mechanism to the preprocessed images for multi-scale feature fusion, and construct a multi-dimensional feature map of the pavement scene;
[0033] S4: Align and fuse visible light image features, infrared thermal image features, and radar point cloud features across modalities, and adopt an adaptive weight allocation strategy to generate a joint feature vector;
[0034] S5: Perform foreign object detection based on the joint feature vector, and successively pass through the preliminary positioning sub-module, semantic segmentation verification sub-module, and radar reflection intensity verification sub-module to generate foreign object coordinates and confidence information;
[0035] S6: Trigger an alarm signal according to the foreign object coordinates, and map the foreign object position to the airport pavement geographic information system through a coordinate conversion algorithm to drive the mechanical cleaning device to perform the cleaning task.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] The FOD recognition system of the present invention synchronously collects multi-modal image data using visible light cameras, infrared cameras, and millimeter-wave radars, including visible light images, infrared thermal images, and radar point cloud data. The multi-modal data reflects the airport pavement conditions from different angles. For example, visible light images provide intuitive visual information, infrared thermal images can detect targets in low light or for special objects, and radar point cloud data is used to detect moving targets. The multi-modal data fusion module aligns and fuses these different modal features across modalities, and adopts an adaptive weight allocation strategy to generate a joint feature vector, giving full play to the advantages of each modal data. Compared with single-modal data recognition, the accuracy of foreign object recognition is greatly improved, and the missed detection and false detection rates are effectively reduced.
[0038] The dynamic environment adaptive preprocessing module addresses the problem of image interference in complex environments and employs a rain and fog removal sub-module, a lighting equalization sub-module, and a blur correction sub-module based on a generative adversarial network. The generator of the generative adversarial network adopts a U-Net structure, which can effectively remove rain and fog, equalize lighting, and correct blur, and output high-definition pavement images. This enables the system to still operate stably under harsh conditions such as rain and fog, uneven lighting, and blur, ensuring the accuracy of subsequent feature extraction and recognition, greatly enhancing the system's adaptability to complex environments, and broadening the application scenarios of the system.
[0039] The multi-attention feature extraction module utilizes a spatial attention mechanism, a channel attention mechanism, and a time series attention mechanism. The spatial attention mechanism adjusts local weights according to the importance of image regions to highlight potential FOD regions; the channel attention mechanism learns the interdependence between channels to enhance key channel features; the time series attention mechanism uses a gated recurrent unit to model the temporal correlation of consecutive frames to capture the dynamic changes of foreign objects. The three mechanisms work together to construct a multi-dimensional feature map, comprehensively and accurately reflecting pavement features, improving the feature extraction ability for small, concealed, and dynamically changing foreign objects, and further enhancing the recognition accuracy.
[0040] The multi-level verification and decision-making module performs foreign object detection based on the joint feature vector, including a preliminary positioning sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module. The preliminary positioning sub-module adopts an improved YOLOv8 network architecture to quickly locate potential foreign objects; the semantic segmentation verification sub-module optimizes the boundary accuracy between the pavement and foreign objects based on the DeepLabv3+ network; the radar reflection intensity verification sub-module eliminates false alarm targets through dynamic threshold setting, detection of moving targets, and geometric consistency verification. The multi-level verification cooperates with each other to verify the detection results from different angles, effectively improving the reliability of foreign object detection and ensuring that the system outputs accurate foreign object coordinates and confidence information.
[0041] The real-time alarm and closed-loop control module triggers an alarm signal based on the foreign object coordinates, and maps the position to the airport pavement geographic information system through a coordinate transformation algorithm, driving the mechanical cleaning device to perform the cleaning task. The coordinate transformation algorithm adopts a robust registration method based on RANSAC to ensure the accuracy of coordinate transformation. This closed-loop control process realizes fast processing from detection to alarm and then to cleaning, greatly shortening the residence time of foreign objects on the pavement, effectively reducing safety risks, and improving the safety and operation efficiency of the airport pavement. Description of the Drawings
[0042] Figure 1 It is the working principle diagram of the FOD recognition system based on machine vision described in the present invention;
[0043] Figure 2It is the workflow diagram of the multi-modal data fusion module;
[0044] Figure 3 It is the workflow diagram of the coordinate transformation of the real-time alarm and closed-loop control module;
[0045] Figure 4 It is the step diagram of the radar reflection intensity verification sub-module. Specific implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figures 1-4 , the present invention relates to a FOD recognition system based on machine vision, which aims to improve the accuracy and efficiency of foreign object identification on airport pavements and ensure the safety of airport operations. The specific implementation manners of the present invention are described in detail below. The system includes:
[0048] Multi-spectral image acquisition module: This module synchronously acquires multi-modal image data of the airport pavement by using a visible light camera, an infrared camera, and a millimeter-wave radar. Among them, the visible light camera is responsible for acquiring the visible light image of the airport pavement, which can clearly present details such as the texture and color of the pavement; the infrared camera obtains an infrared thermal image, and through the thermal radiation differences of different objects, it can more clearly identify some objects that are not easily detected under visible light, such as foreign objects with different temperatures from the pavement; the millimeter-wave radar collects radar point cloud data, which uses the principle of electromagnetic wave reflection to accurately measure information such as the position and speed of the target object. Through the synchronous acquisition of these three devices, comprehensive and rich pavement information is obtained, providing a raw data basis for subsequent analysis and processing.
[0049] Dynamic environment adaptive preprocessing module: This module performs super-resolution reconstruction and denoising processing on the acquired multi-modal image data. In a complex airport environment, the image may be affected by problems such as rain, fog, uneven illumination, and blurring. Therefore, multiple sub-modules are required to optimize the image quality. The rain and fog removal sub-module based on the generative adversarial network uses the characteristics of the generative adversarial network, where the generator and the discriminator learn from each other adversarially, effectively removing rain and fog interference in the image; the illumination equalization sub-module adjusts the illumination of the image through a specific algorithm to make the brightness of each part of the image more uniform and enhance the readability of the image; the blur correction sub-module corrects the blurred image caused by equipment jitter or other reasons, and finally outputs a high-definition pavement image, providing high-quality data for subsequent feature extraction and analysis.
[0050] Multi - attention Feature Extraction Module: It performs multi - scale feature fusion on the pre - processed image, and constructs a multi - dimensional feature map of the runway scene through spatial attention mechanism, channel attention mechanism and time - series attention mechanism. The spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through a convolutional kernel, enabling the model to pay more attention to the spatial location where the target object is located; the channel attention mechanism learns the inter - channel dependence relationship through global average pooling and fully - connected layers, and automatically adjusts the importance of different channels; the time - series attention mechanism uses a gated recurrent unit to model the temporal correlation between consecutive frames, capturing the dynamic change information of the object, so as to more accurately identify foreign objects in motion.
[0051] Multi - modal Data Fusion Module: It performs cross - modal alignment and fusion on visible - light image features, infrared thermal - imaging features and radar point - cloud features. First, it realizes the unification of the visible - light image coordinate system, infrared image coordinate system and radar point - cloud coordinate system through an affine transformation matrix, so that data of different modalities can be fused in the same coordinate system. The fusion strategy adopts a multi - modal confidence - weighted method based on cross - entropy. According to the reliability of different modality data in identifying foreign objects, it dynamically adjusts the contribution ratio of each modality feature in the joint feature vector, generates a more representative joint feature vector, and improves the accuracy of recognition.
[0052] Multi - level Verification and Decision - making Module: It performs foreign - object detection based on the joint feature vector, and this module contains multiple sub - modules. The preliminary localization sub - module uses a specific network architecture to generate initial candidate boxes and their confidence scores; the semantic segmentation verification sub - module performs precise segmentation and verification on the runway area and foreign objects based on a specific network; the radar reflection intensity verification sub - module further improves the detection accuracy through the analysis and verification of radar reflection intensity, and finally generates foreign - object coordinates and confidence information.
[0053] Real - time Alarm and Closed - loop Control Module: It triggers an alarm signal according to the foreign - object coordinates, and at the same time maps the foreign - object position to the airport runway geographic information system through a coordinate transformation algorithm. This coordinate transformation algorithm uses a specific robust registration method to ensure the accuracy of coordinate transformation. Then it drives the mechanical cleaning device to perform the cleaning task, realizes the timely handling of foreign objects, and ensures the safety of the airport runway.
[0054] The technical solution of the present invention will be further described in detail below in conjunction with specific embodiments. Embodiment 1
[0055] In this embodiment, the specific implementation of the generative adversarial network in the dynamic - environment adaptive pre - processing module is elaborated in detail. The generator of the generative adversarial network adopts a U - Net structure, which consists of an encoder and a decoder. The role of the encoder is to extract multi - scale features through multiple layers of convolution. Assume the input image is , the -layer convolutional operation in the encoder can be expressed as . After multiple layers of convolution, the size of the image gradually shrinks, and the feature information is concentrated. The decoder restores the high-resolution image through transposed convolution and skip connections. The transposed convolution operation can be expressed as , where represents the number of layers of transposed convolution. The skip connection connects the feature maps of the corresponding layers in the encoder with the corresponding layers in the decoder, enabling the decoder to utilize the feature information of different scales in the encoder and better restore the image details.
[0056] The discriminator adopts the PatchGAN structure, whose purpose is to distinguish between real images and reconstructed images. The discriminator makes a Patch-level judgment on the input images, determining whether each Patch comes from a real image or a reconstructed image generated by the generator. The loss function includes adversarial loss, perceptual loss, and L1 reconstruction loss. The adversarial loss is used to measure the difference between the image generated by the generator and the real image, prompting the generator to generate results closer to the real image. Its formula is: , where represents the real image, is the data distribution of the real image, is the random noise, is the noise distribution, is the discriminator, is the generator. The perceptual loss measures the difference from the semantic and feature levels of the image. It is based on a pre-trained neural network and optimizes the generator by comparing the differences between the generated image and the real image at the network feature layer. The formula is: , where represents the feature extraction function of the -th layer of the pre-trained network, is the number of elements of the feature map of the -th layer. The L1 reconstruction loss is used to ensure the pixel-level similarity between the generated image and the real image. The formula is: . Through the combined action of these three loss functions, the generative adversarial network can effectively perform operations such as rain and fog removal and super-resolution reconstruction on multi-modal image data. Example 2
[0057] This example is used to describe the specific working methods of the spatial attention mechanism, channel attention mechanism, and time series attention mechanism in the multi-attention feature extraction module.
[0058] The spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through the convolution kernel. Let the input feature map be , the spatial attention mechanism first performs average pooling and max pooling operations on the feature map along the channel dimension, obtaining the average pooling feature map and the max pooling feature map . Then, these two feature maps are concatenated, and a convolutional layer is used to generate the spatial attention weight map , and its formula is: , where is the Sigmoid activation function, represents the convolution operation, represents concatenating the two feature maps along the channel dimension. Finally, the generated spatial attention weight map is multiplied by the original feature map to obtain the enhanced feature map , enabling the model to pay more attention to the spatial location where the target object is located.
[0059] The channel attention mechanism learns the inter-channel dependencies through global average pooling and fully connected layers. For the input feature map , global average pooling is performed to obtain the mean vector along the channel dimension, and its calculation formula is: , where and are the height and width of the feature map respectively. Then, the mean vector passes through two fully connected layers. The first fully connected layer reduces the dimension, and the second fully connected layer increases the dimension. Finally, the channel attention weight vector is obtained through the Sigmoid activation function, and the formula is: , where and represent the two fully connected layers respectively, is the ReLU activation function. The channel attention weight vector is weighted with the original feature map along the channel dimension to obtain the enhanced feature map , realizing the adjustment of the importance of different channels.
[0060] The time series attention mechanism uses a gated recurrent unit (GRU) to model the temporal correlation between consecutive frames. Assuming the input sequence of consecutive frame feature maps is , the input of the GRU is the current frame feature map and the hidden state at the previous moment . Inside the GRU, the information flow is controlled by the update gate and the reset gate . The calculation formula of the update gate is: , and the calculation formula of the reset gate is: , where and is the weight matrix, represents concatenating the current frame feature map and the previous hidden state. Then, through the candidate hidden state to calculate the hidden state at the current moment , and the formula is: , , where is the weight matrix, represents element-wise multiplication. Through the processing of the GRU, the temporal correlation between consecutive frames can be effectively captured, improving the recognition ability of moving objects. Example 3
[0061] This example details the specific implementation of the cross-modal alignment and fusion strategy in the multi-modal data fusion module.
[0062] Cross-modal alignment is achieved through an affine transformation matrix to unify the visible light image coordinate system, the infrared image coordinate system, and the radar point cloud coordinate system. For a point in the visible light image, the corresponding point in the infrared image, and the corresponding point in the radar point cloud coordinate system, the affine transformation matrix can transform them to the same coordinate system. Taking the transformation from the visible light image coordinate system to the unified coordinate system as an example, the transformation formula is: , where is the coordinate after transformation to the unified coordinate system. The affine transformation matrix is obtained by calculating the coordinates of known corresponding points. For example, methods such as the least squares method can be used for solution.
[0063] The fusion strategy adopts a multi-modal confidence-weighted method based on cross-entropy. Let the visible light image feature be , the infrared thermal imaging feature be , and the radar point cloud feature be . Their corresponding confidences are respectively. First, calculate the weighting coefficient of each modal feature. Taking the visible light image feature as an example, its weighting coefficient is calculated by the formula: . Then, generate the joint feature vector through weighted summation, and the formula is: . In this way, according to the reliability of different modal data in identifying foreign objects, the contribution ratio of each modal feature in the joint feature vector is dynamically adjusted, improving the recognition accuracy. Example 4
[0064] The coordinate transformation algorithm in the real-time alarm and closed-loop control module adopts a robust registration method based on RANSAC (Random Sample Consensus algorithm) to map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system. Assume there is a series of foreign object coordinate points in the image coordinate system , and there are corresponding coordinate points in the airport geographic coordinate system . The basic idea of the RANSAC algorithm is to randomly sample to select some point pairs to calculate the affine transformation matrix , and then use this matrix to transform all points and calculate the error between the transformed points and the true geographic coordinate points. The error function is defined as: , where is the total registration error of the coordinate transformation, is the affine transformation matrix, is the image coordinate, is the geographic coordinate, is the Huber loss function, is the noise threshold. The expression of the Huber loss function is: , where is a threshold used to balance the calculation method of the error. When the error is small, the squared error is used, and when the error is large, the linear error is used, which can improve the robustness of the algorithm to noise and outliers.
[0065] In practical applications, the RANSAC algorithm continuously repeats the processes of random sampling, calculating the affine transformation matrix and the error, and selects the affine transformation matrix with the smallest error as the final coordinate transformation matrix. In this way, it can accurately map the coordinates of foreign objects in the image coordinate system to the airport geographic coordinate system, providing accurate position information for the subsequent mechanical cleaning device. Example 5
[0066] This example focuses on describing the specific details of the improved YOLOv8 network architecture adopted by the preliminary positioning sub-module. The backbone network adopts the CSPDarknet53 structure, which reduces the computational amount through cross-stage partial connections. The CSPDarknet53 structure divides the feature map of the backbone network into two parts. One part is directly passed to the next stage, and the other part is merged with the directly passed part after a series of convolutional operations. Assume the input feature map is , and after being processed by a certain stage in the CSPDarknet53 structure, the calculation process of the output feature map can be expressed as: , where and They respectively represent different convolution operations. Through this way of cross-stage partial connection, the computational complexity is effectively reduced without losing too much accuracy, and the running efficiency of the network is improved.
[0067] The neck network introduces a cascaded structure of the spatial attention module and the channel attention module to dynamically enhance the feature response of the target region. The working mode of the spatial attention module is as described in Embodiment 2. By generating a spatial attention weight map to weight the feature map. The channel attention module also generates a channel attention weight vector in the same way as in Embodiment 2 . For the input feature map , after being processed by the neck network, the output feature map has the following calculation formula: . Through this cascaded structure, the features of the target region can be more prominent, and the detection accuracy of the target object can be improved.
[0068] The output layer of the prediction head adopts a composite loss function, including CIoU localization loss, Focal classification loss, and object existence confidence loss, to generate initial candidate boxes and their confidence scores. Its expression is: , where is a balance factor used to adjust the importance of different loss functions; is the improved intersection over union (IoU) loss, which, based on the traditional IoU loss, takes into account factors such as the distance, overlap rate, and scale consistency between the predicted box and the ground truth box. The calculation formula is: , where is the IoU, is the center point distance between the predicted box and the ground truth box , is the diagonal length of the smallest closed region containing the two boxes, is the weight coefficient, is a parameter to measure the aspect ratio consistency; is the Focal classification loss used to solve the problem of sample imbalance. The formula is: , where is the number of samples, is the class weight, is the predicted probability, is the Focal parameter; is the confidence loss used to measure the confidence of whether there is an object in the predicted box. Through this composite loss function, the localization and classification accuracy of the network for the target object can be improved. Embodiment 6
[0069] This embodiment elaborates in detail the specific implementation of the semantic segmentation verification sub-module based on the DeepLabv3+ network. The output layer of the segmentation head adopts a hybrid attention mechanism, combining a spatial gating unit and a channel recalibration module to optimize the boundary accuracy between the road surface area and foreign objects. The spatial gating unit analyzes the information of spatial positions to generate spatial gating weights , and weights the feature map in the spatial dimension. Let the input feature map be , and the feature map after being processed by the spatial gating unit is calculated as follows: . The channel recalibration module generates channel recalibration weights by learning the relationships between channels, and weights the feature map in the channel dimension. The formula is: . Through this hybrid attention mechanism, the boundary features between the road surface area and foreign objects can be more accurately highlighted, improving the segmentation accuracy.
[0070] The post-processing unit eliminates segmentation noise through morphological closing filtering and filters out interference regions with an area smaller than a preset threshold using a connected component analysis algorithm. Morphological closing consists of dilation and erosion operations. Assuming the structuring element is , for the binary image after segmentation, the dilation operation expands the object boundaries in the image. Its calculation formula is: , where is the image area covered by the structuring element when the center of the structuring element is translated to the pixel point in the image . The erosion operation contracts the object boundaries inward. The formula is: . The closing operation that first performs dilation and then erosion can fill small holes inside the object and smooth the object boundaries. The formula is: .
[0071] The connected component analysis algorithm marks the connected regions in the image and calculates the area of each region. Let be a certain connected region, and its area is calculated by counting the number of pixel points in the region . Compare the area with the preset threshold . If , then this region is determined as an interference region and filtered, thus further improving the accuracy of the segmentation result. Embodiment 7
[0072] The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on the statistical analysis of historical radar data. The threshold is calculated as follows: . Among them, is the dynamic reflection intensity threshold, which will be dynamically adjusted with the change of time ; is the time window The mean value of the reflection intensity within reflects the average level of the radar reflection intensity during this time period. The calculation method is to divide the sum of all reflection intensity values within the time window by the number of reflection intensity values , that is ; is the standard deviation, which is used to measure the dispersion degree of the reflection intensity values around the mean value. The calculation formula is ; is the confidence coefficient, which is set according to the requirements and experience of the actual scene and is used to adjust the strictness of the threshold. By this way of dynamically setting the threshold, different airport environments and weather conditions can be adapted.
[0073] The moving target detection unit calculates the radial velocity of the foreign object through the Doppler frequency shift and predicts the trajectory by combining the Kalman filter. When the foreign object moves relative to the radar, a Doppler frequency shift will be generated. According to the principle of the Doppler effect, the relationship between it and the radial velocity of the foreign object is , where is the wavelength of the electromagnetic wave emitted by the radar. By measuring the Doppler frequency shift , the radial velocity of the foreign object can be calculated.
[0074] The Kalman filter is used to predict the trajectory of the foreign object. Its state equation is defined as: . Among them, is the state vector, which contains information such as the position and velocity of the foreign object at moment; is the state vector at the next moment; is the state transition matrix, which is determined according to the motion model of the system and is used to describe the transition relationship of the state vector from the current moment to the next moment; is the process noise, which is used to represent the uncertain factors existing in the system. Usually, it is assumed to follow a Gaussian distribution.
[0075] The verification logic unit verifies the geometric consistency of the visual positioning result by constructing an affine transformation equation between the image coordinates and the radar coordinates, and eliminates the false alarm targets with excessive spatial offset. Assume the image coordinates are , and the radar coordinates are , the affine transformation equation is: , where is the affine transformation matrix. By calculating the difference between the coordinates of foreign objects obtained through visual positioning after affine transformation and the radar measurement coordinates, if the difference exceeds a certain threshold, the target is determined as a false alarm target and excluded, thereby improving the accuracy of FOD recognition.
[0076] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.
[0077] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An FOD recognition system based on machine vision, characterized in that, It includes the following modules: A multi-spectral image acquisition module, which is used to synchronously acquire multi-modal image data of the airport pavement through a visible light camera, an infrared camera, and a millimeter wave radar, including visible light images, infrared thermal images, and radar point cloud data; A dynamic environment adaptive preprocessing module, which is used to perform super-resolution reconstruction and denoising processing on the multi-modal image data, including a rain and fog removal sub-module, a lighting equalization sub-module, and a blur correction sub-module based on a generative adversarial network, and outputs a high-definition pavement image; A multi-attention feature extraction module, which is used to perform multi-scale feature fusion on the preprocessed image, including a spatial attention mechanism, a channel attention mechanism, and a time series attention mechanism, and constructs a multi-dimensional feature map of the pavement scene; A multi-modal data fusion module, which is used to perform cross-modal alignment and fusion of visible light image features, infrared thermal image features, and radar point cloud features, and generates a joint feature vector through an adaptive weight allocation strategy; A multi-level verification and decision-making module, which is used to detect foreign objects based on the joint feature vector, including a preliminary positioning sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module, and generates foreign object coordinates and confidence information; A real-time alarm and closed-loop control module, which is used to trigger an alarm signal according to the foreign object coordinates, and maps the foreign object position to the airport pavement geographic information system through a coordinate transformation algorithm, and drives a mechanical cleaning device to perform a cleaning task; The radar reflection intensity verification sub-module includes: The reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on the statistical analysis of historical radar data, and the threshold is calculated as follows: where is the dynamic reflection intensity threshold, is the time window is the mean value of the reflection intensity within the time window, is the standard deviation, is the confidence coefficient; The moving target detection unit calculates the radial velocity of foreign objects through Doppler frequency shift, combines Kalman filtering to predict the trajectory, and the state equation is defined as: Where is the state vector, is the state vector at the next moment, is the state transition matrix, is the process noise; A verification logic unit, which verifies the geometric consistency of the visual positioning result by constructing an affine transformation equation between the image coordinates and the radar coordinates, and eliminates false alarm targets with excessive spatial offset.
2. The system according to claim 1, wherein In the dynamic environment adaptive preprocessing module, the generator of the generative adversarial network adopts a U-Net structure, including an encoder and a decoder. The encoder extracts multi-scale features through multiple layers of convolution, and the decoder restores the high-resolution image through transposed convolution and skip connections; The discriminator adopts a PatchGAN structure to distinguish between real images and reconstructed images; the loss function includes adversarial loss, perceptual loss, and L1 reconstruction loss.
3. The system according to claim 1, wherein In the multi-attention feature extraction module, the spatial attention mechanism dynamically adjusts the local weight distribution of the feature map through a convolutional kernel, the channel attention mechanism learns the inter-channel dependence through global average pooling and a fully connected layer, and the time series attention mechanism uses a gated recurrent unit to model the temporal correlation between consecutive frames.
4. The system according to claim 1, wherein In the multi-modal data fusion module, the cross-modal alignment realizes the unification of the visible light image coordinate system, the infrared image coordinate system, and the radar point cloud coordinate system through an affine transformation matrix, and the fusion strategy adopts a multi-modal confidence weighted method based on cross-entropy to dynamically adjust the contribution ratio of each modal feature in the joint feature vector.
5. The system according to claim 1, characterized in that, In the real-time alarm and closed-loop control module, the coordinate transformation algorithm adopts a robust registration method based on RANSAC to map the coordinates of foreign objects in the image coordinate system to the airport geographical coordinate system. The error function is defined as: where is the total registration error of coordinate transformation, is the affine transformation matrix, is the image coordinate, is the geographical coordinate, is the Huber loss function, is the noise threshold.
6. The system according to claim 1, wherein The preliminary positioning sub-module adopts an improved YOLOv8 network architecture, including: the backbone network adopts a CSPDarknet53 structure to reduce the computational amount through cross-stage partial connection; The neck network introduces a cascaded structure of a spatial attention module and a channel attention module to dynamically enhance the feature response of the target area; The prediction head output layer adopts a composite loss function, including CIoU localization loss, Focal classification loss, and object existence confidence loss, to generate initial candidate boxes and their confidence scores; its expression is: Among them, is the balance factor, is the improved intersection over union loss, is the Focal classification loss, is the confidence loss.
7. The system according to claim 6, characterized in that, The semantic segmentation verification sub-module is constructed based on the DeepLabv3+ network, including: the segmentation head output layer adopts a hybrid attention mechanism, combining a spatial gating unit and a channel recalibration module to optimize the boundary accuracy of the pavement area and foreign objects; The post-processing unit filters out segmentation noise through morphological closing operation filtering and uses a connected component analysis algorithm to filter out interference regions with an area smaller than a preset threshold.
8. The system according to claim 1, characterized in that, In the multi-spectral image acquisition module, the millimeter-wave radar adopts a frequency-modulated continuous wave mode to detect moving foreign objects through Doppler frequency shift and separate them from static targets.
9. A method for identifying FOD based on machine vision, characterized in that, It includes the following steps: S1: Use a visible light camera, an infrared camera, and a millimeter-wave radar to synchronously collect multi-modal image data of the airport pavement, including visible light images, infrared thermal images, and radar point cloud data; S2: Perform super-resolution reconstruction and denoising processing on the collected multi-modal image data. Specifically, operate through a rain and fog removal sub-module, an illumination equalization sub-module, and a blur correction sub-module based on a generative adversarial network to output a high-definition pavement image; S3: Apply a spatial attention mechanism, a channel attention mechanism, and a time series attention mechanism to the preprocessed images for multi-scale feature fusion to construct a multi-dimensional feature map of the pavement scene; S4: Align and fuse the visible light image features, infrared thermal image features, and radar point cloud features across modalities, and adopt an adaptive weight allocation strategy to generate a joint feature vector; S5: Perform foreign object detection based on the joint feature vector, successively passing through a preliminary localization sub-module, a semantic segmentation verification sub-module, and a radar reflection intensity verification sub-module to generate foreign object coordinates and confidence information; S6: Trigger an alarm signal according to the foreign object coordinates, and map the foreign object position to the airport pavement geographic information system through a coordinate transformation algorithm to drive the mechanical cleaning device to perform the cleaning task; The radar reflection intensity verification sub-module includes: A reflection intensity threshold dynamic setting unit generates a scene-adaptive intensity threshold curve based on statistical analysis of historical radar data. The threshold is calculated as follows: where is the dynamic reflection intensity threshold, is the time window is the mean value of the reflection intensity within the time window, is the standard deviation, is the confidence coefficient; Moving target detection unit, calculates the radial velocity of foreign objects through Doppler frequency shift, combines Kalman filtering to predict the trajectory, and the state equation is defined as: Wherein, Is the state vector, Is the state vector at the next moment, Is the state transition matrix, Is the process noise; A verification logic unit that geometrically verifies the visual positioning results by constructing an affine transformation equation between image coordinates and radar coordinates, and eliminates false alarm targets with excessive spatial offset.
Citation Information
Patent Citations
FOD detection method, device and system
CN112285706A
Road safety monitoring method based on LiDAR and thermal infrared visible light information fusion
CN118038226A
Cited By
Gate inclination angle detection method based on machine vision
CN122313380A