Airport runway coverage detection system and method based on radar and vision fusion

CN122672032APending Publication Date: 2026-09-01CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610762717.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

目前,对覆盖物厚度测量主要依靠毫米波雷达进行,以固定的介电常数作为输入进行计算;机场跑道表面往往出现覆盖物共存的现象,因此固定的介电常数不再满足机场跑道检测的要求,缺少一种有效的机场跑道覆盖物厚度与面积测量的方法

Benefits of technology

[0019]有益效果:本发明的基于雷达与视觉融合的机场跑道覆盖物检测系统及方法,通过毫米波雷达与相机的融合,对机场跑道的覆盖物厚度和面积进行检测;在检测覆盖物的厚度时,将相机所采集到的图像数据输入到预测算法模型中,动态输出覆盖物的类型和相应的介电常数,并基于相应覆盖物的介电常数和电磁波传播的双程时间差计算输出覆盖物厚度数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672032A_ABST
    Figure CN122672032A_ABST
Patent Text Reader

Abstract

This invention discloses an airport runway cover detection system and method based on radar and vision fusion. When detecting covers, image data of the covers is acquired and input into a cover prediction and segmentation model. The model outputs the cover category, the cover's position in the image, and the cover mask pixels, and calculates the cover area. Visual guidance is provided to the millimeter-wave radar based on the cover's position in the image. Image data is also input into a cover dielectric constant calculation and prediction model, dynamically outputting the dielectric constant of the corresponding cover. The radar signal is processed to obtain the two-way time difference of electromagnetic wave propagation, and the cover thickness is calculated. This invention uses a camera to acquire image data of the covers, inputs the image data into the algorithm model, dynamically outputs the cover type and corresponding dielectric constant, and then combines this with the echo signal measured by millimeter-wave radar to calculate the area and thickness of the cover.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cover inspection, and more particularly to an airport runway cover inspection system and method based on radar and vision fusion. Background Technology

[0002] Airport runways are the core infrastructure for aircraft takeoff and landing, and their surface condition directly affects aircraft safety. After rain or snow, airport runway surfaces become covered with water, snow, or ice, which reduces runway smoothness, lowers the coefficient of friction, and decreases its anti-skid performance, potentially causing tire slippage and loss of aircraft control. Therefore, airport runway condition monitoring is of great importance to ensuring aviation safety.

[0003] According to the "Rules for Assessment and Notification of Runway Surface Condition at Transport Airports" issued by the Civil Aviation Administration of China, the main coverings on airport runways include water, dry snow, wet snow, slush, and ice. Notification is required when the area and depth of these coverings exceed specified values. Currently, the thickness of these coverings is primarily measured using millimeter-wave radar, with a fixed dielectric constant as input for calculation. However, airport runway surfaces often exhibit a coexistence of various coverings; therefore, a fixed dielectric constant is no longer sufficient for airport runway inspection requirements, and an effective method for measuring the thickness and area of ​​airport runway coverings is lacking. Summary of the Invention

[0004] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides an airport runway cover detection system and method based on radar and vision fusion. The system uses a camera to collect image data of the cover, inputs the image data into an algorithm model, and dynamically outputs the type of the cover and its corresponding dielectric constant. Then, it combines the echo signal measured by millimeter-wave radar to calculate the area and thickness of the cover.

[0005] Technical Solution: To achieve the above objectives, the present invention provides an airport runway cover detection method based on radar and vision fusion. This method constructs a cover prediction and segmentation model and a cover dielectric constant calculation prediction model, and trains both models. When detecting airport runway covers, image data of the covers is acquired using a camera. The image data is input into the cover prediction and segmentation model, which outputs the cover category, the cover's position in the image, and the cover mask pixels. The cover area data is calculated and output based on the cover mask pixels. The millimeter-wave radar is visually guided based on the cover's position in the image, causing the radar to transmit signals towards the corresponding cover. The image data is also input into the cover dielectric constant calculation prediction model, which dynamically outputs the dielectric constant of the corresponding cover based on its type. The radar signal is processed to obtain the two-way time difference of electromagnetic wave propagation. The cover thickness data is calculated and output based on the dielectric constant of the corresponding cover and the two-way time difference of electromagnetic wave propagation.

[0006] Furthermore, the overlay prediction and segmentation model is the RSC-YOLOv8_UT algorithm model, which is an improvement upon the YOLOv8 algorithm. The backbone network of the RSC-YOLOv8_UT algorithm model adopts an improved CSPDarknet-53 network structure, replacing the C2f module of the backbone network with the C2f_UIB module. The C2f_UIB module retains the cross-stage splitting and residual linking framework of the original C2f module and uses the general inversion of the MobileNet4 algorithm. The bottleneck UIB module replaces the original Bottleneck unit based on standard convolution; the neck network of the RSC-YOLOv8_UT algorithm model adopts an improved PANet bidirectional fusion architecture, replacing the C2f module of the neck network with the C2f_UT module; the C2f_UT module retains the cross-stage splitting and residual connection framework of the original C2f module, replaces the original Bottleneck unit based on standard convolution with the general inverted bottleneck UIB module of the MobileNet4 algorithm, and incorporates the Triplet Attention mechanism; the head network of the RSC-YOLOv8_UT algorithm model adopts a three-branch parallel decoupled detection head and independent Proto prototype mask branch architecture, the first branch is the classification branch, the second branch is the bounding box regression branch, and the third branch is the mask coefficient branch.

[0007] Furthermore, the UIB module includes pointwise ascending convolution, pointwise descending convolution, and two optional depthwise convolutions placed before and between the pointwise ascending and descending convolutions, respectively; through the two optional depthwise convolutions, four different lightweight architectures can be formed: IB, ConvNeXt, ExtraDW, and FFN.

[0008] Furthermore, the dielectric constant calculation and prediction model for the covering material is the Perm-YOLOv8_UT algorithm model, which is an improvement upon the YOLOv8 algorithm. The backbone network of the Perm-YOLOv8_UT algorithm model includes an image data input branch and a dielectric constant input branch. The image data branch adopts the backbone network part of the RSC-YOLOv8_UT algorithm model, and the dielectric constant input branch includes one-dimensional batch normalization, a linear fully connected layer, and a SiLU activation function. The neck network of the Perm-YOLOv8_UT algorithm model, based on the neck network of the RSC-YOLOv8_UT deep learning algorithm, adds a neck network cross-modal alignment and fusion module for fusing image features and dielectric constant. The neck network cross-modal alignment and fusion module is abbreviated as CAFM module. The head network of the Perm-YOLOv8_UT algorithm model includes a classification branch, a bounding box regression branch, and a dielectric constant regression branch. The classification branch and the bounding box regression branch use the detection head of YOLOv8 object detection and output the overlay category, category confidence, and bounding box coordinates. The dielectric constant regression branch includes a convolutional CBS module, a Conv convolutional module, and an Avgpool global pooling module, and outputs the predicted dielectric constant of the overlay.

[0009] Furthermore, the neck network of the Perm-YOLOv8_UT algorithm model includes four C2f_UT modules, and the CAFM module is added after the fourth C2f_UT module; the CAFM includes an image feature branch and a dielectric constant feature branch, and the CAFM module includes a global average pooling module, a global feature projection module, a dimension reshaping module, a fully connected layer module, and a dimension reduction convolution module.

[0010] Furthermore, during the training of the Perm-YOLOv8_UT algorithm model, the output of the GAFM module is used as the input of the head network, and during inference, the output of the fourth C2f_UT module is used as the input of the head network.

[0011] Furthermore, the millimeter-wave radar is mounted on a radar turntable, and the visual-guided dynamic positioning of the millimeter-wave radar is achieved by establishing a mapping relationship between the camera coordinate system, the radar turntable coordinate system, and the millimeter-wave radar coordinate system.

[0012] Furthermore, the formula for calculating the area data of the covering is as follows:

[0013] , , ;

[0014] in, The actual area of ​​the covering. This represents the ratio of the actual size in the real-world scenario to the size of the captured image. To capture the area of ​​the mask pixels of the overlay in the image; This represents the total area of ​​the scene captured by the camera. This represents the total pixel area in the image. The coverage rate of the covering.

[0015] Furthermore, the formula for calculating the thickness data of the covering is as follows:

[0016] ; ;

[0017] in, The speed of electromagnetic wave propagation. Let be the speed of light in a vacuum. The dielectric constant of the covering material; To measure the thickness of the covering; This is the time difference between the two-way propagation of electromagnetic waves.

[0018] Furthermore, the airport runway cover detection system based on radar and vision fusion includes an inspection vehicle equipped with a camera, a radar turntable, a millimeter-wave radar, and a frame encoder. The camera is used to acquire image data of the cover, the radar turntable is used to control the movement of the millimeter-wave radar, the millimeter-wave radar is used to acquire echo signal data of the cover, and the frame encoder is used to align the camera and the millimeter-wave radar in spatial position to achieve the correspondence between the image data of the camera and the signal data of the millimeter-wave radar.

[0019] Beneficial effects: The airport runway cover detection system and method based on radar and vision fusion of the present invention detects the thickness and area of ​​airport runway covers by fusing millimeter-wave radar and camera; when detecting the thickness of the cover, the image data collected by the camera is input into the prediction algorithm model, dynamically outputting the type of cover and the corresponding dielectric constant, and calculating and outputting the cover thickness data based on the dielectric constant of the corresponding cover and the two-way time difference of electromagnetic wave propagation. Attached Figure Description

[0020] Appendix Figure 1 This is a side view of the inspection vehicle;

[0021] Appendix Figure 2This is a front view of the inspection vehicle;

[0022] Appendix Figure 3 This is a schematic diagram of the frame encoder structure;

[0023] Appendix Figure 4 Here is a flowchart of the covering detection method;

[0024] Appendix Figure 5 The network structure diagram of the RSC-YOLOv8_UT algorithm is shown below;

[0025] Appendix Figure 6 Here is the structure diagram of the C2f_UIB module;

[0026] Appendix Figure 7 Here is a diagram of the UIB module structure;

[0027] Appendix Figure 8 Here is a diagram of the C2f_UT module structure;

[0028] Appendix Figure 9 The diagram shows the Perm-YOLOv8_UT network structure.

[0029] Appendix Figure 10 Here is a diagram of the CAFM module structure;

[0030] Appendix Figure 11 A schematic diagram illustrating the principle of using millimeter-wave radar to measure the thickness of coverings. Detailed Implementation

[0031] The invention will now be further described with reference to the accompanying drawings.

[0032] As attached Figures 1 to 11 The aforementioned airport runway cover detection method based on radar and vision fusion constructs a cover prediction and segmentation model and a cover dielectric constant calculation and prediction model, and trains the two models.

[0033] When detecting airport runway coverings, camera 2 is used to acquire image data of the coverings; the image data is input into the covering prediction and segmentation model, which outputs the category of the covering, the position of the covering in the image, and the covering mask pixels; the covering area data is calculated and output based on the covering mask pixels; the millimeter-wave radar 6 is visually guided based on the position of the covering in the image, so that the millimeter-wave radar 6 transmits signals toward the corresponding covering.

[0034] Image data is also input into the dielectric constant calculation and prediction model of the covering. The model dynamically outputs the dielectric constant of the covering according to the type of covering, processes the radar signal, and obtains the two-way time difference of electromagnetic wave propagation. Based on the dielectric constant of the covering and the two-way time difference of electromagnetic wave propagation, the model calculates and outputs the thickness data of the covering.

[0035] The overlay prediction and segmentation model is the RSC-YOLOv8_UT algorithm model, which is an improvement on the YOLOv8 algorithm. (See attached image.) Figure 5 As shown, the network structure of the RSC-YOLOv8_UT algorithm model is mainly divided into three parts: the backbone network, the neck network, and the head network.

[0036] The backbone network of the RSC-YOLOv8_UT algorithm model adopts an improved CSPDarknet-53 network structure, replacing the C2f modules with C2f_UIB modules. Specifically, the backbone network of the RSC-YOLOv8_UT algorithm model consists of the following modules arranged in sequence: first convolutional CBS module, second convolutional CBS module, first C2f_UIB module, third convolutional CBS module, second C2f_UIB module, fourth convolutional CBS module, third C2f_UIB module, fifth convolutional CBS module, fourth C2f_UIB module, and spatial pyramid pooling SPPF module.

[0037] The Convolutional Baseline (CBS) module is primarily responsible for extracting features from image data and forms the basis for learning airport runway overlays. The C2f_UIB module lightweights the backbone network, performs further deep feature extraction on the image data, and fuses shallow and deep features. The Spatial Pyramid Pooling (SPPF) module acts as a bridge between the backbone and the neck network, constructing a global receptive field of view while capturing features of both small and large targets.

[0038] As attached Figure 6 As shown, the C2f_UIB module retains the cross-stage splitting and residual linking framework of the original C2f module, and replaces the original Bottleneck unit based on standard convolution with the general inverted bottleneck UIB module of the MobileNet4 algorithm to achieve lightweight feature extraction.

[0039] Specifically, the C2f_UIB module includes a dimensionality reduction convolution unit, a direct connection shortcut branch, a UIB feature extraction unit, a Concat channel concatenation unit, and an up-dimensional convolution unit; firstly, 1 The 1:1 dimensionality reduction convolution compresses the channels of the input feature map, reducing redundant channels. Then, the Split module splits the input feature map into two independent branches at a 1:1 channel dimension. The direct shortcut branch connects directly from input to output, avoiding the gradient vanishing problem in deep learning networks. The UIB feature extraction unit can be stacked n times to extract deep features from the input feature map. The outputs of the two independent branches are merged and concatenated in the Concat channel splicing unit, preserving the original feature information while fusing the extracted deep features. Finally, after passing through a 1:1 channel concatenation module... A dimensionality-reducing convolution of 1 is used to increase the dimensionality so that the number of output channels is the same as the number of input channels.

[0040] With the above improvements, the number of parameters and computational load are greatly reduced without sacrificing feature extraction capabilities, and the ability to extract spatial features from images is enhanced. It can be deployed on edge devices, making it convenient for industrial application.

[0041] The C2f module uses standard convolutions for feature extraction, and the number of parameters for a single standard convolution is [number missing]. ,in The kernel size is the convolution kernel size. and This represents the number of input channels. Therefore, for processing a high-order channel... Width For feature maps, the computational cost is As the number of standard convolutions and channels increases, the computational cost increases quadratically. However, the UIB feature extraction unit in the C2f_UIB module uses depthwise separable convolutions instead of standard convolutions. Depthwise separable convolutions first perform independent spatial convolutions on each input channel, with a parameter count of... Then, for the input Perform pointwise convolution on each channel, with the number of parameters being: Therefore, the number of parameters for a single depth-separable convolution is The ratio of the number of parameters in a single depthwise separable convolution to that of a standard convolution is approximately... Its value is much less than 1. For example, when K=3 and Cout=256, the number of parameters of the depthwise separable convolution is only 11.5% of that of the standard convolution, which is nearly 9 times the number of parameters. The more convolutional layers there are and the more feature maps are processed, the more obvious the effect will be.

[0042] This embodiment employs the ExtraDW architecture, which performs two feature extraction steps compared to traditional standard convolution. The input feature map first undergoes a pre-depthwise separable convolution to extract features such as the edge contours and local textures of the covering in low-dimensional space. Then, it undergoes a pointwise upscaling convolution to increase the dimensionality, and finally, an intermediate depthwise separable convolution extracts secondary spatial features of the overall shape of the covering in high-dimensional space. Simultaneously, the cross-stage splitting and residual connection framework of the original C2f module are retained, providing structural compensation for feature extraction.

[0043] In summary, the C2f_UIB module greatly reduces the number of parameters and computational cost without sacrificing feature extraction capabilities, enabling its RSC-YOLOv8_UT algorithm model to be deployed on edge devices.

[0044] As attached Figure 7 As shown, the UIB module includes pointwise ascending convolution, pointwise descending convolution, and two optional depthwise convolutions placed before and between the pointwise ascending and descending convolutions, respectively. These two optional depthwise convolutions can form four different lightweight architectures: the traditional IB, ConvNeXt, ExtraDW, and FFN. In this embodiment, the ExtraDW architecture is used to replace the original Bottleneck module. The input feature map first undergoes depthwise separable convolution for low-dimensional feature extraction, capturing detailed information of the target in advance. Then, pointwise ascending convolution maps the feature map to a high-dimensional space, enhancing its spatial feature representation capability. Next, depthwise separable convolution is used for high-dimensional feature extraction. Finally, pointwise descending convolution reduces the number of channels to match the number of channels input to subsequent modules.

[0045] The neck network of the RSC-YOLOv8_UT algorithm model adopts an improved PANet bidirectional fusion architecture, and the C2f module of the neck network is replaced with a C2f_UT module. Specifically, the neck network of the RSC-YOLOv8_UT algorithm model includes a first upsampling module, a first fully connected layer module, a first C2f_UT module, a second upsampling module, a second fully connected layer module, a second C2f_UT module, a first convolutional CBS module, a third fully connected layer module, a third C2f_UT module, a second convolutional CBS module, a fourth fully connected layer module, and a fourth C2f_UT module.

[0046] The upsampling module amplifies deep, small feature maps, preparing them for subsequent channel concatenation at different scales. The fully connected layer module concatenates the upsampled semantic features with the features directly input from the backbone network, fusing their information. The C2f_UT module performs depth extraction of target features while maintaining lightweight design, enhancing the extraction of strip-shaped target features. The convolutional CBS module handles channel fusion and dimension alignment, ensuring that different feature maps have the same channel dimension, preparing for feature map concatenation.

[0047] As attached Figure 8 As shown, the C2f_UT module retains the cross-stage splitting and residual linking framework of the original C2f module, replaces the original Bottleneck unit based on standard convolution with the general inverted bottleneck UIB module of the MobileNet4 algorithm, and incorporates the Triplet Attention mechanism to enhance the target features of the overlay, suppress non-feature information, and solve the problem of imbalance between positive and negative samples.

[0048] The original C2f module only possesses general local spatial feature extraction capabilities, failing to distinguish the importance differences between the covering region and the background region, showing insufficient attention to the covering region, resulting in excessive differences between positive and negative samples during training, and the square receptive field of the standard convolution cannot adapt to the striped structure of the covering, thus limiting its adaptability. By replacing the Bottleneck unit with the UIB module, the fused feature map undergoes two enhancement extractions, further extracting discriminative features. Triplet Attention is a three-dimensional cross-attention mechanism. Each of the three branches performs a specific function, establishing pairwise interactive attention dependencies between the channel dimension (C), height dimension (H), and width dimension (W). In the height dimension, it learns the vertical features of the covering object, assigning higher weights to corresponding height ranges to strengthen the vertical features. In the width dimension, it learns the horizontal features of the covering object, assigning higher weights to corresponding width ranges to strengthen the horizontal features. In the channel dimension, it directly weights the spatial location of the covering object, enhancing the corresponding capabilities of the covered area. The three branches can perform weighted processing on the corresponding channel and spatial location based on the actual extension direction of the covering object, assigning higher weights to the covered area, suppressing background information, and widening the difference between background and target features, thereby significantly enhancing the features of the covering object and solving the problem of positive and negative sample imbalance during training.

[0049] Specifically, the C2f_UT module includes a dimensionality reduction convolution unit, a direct connection shortcut branch, a UIB feature extraction unit, a Concat channel concatenation unit, a Triplet Attention mechanism module, and an up-dimensional convolution unit; first, it goes through 1 The 1:1 dimensionality reduction convolution compresses the channels of the input feature map, reducing redundant channels. Then, the Split module splits the input feature map into two independent branches at a 1:1 channel dimension. The direct shortcut branch connects directly from input to output, avoiding the gradient vanishing problem in deep learning networks. The UIB feature extraction unit can be stacked n times to extract deep features from the input feature map. The outputs of the two independent branches are merged and concatenated in the Concat channel splicing unit, preserving the original feature information while fusing the extracted deep features. The fused feature map is then input into the Triplet Attention module for feature enhancement, improving the model's directional awareness. Finally, after passing through a 1:1 channel concatenation module... A dimensionality-reducing convolution of 1 is used to increase the dimensionality so that the number of output channels is the same as the number of input channels.

[0050] Through the above improvements, the neck network achieves lightweight processing and incorporates the Triplet Attention mechanism to enhance the model's feature extraction capabilities.

[0051] The head network of the RSC-YOLOv8_UT algorithm model is the final output of various overlay predictions and instance segmentation. It takes the feature map formed by the neck network as input, performs object detection and pixel-by-pixel instance segmentation tasks, and outputs overlay types, category confidence, bounding boxes, and instance segmentation masks.

[0052] The RSC-YOLOv8_UT algorithm model employs a three-branch parallel decoupled detection head and independent Proto prototype mask branch architecture. The first branch is the classification branch, performing multi-contamination category prediction. The second branch is the bounding box regression branch, accurately predicting the bounding box coordinates of the covered objects and outputting the object's position and size in the image. The third branch is the mask coefficient branch, predicting corresponding mask weight coefficients for each covered object, providing a foundation for accurate instance segmentation. The independent Proto prototype mask branch is a module specific to instance segmentation, taking only the output P3 features from the second C2f_UT in the neck network as input to learn the features of the covered objects and form a shared set of prototype mask base shapes.

[0053] To enable the RSC-YOLOv8_UT deep learning algorithm to predict and segment various coverings, the RSC-YOLOv8_UT deep learning algorithm is trained. Taking the detection of coverings on airport runways as an example, the specific training process includes the following steps.

[0054] Step 1: Create training, validation, and test sets for the cover data. In this embodiment, the original cover data dataset is collected on-site on the airport runway after rain or snow. The collected video frames are sliced ​​at 20-frame intervals to form image data containing different types of cover data. To expand the original dataset, data augmentation methods such as random rotation, adding noise, translation, and Gaussian blur are used to process the original dataset, forming an airport runway cover data dataset containing different cover categories and multiple data augmentation fusions.

[0055] Airport runway cover materials include five categories: dry snow, wet snow, slush, ice, and water accumulation. The Labelme image data annotation software was used to annotate the cover material categories and outlines in the airport runway cover material dataset, generating a JSON file containing the cover material's category and outline coordinates. A Python script was used to convert the JSON format into a TXT format suitable for YOLO algorithm training. Finally, a Python script was used to partition the dataset into training, validation, and test sets in an 8:1:1 ratio.

[0056] Step 2: Set the hyperparameters for the training model and train the RSC-YOLOv8_UT deep learning algorithm. Input the image data from the airport runway overlay dataset in Step 1 into the RSC-YOLOv8_UT deep learning algorithm for training, and save the optimal model and the model from the last iteration. Set the training epochs to 1000 epochs and the image size to 640. The batch size for each iteration is set to 8, and a dynamic learning rate is used. Cosine annealing learning rate scheduling is enabled, with an initial learning rate of 0.005 and a final learning rate of 0.01. After setting the hyperparameters for training the network, GPU training is performed in a YOLOv8 virtual environment. After training, the best overlay prediction and segmentation model can be used for prediction, outputting the overlay category and overlay mask pixels.

[0057] The dielectric constant calculation and prediction model for the covering material is the Perm-YOLOv8_UT algorithm model, which is an improvement on the YOLOv8 algorithm. The network structure of the Perm-YOLOv8_UT algorithm model is mainly divided into three parts: the backbone network, the neck network, and the head network.

[0058] As attached Figure 9As shown, the backbone network of the Perm-YOLOv8_UT algorithm model has two input branches: an image data input branch and a dielectric constant input branch. The image data branch uses the backbone network portion of the RSC-YOLOv8_UT algorithm model to extract features from the overlay image, and its output is input to the neck network. The dielectric constant input branch includes a one-dimensional batch normalized layer (BatchNorm1d), a linear fully connected layer (Linear), and a SiLU activation function.

[0059] The dielectric constant of the covering material is input as a single-dimensional vector. This vector is normalized using BN1d to prevent excessive differences in the dielectric constant's numerical range, which could affect training convergence stability and speed. The single-dimensional dielectric constant vector is then increased in dimensionality using Linear, enabling the model to learn higher-dimensional features of the dielectric constant. This is followed by further normalization using BN1d to accelerate training convergence. Finally, a non-linear SiLU activation function is introduced, allowing the model to learn more complex non-linear mappings and enhancing its generalization ability. This process from Linear to SiLU activation is superimposed twice to obtain the output of the dielectric constant branch, which is then used as input to the Cross-modal Alignment and Fusion Module of the neck network.

[0060] The neck network of the Perm-YOLOv8_UT algorithm model, based on the neck network of the RSC-YOLOv8_UT deep learning algorithm, adds a neck network cross-modal alignment and fusion module for fusing image features and dielectric constant. This neck network cross-modal alignment and fusion module is abbreviated as CAFM module.

[0061] The neck network of the Perm-YOLOv8_UT algorithm model includes four C2f_UT modules, with the CAFM module added after the fourth C2f_UT module. (See attached diagram) Figure 10 As shown, the CAFM module includes a global average pooling module, a global feature projection module, a dimension reshaping module, a fully connected layer module, and a dimension reduction convolution module, wherein the global feature projection module is abbreviated as GFPM module.

[0062] The CAFM module includes an image feature branch and a dielectric constant feature branch. The image feature map output from the fourth C2f_UT module is used as input to the image feature branch. A global average pooling module extracts the global vector from the image feature map, which is then input to the projection layer via the GFPM module. The dielectric constant feature vector output from the backbone network is used as input to the dielectric constant feature branch. The GFPM module inputs the dielectric constant vector to the projection layer, where it undergoes spatial broadcasting of the dielectric constant feature vector through a dimension reshaping module, aligning its spatial dimension with that of the image feature vector. A Concat fully connected layer then concatenates and fuses the image feature vector and the reshaped dielectric constant feature vector, integrating them into the image features so that each pixel in the image shares the same dielectric constant. Finally, a dimension reduction convolution (CBS) is applied to compress the number of channels, matching the input channel number of subsequent modules.

[0063] The GFPM module includes a linear fully connected layer (Linear) and a one-dimensional batch normalization layer (BatchNorm1d, BN1d), which maps the global feature vector and dielectric constant feature vector of the image to the same space, facilitating feature fusion.

[0064] The head network of the Perm-YOLOv8_UT algorithm model is the task execution end that predicts the category of various coverings and outputs the corresponding dielectric constant of the coverings. The head network of the Perm-YOLOv8_UT algorithm model includes a classification branch, a bounding box regression branch, and a dielectric constant regression branch. The classification and bounding box regression branches use the detection head of YOLOv8 object detection, outputting the covering category, category confidence, and bounding box coordinates. The newly added dielectric constant regression branch includes a convolutional CBS module, a Conv convolutional module, and an Avgpool global pooling module, outputting the predicted dielectric constant of the covering.

[0065] It should be noted that when training the Perm-YOLOv8_UT algorithm model, the output of the GAFM module is used as the input to the head network. During inference, the outputs of the second, third, and fourth C2f_UT modules are used as the input to the head network. To train the Perm-YOLOv8_UT deep learning network to learn the mapping relationship between the dielectric constant and the overlay image, the input model has two parts of data: the dielectric constant and the overlay image. These two types of data are fused by the GAFM module and used as the output of the GAFM module, which is then input to the detection head. However, during inference, the trained model is used, and only the overlay image is input. Therefore, there is no fusion process with the dielectric constant, and the head network only receives the outputs of the second, third, and fourth C2f_UT modules.

[0066] To enable the Perm-YOLOv8_UT deep learning network to identify various coverings and output the dielectric constant of the corresponding coverings, the Perm-YOLOv8_UT deep learning network is trained. Taking the detection of airport runway coverings as an example, the specific training process includes the following steps.

[0067] Step 1: Construct a paired dataset of cover image data and dielectric constant. In this embodiment, the cover image dataset includes a standard cover dataset prepared in the laboratory and an image dataset collected in the field after rain and snow. The dielectric constant of the corresponding cover is calculated by using the thickness measurement principle of millimeter-wave radar to obtain the two-way time difference of electromagnetic wave propagation and the manually measured cover thickness. The dielectric constant is calculated for both the standard dataset prepared in the laboratory and the dataset collected in the field. The standard dataset prepared in the laboratory serves as a verification of the field-collected data. Each image data contains only one type of cover, and the dielectric constant data is saved as a txt file named after the image data, forming a dataset where each cover is paired with its dielectric constant. The image data annotation software LabelImage is used to annotate the cover category and location in the airport runway cover dataset, generating a txt file containing the cover category and location coordinates. Finally, a dataset partitioning script is written using Python to divide the airport runway cover dataset into a training set, a validation set, and a test set in an 8:1:1 ratio.

[0068] Step 2: Set training hyperparameters and train the Perm-YOLOv8_UT algorithm. Input the image data from the airport runway overlay dataset in Step 1 into the RSC-YOLOv8_UT deep learning algorithm for training, and save the optimal model and the model from the last iteration; set the training epochs to 1000 epochs and the image size to 640. The batch size for each iteration is set to 8, and the learning rate is dynamically adjusted. Cosine annealing learning rate scheduling is enabled, with an initial learning rate of 0.005 and a final learning rate of 0.01. After setting the hyperparameters for training the network, GPU training is enabled in the YOLOv8 virtual environment. First, the image input branch in the backbone network is frozen, and the dielectric constant input branch is warmed up to allow the model to learn the mapping relationship between image features and dielectric constant. After 15 rounds, all networks are trained. After training, the best dielectric constant prediction model for the covering can be used to predict and output the covering category and the dielectric constant of the covering.

[0069] The airport runway cover detection system based on radar and vision fusion includes an inspection vehicle 1 as the basic carrier, which is equipped with a camera 2, a radar turntable 5, a millimeter-wave radar 6, and a frame encoder 4. The millimeter-wave radar 6 is used to collect echo signal data of the cover. The camera 2 is a visible light camera used to collect image data of the cover. The millimeter-wave radar 6 is mounted on the radar turntable 5, which controls the movement of the millimeter-wave radar 6. The frame encoder 4 is used to align the camera 2 and the millimeter-wave radar 6 in space, achieving a correspondence between the image data from the camera 2 and the signal data from the millimeter-wave radar 6.

[0070] The frame encoder 4 provides a position reference for the sampling data of the visible light camera and millimeter-wave radar 6 by measuring the mileage of the inspection vehicle 1. The output shaft of the frame encoder 4 mates with the hub hole of the measuring wheel 13, converting the angular displacement of the measuring wheel 13 into pulse signals. The total number of pulses output by the frame encoder is linearly related to the number of revolutions of the measuring wheel 13.

[0071] ;

[0072] in, This represents the total number of pulses output by frame encoder 4. This refers to the number of pulses output per revolution by frame encoder 4, expressed in pulses per revolution. The rotational speed of wheel 13 is measured in r / s.

[0073] The mileage of the inspection vehicle 1 can be calculated by the total number of pulses output by the frame encoder 4 and the circumference of the measuring wheel 13:

[0074] ;

[0075] in, The distance traveled by inspection vehicle 1 is in meters. The measurement is for the circumference of wheel 13, in meters (m).

[0076] The frame encoder 4 sends the number of pulses to the industrial control computer in real time. When the total number of pulses output by the frame encoder 4 reaches the set threshold, the industrial control computer simultaneously sends a data sampling command to the visible light camera and the millimeter-wave radar 6, so that the visible light camera and the millimeter-wave radar 6 collect data in the same position, thereby aligning the visible light camera and the millimeter-wave radar 6 in spatial position.

[0077] As attached Figure 1 and 2 In one embodiment shown, both the camera 2 and the radar turntable 5 are fixed on the mounting platform 3, which is placed on the roof of the inspection vehicle 1. The frame encoder 4 is mounted on the rear side of the inspection vehicle 1. The mounting height and angle of the camera 2 can be adjusted according to the height of the inspection vehicle 1 and the actual conditions of the airport runway.

[0078] When conducting airport runway covering inspections, inspection vehicle 1 conducts cyclical inspections along the centerline of the airport runway. Through the detection devices and systems mounted on inspection vehicle 1, it detects the multiple coverings that cover the airport runway, collecting information on their location, thickness, and area. This data serves as a reliable basis for airport runway status early warning and dispatching.

[0079] As attached Figure 3 As shown, the frame encoder 4 includes a mounting bracket 10, a support plate 8, an electric push rod 11, a measuring wheel 13, an encoder rocker arm 15, and a differential encoder 7. The frame encoder 4 is used to align the camera 2 and the millimeter-wave radar 6 in space, achieving a one-to-one correspondence between the image data of the visible light camera and the signal data of the millimeter-wave radar 6, forming paired data. The mounting bracket 10 is fixed to the rear side of the inspection vehicle 1 to support the entire frame encoder 4. The support plate 8 is welded to the mounting bracket 10. One end of the encoder rocker arm 15 is connected to the first turntable base 9 on the support plate 8, and the other end of the encoder rocker arm 15 is connected to the fixed plate 14 via a rotating shaft. The differential encoder 7 is mounted on the fixed plate 14, and the output shaft of the differential encoder 7 is fixedly connected to the measuring wheel 13 to achieve rotation at the same speed. The bottom end of the electric push rod is connected to the second turntable base 12 on the support plate 8, and the front end of the electric push rod is connected to the encoder rocker arm 15 via a pin. When the inspection vehicle 1 is in operation, the rocker arm swings under the action of the electric actuator, thereby keeping the measuring wheel 13 in close contact with the ground, and can be adjusted in real time according to the terrain slope. After the inspection vehicle 1 completes its operation, the rocker arm is retracted to avoid unnecessary friction on the measuring wheel 13.

[0080] The millimeter-wave radar 6 is mounted on the radar turntable 5. Therefore, the dynamic positioning of the millimeter-wave radar 6 guided by vision is achieved by establishing a mapping relationship between the camera coordinate system, the radar turntable coordinate system, and the millimeter-wave radar coordinate system. Taking a visible light camera as an example, the specific steps for establishing the mapping relationship are as follows.

[0081] Step 1: Visible Light Camera Calibration. The Zhang calibration method is used to calibrate the visible light camera, obtaining the intrinsic and extrinsic parameter matrices of camera 2, as well as the distortion coefficients of the visible light camera lens. The visible light camera is fixed, and the checkerboard calibration board is placed at different heights, distances, and angles to capture 20 valid checkerboard images. The captured images are filtered and grayscaled to improve image quality. The Harris corner detection algorithm is used to extract the feature points of the checkerboard, and sub-pixel corner refinement is used to improve the accuracy of corner coordinates. The world coordinates and corresponding pixel coordinates of the inner corners of the checkerboard calibration board are used as input to solve for the intrinsic and extrinsic parameter matrices of camera 2, as well as the distortion coefficients of the visible light camera lens.

[0082] Step 2: Define the coordinate system. The core purpose of this step is to unify the standards and rules of the coordinate system and eliminate ambiguities in its definition. All coordinate systems follow the right-hand Cartesian coordinate system rule and the right-hand rotation rule. The specific coordinate systems are defined as follows: Camera coordinate system: The origin is the center of the lens of camera 2. The z-axis is forward along the optical axis of camera 2, the x-axis is parallel to the imaging plane and to the right, and the y-axis is parallel to the imaging plane and perpendicular to the downward direction. This coordinate system is used to describe the position of the imaging target relative to camera 2. Radar turntable coordinate system: The origin is the orthogonal intersection of the horizontal rotation axis and the elevation axis of the turntable. The z-axis is along the normal to the front of the turntable, and the x-axis coincides with the horizontal rotation axis, upward, and counterclockwise rotation around the x-axis is the positive direction. The y-axis coincides with the elevation axis, upward, and rotation is the positive direction.

[0083] Step 3: Establish the mapping relationship between the camera coordinate system and the radar turntable coordinate system. A one-to-one mapping relationship between the camera coordinate system and the radar turntable coordinate system is established using hand-eye calibration, enabling the turntable to dynamically rotate based on image data from the visible light camera. The checkerboard calibration plate is placed horizontally on the airport runway, ensuring it appears in the center of the visible light camera's field of view and occupies 25%-50% of the entire image area, ensuring the accuracy of feature point extraction. The radar turntable 5 is controlled to rotate in 20 different preset attitudes. The horizontal rotation angle of the radar turntable 5 ranges from -60° to 60°, and the pitch angle ranges from -30° to 30°, with a 5° difference between adjacent attitudes. For each attitude, a checkerboard image is acquired through the visible light camera, and the horizontal rotation angle and pitch angle of the current attitude are recorded by the angle encoder of the radar turntable 5, forming paired data between the checkerboard image and the radar turntable 5's rotation angle. The rigid transformation matrix between the camera coordinate system and the radar turntable coordinate system is solved using the 20 sets of paired data to obtain the mapping relationship between the two.

[0084] The mapping relationship between the camera coordinate system and the radar turntable coordinate system established in this embodiment can be directly used for the dynamic positioning of the millimeter-wave radar 6 using visual guidance of the covering. When the visible light camera captures the covering, the industrial control computer can calculate the required rotation angle of the radar turntable 5 through the established coordinate system mapping relationship, and send the rotation angle data to the angle encoder of the radar turntable 5 to realize the dynamic rotation of the radar turntable 5, so that the electromagnetic beam emitted by the millimeter-wave radar 6 is aligned with the position coordinate center of the covering, making the detection range of the millimeter-wave radar 6 more accurate and the detection efficiency higher.

[0085] The core formula for area calculation is used to obtain the actual area of ​​the covering and the ratio of the covering area to the actual scene area. For example, the ratio of the covering area to the airport runway area, i.e., the coverage rate of the covering, is as follows:

[0086] , , ;

[0087] in, The actual area of ​​the covering. This represents the ratio of the actual size of an airport runway in a real-world scenario to the size of the captured image. To capture the area of ​​the mask pixels of the overlay in the image; The total area of ​​the airport runway scene captured by camera 2. This represents the total pixel area in the image. The coverage rate of the covering.

[0088] When processing the radar signal of the echo signal, the conventional FMCW millimeter-wave radar 6-echo processing method in this field is adopted, but not limited to conventional radar processing methods, in order to extract the two-way propagation time difference of the reflected echo.

[0089] As attached Figure 11 In one embodiment shown, the millimeter-wave radar 6 measures the thickness of airport runway coverings by emitting high-frequency electromagnetic waves. The millimeter-wave radar 6 emits high-frequency millimeter waves towards the airport runway. When the inspection vehicle 1 detects a covering, the covering interface scatters the millimeter waves, forming a reflected echo A1. Simultaneously, the millimeter waves penetrate the contaminated covering and illuminate the airport runway surface. The airport runway also scatters the millimeter waves and penetrates the upper interface of the covering, forming a reflected echo A2. The millimeter-wave radar 6 transmits the echoes to an industrial control computer for radar signal processing, extracting the two-way propagation time difference between the two reflected echoes.

[0090] The core formula for thickness measurement is derived from the thickness measurement principle of millimeter-wave radar 6. The electromagnetic waves emitted by millimeter-wave radar 6 propagate in a straight line in a uniform magnetic medium, and their propagation speed is only related to the dielectric constant of the medium they pass through. Therefore, the formula for calculating the thickness data of the covering is as follows:

[0091] ; ;

[0092] in, The speed of electromagnetic wave propagation. Let be the speed of light in a vacuum. The dielectric constant of the covering material; To measure the thickness of the covering; This is the time difference between the two-way propagation of electromagnetic waves.

[0093] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting airport runway cover based on radar and vision fusion, characterized in that: Construct a prediction and segmentation model for the covering material and a prediction model for the dielectric constant of the covering material, and train both models. When detecting airport runway coverings, the camera (2) is used to collect image data of the coverings; the image data is input into the covering prediction and segmentation model, and the covering prediction and segmentation model outputs the category of the covering, the position of the covering in the image, and the covering mask pixels; the covering area data is calculated and output based on the covering mask pixels; the millimeter-wave radar (6) is visually guided based on the position of the covering in the image, so that the millimeter-wave radar (6) transmits signals toward the corresponding covering; Image data is also input into the dielectric constant calculation and prediction model of the covering. The model dynamically outputs the dielectric constant of the covering according to the type of covering, processes the radar signal, and obtains the two-way time difference of electromagnetic wave propagation. Based on the dielectric constant of the covering and the two-way time difference of electromagnetic wave propagation, the model calculates and outputs the thickness data of the covering.

2. The airport runway coverage detection method based on radar and vision fusion according to claim 1, characterized in that: The overlay prediction and segmentation model is the RSC-YOLOv8_UT algorithm model, which is an improvement on the YOLOv8 algorithm. The backbone network of the RSC-YOLOv8_UT algorithm model adopts an improved CSPDarknet-53 network structure, replacing the C2f module of the backbone network with the C2f_UIB module. The C2f_UIB module retains the cross-stage splitting and residual linking framework of the original C2f module, and replaces the original Bottleneck unit based on standard convolution with the general inverted bottleneck UIB module of the MobileNet4 algorithm. The neck network of the RSC-YOLOv8_UT algorithm model adopts an improved PANet bidirectional fusion architecture, replacing the C2f module of the neck network with the C2f_UT module. The C2f_UT module retains the cross-stage splitting and residual connection framework of the original C2f module, replaces the original Bottleneck unit based on standard convolution with the general inverted bottleneck UIB module of the MobileNet4 algorithm, and incorporates the Triplet Attention mechanism. The head network of the RSC-YOLOv8_UT algorithm model adopts a three-branch parallel decoupled detection head and independent Proto prototype mask branch architecture. The first branch is the classification branch, the second branch is the bounding box regression branch, and the third branch is the mask coefficient branch.

3. The airport runway coverage detection method based on radar and vision fusion according to claim 2, characterized in that: The UIB module includes pointwise ascending convolution, pointwise descending convolution, and two optional depthwise convolutions placed before and between the pointwise ascending and descending convolutions, respectively. Through the two optional depthwise convolutions, four different lightweight architectures can be formed: IB, ConvNeXt, ExtraDW, and FFN.

4. The airport runway cover detection method based on radar and vision fusion according to claim 2, characterized in that: The dielectric constant calculation and prediction model for the covering material is the Perm-YOLOv8_UT algorithm model, which is an improvement on the YOLOv8 algorithm. The backbone network of the Perm-YOLOv8_UT algorithm model includes an image data input branch and a dielectric constant input branch. The image data branch follows the backbone network part of the RSC-YOLOv8_UT algorithm model. The dielectric constant input branch includes one-dimensional batch normalization, a linear fully connected layer, and a SiLU activation function. The neck network of the Perm-YOLOv8_UT algorithm model, based on the neck network of the RSC-YOLOv8_UT deep learning algorithm, adds a neck network cross-modal alignment fusion module for fusing image features and dielectric constant; the neck network cross-modal alignment fusion module is abbreviated as CAFM module; The head network of the Perm-YOLOv8_UT algorithm model includes a classification branch, a bounding box regression branch, and a dielectric constant regression branch. The classification and bounding box regression branches use the same detection head as YOLOv8's object detection, outputting the overlay category, category confidence, and bounding box coordinates; the dielectric constant regression branch includes a convolutional CBS module, a Conv convolutional module, and an Avgpool global pooling module, outputting the predicted dielectric constant of the overlay.

5. The airport runway cover detection method based on radar and vision fusion according to claim 4, characterized in that: The neck network of the Perm-YOLOv8_UT algorithm model includes four C2f_UT modules, and the CAFM module is added after the fourth C2f_UT module. The CAFM includes an image feature branch and a dielectric constant feature branch, and the CAFM module includes a global average pooling module, a global feature projection module, a dimension reshaping module, a fully connected layer module, and a dimension reduction convolution module.

6. The airport runway cover detection method based on radar and vision fusion according to claim 4, characterized in that: When training the Perm-YOLOv8_UT algorithm model, the output of the GAFM module is used as the input of the head network, and during inference, the output of the fourth C2f_UT module is used as the input of the head network.

7. The airport runway coverage detection method based on radar and vision fusion according to claim 1, characterized in that: The millimeter-wave radar (6) is mounted on the radar turntable (5). The dynamic positioning of the vision-guided millimeter-wave radar (6) is achieved by establishing a mapping relationship between the camera coordinate system, the radar turntable coordinate system and the millimeter-wave radar coordinate system.

8. The airport runway cover detection method based on radar and vision fusion according to claim 1, characterized in that: The formula for calculating the area of ​​the covering is: , , ; in, The actual area of ​​the covering. This represents the ratio of the actual size in the real-world scenario to the size of the captured image. To capture the area of ​​the mask pixels of the overlay in the image; The total area of ​​the scene captured by camera (2), This represents the total pixel area in the image. The coverage rate of the covering.

9. The airport runway coverage detection method based on radar and vision fusion according to claim 1, characterized in that: The formula for calculating the thickness of the covering material is: ; ; in, The speed of electromagnetic wave propagation. Let be the speed of light in a vacuum. The dielectric constant of the covering material; To measure the thickness of the covering; This is the two-way propagation time difference of electromagnetic waves.

10. The airport runway cover detection system based on radar and vision fusion according to any one of claims 1 to 9, characterized in that: The system includes an inspection vehicle (1), which is equipped with a camera (2), a radar turntable (5), a millimeter-wave radar (6), and a frame encoder (4). The camera (2) is used to collect image data of the cover, the radar turntable (5) is used to control the activity of the millimeter-wave radar (6), the millimeter-wave radar (6) is used to collect echo signal data of the cover, and the frame encoder (4) is used to align the camera (2) and the millimeter-wave radar (6) in space to achieve the correspondence between the image data of the camera (2) and the signal data of the millimeter-wave radar (6).