Lightweight detection method for airport scene target
By introducing RCBottleNeck and MPConv into the YOLOV5 network, combining simOTA loss function and heating-distillation training, the problem of difficulty in compatibility with speed and accuracy in airport scene monitoring is solved, and the rapid accuracy of lightweight detection is achieved.
Patent Information
- Application Number
- CN202510225887.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-18
AI Technical Summary
In airport scene monitoring technology, existing methods are difficult to find a balance between ensuring speed and accuracy, resulting in the inability to monitor abnormal situations in real time or the problem of degraded detection quality.
Based on the YOLOV5 network, a repeat crossover bottleneck layer (RCBottleNeck) and a multi-level perceptual convolutional layer (MPConv) are designed, combining the improved simOTA loss function and heating-distillation cycle training process to optimize the lightweight model to improve detection speed and accuracy.
It realizes the rapid and accurate detection of airport scene targets on edge devices, and improves the real-time monitoring capabilities of safe and stable operation of the airport.
Smart Images

Figure CN120339970A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of airport surface monitoring, and relates to a lightweight detection method for airport surface targets. Background Art
[0002] With the development of the civil aviation industry, the demand for air travel and freight transportation has been increasing year by year. The sharp increase in the number of flights has put forward high requirements for the safe and stable operation of airports. Before takeoff and after landing, aircraft need to receive various ground services on the apron, such as refueling, passenger shuttle, cargo transportation, and cleaning services. These services involve the spatio-temporal transfer and complex interactions of various entities such as vehicles, personnel, goods, and luggage. With the increase in takeoff and landing flights, the airport flight area environment has become increasingly complex, and airport entities show the characteristics of large-scale and high-frequency movement, which greatly increases the difficulty of surface monitoring and operation and maintenance management.
[0003] Currently, the problem faced by airport surface monitoring technology is that it is difficult to balance accuracy and speed. The most common means of airport surface monitoring is to use video surveillance cameras to obtain data and then rely on algorithms for target detection. Due to the limited computing power of cameras, large algorithms cannot be deployed. One method is to deploy algorithms at the backend and transmit the collected images to the backend for processing. This method can ensure recognition accuracy and monitoring accuracy. However, in practical applications, due to limitations in data transmission speed, transmission delay, and transmission bandwidth, this distributed data collection and centralized data processing have bottlenecks in processing speed and computing resources, making it difficult to monitor sudden abnormal situations in real time and being unfavorable for the speed of airport operation and maintenance. To ensure speed, another method is to lightweight the algorithm for deployment at the camera edge end, but this often leads to a reduction in detection quality and is prone to delays or missed detections in the detection of abnormal phenomena, abnormal behaviors, or small targets. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a lightweight detection method for airport surface targets in view of the deficiencies of the prior art, including the following steps:
[0005] Step 1, design a repeated cross bottleneck layer RCBottleNeck;
[0006] Step 2, design a multi-level perception convolutional layer MPConv;
[0007] Step 3, replace the corresponding parts in the YOLOV5 network structure with the repeated cross bottleneck layer RCBottleNeck and the multi-level perception convolutional layer MPConv to form a lightweight YOLOV5 network;
[0008] Step 4: Design an improved simplified optimal transport assignment (simOTA) loss function for the lightweight YOLOV5 network;
[0009] Step 5: Design a heating and distillation cyclic training process for the lightweight YOLOV5 network to obtain a trained lightweight YOLOV5 network;
[0010] Step 6: Use the trained lightweight YOLOV5 network to perform object detection on the airport scene to complete the lightweight detection of airport scene objects.
[0011] Step 1 includes: The repeated cross bottleneck layer (RCBottleNeck) replaces all 3×3 convolutions in the Bottleneck module of the C3 structure in the BackBone network with a 3×1 convolution followed by a 1×3 convolution.
[0012] Step 2 includes:
[0013] For the feature with an input channel size of CH1, the multi-level perception convolutional layer (MPConv) first uses a standard convolution to map the feature to a global feature with a channel size of Then, it uses a depthwise separable convolution to obtain a pixel feature with a channel size of At the same time, it performs grouped convolution with a group number of g to obtain a local feature with a channel of After concatenating the global feature, pixel feature, and local feature and performing a shuffle (i.e., interleaving and combining the three features in the order of global - pixel - local to form a feature value with a channel size of CH2), a feature value with a channel of CH2 is obtained, and the time complexity GFLOPs GSConv ~O is:
[0014]
[0015] where W and H are the feature width and height respectively, and K1 and K2 are the convolution kernel sizes.
[0016] Step 3 includes: The YOLOV5 network structure includes an input end, a backbone network (Backbone), a neck network (Neck), and a detection head (Head);
[0017] The backbone network (Backbone) is used to extract image features, the detection head (Head) is used to output prediction boxes, i.e., the results of object detection, and the neck network (Neck) is a transitional part connecting the backbone network (Backbone) and the detection head (Head), which is used to better utilize the features extracted by the backbone network;
[0018] The Neck network includes a standard convolutional layer and an upsampling layer;
[0019] Replace the standard convolutional layer with a multi-level perception convolutional layer MPConv;
[0020] The Backbone network includes a C3 module, which is used to learn residual features, and replace the Bottleneck in the C3 module of the Backbone network with a repeated cross bottleneck layer RCBottleNeck.
[0021] Step 4 includes:
[0022] Step 4-1, pre-screen the prediction boxes;
[0023] Step 4-2, calculate the classification loss and regression loss for the pre-screened prediction boxes and all ground truth boxes to obtain a cost matrix C and an intersection over union iou matrix M;
[0024] Step 4-3, traverse each ground truth box, and obtain the number k of corresponding positive prediction boxes from the cost matrix C j ;
[0025] Step 4-4, reassign prediction boxes for each ground truth box to form an assignment matrix P;
[0026] Step 4-5, further screen the prediction boxes corresponding to different ground truth boxes that are repeated to ensure that one prediction box only corresponds to one ground truth box;
[0027] Step 4-6, obtain positive samples and corresponding ground truth boxes through Steps 4-1 to 4-5, and classify the remaining prediction boxes as negative samples to obtain the loss function.
[0028] Step 4-1 includes:
[0029] Draw a box of 160 pixels × 160 pixels (currently, the network outputs a 1×1 feature value for an input image of 32×32, and in this algorithm, a 5×5 box is taken around the center of the final feature value (that is, 2 feature values are taken outward from the center feature value in the up, down, left, and right directions), and the original pixel value is deduced to be 160 pixels × 160 pixels) at the center of the ground truth box, calculate the intersection area between the box and the ground truth box, and only the prediction boxes that have an intersection with the intersection area are the pre-screened prediction boxes.
[0030] Step 4-2 includes:
[0031] The cost c between the i-th pre-screened prediction box and the j-th ground truth box ij is defined as:
[0032]
[0033] where, is the classification loss generated by the i-th pre-screened prediction box and the j-th ground truth box, is the regression loss generated by the prediction box and the ground truth box, and σ is the proportionality parameter;
[0034] The cost c calculated by traversing all pre-screened prediction boxes and ground truth boxes ij , obtaining the cost matrix C;
[0035] Traverse all pre-screened prediction boxes and all ground truth boxes to calculate the intersection over union (IoU), obtaining the IoU matrix M. The element m in the i-th row and j-th column of the IoU matrix M ij represents the IoU between the i-th pre-screened prediction box and the j-th ground truth box;
[0036] The intersection over union m ij is calculated as follows: Find the intersection area between the i-th pre-screened prediction box and the j-th ground truth box. The intersection area is the overlapping part of the i-th pre-screened prediction box and the j-th ground truth box. Calculate the area of the intersection region: Find the total area covered by the i-th pre-screened prediction box and the j-th ground truth box, which includes the parts unique to each box plus their common intersection part; Divide the area of the intersection by the area of the union to obtain the intersection over union.
[0037] Step 4-3 includes:
[0038] For the j-th ground truth box, take the top X1 (usually 10) prediction boxes with the largest IoU with the j-th ground truth box. Calculate the IoU between the X1 prediction boxes and the j-th ground truth box, accumulate the results and round up. The obtained result is the number k of positive prediction boxes corresponding to the j-th ground truth box j .
[0039] Step 4-4 includes:
[0040] Traverse all ground truth boxes. For the j-th ground truth box, query the cost matrix and assign the top k j prediction boxes with the smallest cost to the j-th ground truth box;
[0041] Traverse all pre-screened prediction boxes and all ground truth boxes to check the assignment relationship (because in the previous step "Traverse all ground truth boxes. For the j-th ground truth box, query the cost matrix and assign the top k j prediction boxes with the smallest cost to it;") the k j prediction boxes have been assigned to the j-th ground truth box. For these k j prediction boxes, they have an assignment relationship with the j-th ground truth box. Specifically, in this step, each element p ij represents whether the i-th pre-screened prediction box is assigned to the j-th ground truth box. If it is, p ij takes 1. If not, pij Take 0. (That is the same meaning), and obtain the assignment matrix P, where each element p ij indicates whether the i-th predicted bounding box after pre-screening is assigned to the j-th ground truth bounding box. If it is, p ij takes 1. If not, p ij takes 0.
[0042] Step 4-5 includes:
[0043] Step 4-5-1: According to the assignment matrix P, screen which predicted bounding boxes are assigned to more than two ground truth bounding boxes at the same time, and remove the redundant assignment relationships: Assume that the i-th predicted bounding box is assigned to more than two ground truth bounding boxes. Query the cost matrix, and only retain the assignment relationship of the ground truth bounding box with the minimum cost, and modify the corresponding values in the assignment matrix P;
[0044] Step 4-5-2: Traverse all ground truth bounding boxes, and according to the assignment matrix P, screen whether the number of positive samples assigned to all ground truth bounding boxes is less than the number of positive samples k j . If so, execute Step 4-5-5. Otherwise, execute Step 4-5-3;
[0045] Step 4-5-3: Traverse the predicted bounding boxes missing from the assignment of all ground truth bounding boxes. For the j-th ground truth bounding box, assume that the j-th ground truth bounding box has k j ′ assigned predicted bounding boxes, and there should be k j assigned predicted bounding boxes, and k j >k j ′ . Query the cost matrix, and assign the k j -k j ′ predicted bounding boxes with the minimum unassigned cost to the j-th ground truth bounding box, and modify the corresponding values in the assignment matrix P;
[0046] Step 4-5-4: Return to Step 4-5-1;
[0047] Step 4-5-5: Output the assignment relationship between the predicted bounding boxes and the ground truth bounding boxes;
[0048] Step 4-6 includes:
[0049] Step 4-6-1: Combine all the ground truth bounding boxes and the corresponding predicted bounding boxes into more than two groups of training data. Assume that the j-th ground truth bounding box and the corresponding k j positive samples and n j negative samples form a group of training data;
[0050] Step 4-6-2: Define the loss function Loss of the j-th ground truth bounding box as: j Define as:
[0051]
[0052] Among them, L cls is the classification loss; L reg is the regression loss; L obj is the confidence loss; λ is the balance coefficient of the regression loss;
[0053] Step 4-6-3: The total loss function Loss total is defined as:
[0054]
[0055] Among them, N gd is the number of ground truth boxes in the training data;
[0056] Step 5 includes:
[0057] Step 5-1, heating stage: Train the teacher model and the student model once using the same training parameters and loss function (the training parameters refer to the learning rate, and the loss function refers to the values of each proportional parameter in the loss function); the teacher model adopts the YOLOV5 model, and the student model adopts the lightweight YOLOV5 network;
[0058] Step 5-2, freeze the teacher model, and set the learning rate of the student model to one-tenth of that in Step 5-1 (this is an empirical value, generally required to be much lower than the previous learning rate. The reason is that the results output by the teacher model are not absolutely reliable and cannot be fully trusted. Therefore, the learning rate of the student model should be greatly reduced to prevent being overly interfered by the results of the teacher model and deviating from the correct results);
[0059] Step 5-3, distillation stage: For the same image, use the teacher model to predict to obtain the predicted boxes, and use the ground truth boxes and the teacher's predicted boxes to calculate the loss function by weighted calculation to fine-tune the student model (train all the training data once); in the distillation stage, use the ground truth boxes and the teacher's predicted boxes to calculate the loss function L dis :
[0060] L dis = αL soft + L hard (6)
[0061] Among them, L dis is the distillation loss, L soft is the soft target loss, that is, using the predicted boxes generated by the teacher network as the ground truth boxes to calculate the loss with the predicted boxes generated by the student network; L hardis the hard target loss, that is, the loss is calculated using the ground truth boxes and the predicted boxes generated by the student network; α is the proportionality coefficient;
[0062] For L soft and L hard , the improved simOTA loss function in Equation (3) is adopted. The difference is that in L soft , the following formula is used to calculate the sample probability generated by the teacher model:
[0063]
[0064] where p si is the probability of a certain type of sample output by the teacher model; z i refers to the weight of the i-th type of sample output by the teacher model; T is the distillation temperature; exp is the natural exponential function; Step 5-4, repeat Step 5-1 to Step 5-3 until the preset number of iterations is reached;
[0065] Step 6 includes:
[0066] Use the lightweight improved model after heating and distillation cycle training to detect airport surface targets, and output the coordinates and categories of the targets.
[0067] The present invention has the following beneficial effects:
[0068] 1. Aiming at the problem that the current complex and large target detection algorithms are difficult to be deployed to the edge side and cannot be quickly and real-time monitored. Based on the YOLOV5 network, the present invention designs a C3 module based on the repeated cross bottleneck layer (RCBottleNeck) in the backbone network. Through two-layer cross one-dimensional convolution, while establishing the correlation between feature weights and spatial positions, the number of network parameters and time complexity are reduced; a multi-level perception convolution (MPConv) is constructed in the neck network, and the multi-layer combination of pixel features, local features, and global features is used to replace the global features, improving the operation speed of the model.
[0069] 2. Aiming at the problem of the inevitable decline in monitoring accuracy caused by lightweight networks, the present invention adopts the simOTA loss function to alleviate the harmful gradients caused by the inappropriate allocation of predicted boxes; at the same time, the present invention proposes a heating-distillation cycle training process, through the cycle of normal training and knowledge distillation fine-tuning process, making the lightweight model perform approximately the same as the normal (non-lightweight) model.
[0070] In short, the model can ensure a certain accuracy while being lightweight enough to be deployed to marginalized devices; it can provide a method for quickly and accurately monitoring airport surface targets in real scenarios, providing auxiliary support for the safe and stable operation of airports, and having certain practical value. Description of the Drawings
[0071] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0072] Figure 1 This is a diagram of the overall architecture of the method of the present invention.
[0073] Figure 2 Schematic diagram of the calculation process of the Focus module used in YOLOv5.
[0074] Figure 3 It is a C3 structure based on RCBottleNeck.
[0075] Figure 4 This is the MPConv network structure diagram.
[0076] Figure 5 Schematic diagram of the heating-distillation cycle training process.
[0077] Figure 6 A sample diagram of the airport surface monitoring dataset.
[0078] Figure 7 This is an example of airport surface target detection results. DETAILED DESCRIPTION
[0079] The embodiment of the present invention is aimed at the task of airport scene monitoring. In order to achieve both accuracy and rapidity in monitoring, a lightweight detection method for airport scene targets is designed. On the basis of the YOLOV5 network, a repeated cross-bottleneck layer (RCBottleNeck) is designed in the C3 module of the backbone network, and a multi-level perceptual convolution (MPConv) is constructed in the neck network to achieve lightweight and ensure rapid monitoring. At the same time, the present invention optimizes the loss function and training method, uses the simOTA loss function to alleviate harmful gradients, and proposes a heating-distillation cycle training process to address the problem of decreased accuracy due to lightweighting. The present invention can provide a method for quickly and accurately monitoring airport scene targets in real scenarios, provide auxiliary support for the safe and stable operation of airports, and has certain practical value. The overall model structure diagram of the present invention is shown in the attached figure. Figure 1 As shown. The Focus module calculation is shown in Figure 2 The original image (resolution W×H×3) is input into the Focus structure and sliced to become After concatenation, the feature maps are subjected to a 3×3 convolution operation and finally become feature map.
[0080] The lightweight detection algorithm of an airport surface target of the present invention is mainly implemented by the following steps.
[0081] Step 1: Design a repeated cross bottleneck layer (RCBottleNeck), as Figure 3 shown
[0082] Use a 3×1 convolution and a 1×3 convolution to replace the 3×3 convolution in the Bottleneck module of the standard YOLOV5 network. While directly reducing the number of parameters and time complexity, the combination of the two can still enable each pixel feature value to obtain the surrounding pixel features. The improved model has the characteristics of light weight. At the same time, this structure naturally constructs the relationship between feature weight and spatial position, that is, the central feature weight > the row and column feature weight > the corner feature weight, further enhancing the attention of the neural network to the central pixels.
[0083] Step 2: Design a multi-level perception convolution layer (MPConv), as Figure 4 shown:
[0084] The YOLOV5 network structure includes four parts: an input end, a backbone network, a neck network, and a detection head. The neck network is mainly composed of a standard convolution layer and an upsampling layer. For the standard convolution layer, its input channel size is the feature of CH1, and it maps it to the feature value with the channel size of CH2. Its time complexity is:
[0085] GFLOPs SC ~O(WHK1K2CH1CH2) (1)
[0086] where W and H are the sizes of the features, and K1 and K2 are the sizes of the convolution kernels.
[0087] For the time complexity of the standard convolution, it can be seen that the larger the number of output channels, the higher the time complexity. To effectively reduce the time complexity, the depthwise separable convolution (DWConv) performs depth convolution operations on each input channel separately, and at the same time uses a 1×1 convolution kernel to perform pointwise convolution operations at each position, and linearly combines the outputs of the depth convolution, but it completely cuts off the hidden connections between each channel, resulting in a decline in feature extraction ability and network accuracy. The grouped convolution (GConv) is between the standard convolution and the depthwise separable convolution. By performing block convolution on the features, it obtains local features of each group, and has certain feature extraction ability and time performance.
[0088] The MPConv designed in the present invention comprehensively considers the characteristics of the standard convolution, grouped convolution, and depthwise separable convolution, thereby constructing a multi-level perception convolution structure that simultaneously extracts pixel features, local features, and global features, and effectively reduces the time complexity at the same time. For the feature with the input channel size of CH1, first use the standard convolution to map it to the channel as The global features are then used with depthwise separable convolutions to obtain pixel features with a channel of At the same time, grouped convolutions with a grouping of g are performed to obtain local features with a channel of After the three are concatenated and shuffled, eigenvalue with a channel of CH2 is obtained, and its time complexity is:
[0089]
[0090] By replacing the standard convolution with the MPConv proposed in the present invention, the time complexity can be reduced while retaining better feature extraction capabilities.
[0091] In step 3, the corresponding parts in the YOLOV5 network structure are replaced with RCBottleNeck and MPConv to form a lightweight YOLOV5 network;
[0092] The present invention is based on the YOLOV5 network structure. The YOLOV5 network structure includes four parts: an input end, a backbone network, a neck network, and a detection head. The backbone network is used to extract image features, the detection head is used to output prediction boxes, that is, the results of object detection, and the neck network is a transitional part connecting the backbone network and the detection head, which is used to better utilize the features extracted by the backbone network.
[0093] Among them, there are a large number of C3 modules in the backbone network, and its function is to learn residual features. The C3 module is divided into two branches. One branch is stacked with BottleNeck after using standard convolution operations, and the other branch is only obtained by using standard convolution operations. The finally output features are the result of concatenating the two branches. Due to the existence of a large number of BottleNeck stacks, the parameter quantity of this structure is large, resulting in a slow detection speed. At the same time, these operations do not improve the accuracy of the model on small models. Replacing the BottleNeck in the C3 module of the backbone network in the YOLOV5 network structure with RCBottleNeck directly reduces the parameter quantity and time complexity, and has the characteristics of lightweight.
[0094] The neck network is mainly composed of standard convolutional layers and upsampling layers. By replacing the standard convolutions in the neck network with the MPConv proposed in the present invention, the time complexity can be reduced while retaining better feature extraction capabilities. However, if all stages of the model are replaced with MPConv, the network layer of the model will be deeper, and the deep layer will exacerbate the resistance to the data stream, significantly increasing the inference time. For the features at the Neck network, the channel dimension reaches the maximum, while the width and height dimensions reach the minimum. Therefore, the present invention only uses MPConv in the neck network.
[0095] Step 4: Design an improved simOTA loss function for the lightweight YOLOV5 network;
[0096] During the training process of YOLOV5, to calculate the loss function, it is necessary to match the predicted bounding boxes with the ground truth bounding boxes one by one. The label assignment method adopted by YOLOV5 is static assignment, but the static assignment strategy does not consider the situations of the ground truth bounding boxes having different sizes, shapes, or occlusions, resulting in a decline in the model's performance. To alleviate this situation, the present invention adopts an improved simOTA loss function, which specifically includes the following steps:
[0097] Step 4-1: Pre-screen the predicted bounding boxes;
[0098] Draw a square box of 160 pixels × 160 pixels centered on the ground truth bounding box (i.e., the size on the feature map is 5×5 because the scaling ratio of the YOLOV5 feature map is 32), calculate the intersection area between this square box and the ground truth bounding box, and only the predicted bounding boxes that have an intersection with this area are the pre-screened predicted bounding boxes.
[0099] Step 4-2: Calculate the classification loss and regression loss for the pre-screened predicted bounding boxes and all ground truth bounding boxes to obtain the cost matrix C and the IoU matrix M;
[0100] The cost matrix C described in Step 4-2 specifically includes:
[0101] simOTA regards the policy matching problem as an optimal transport problem. That is, assign the predicted bounding boxes to the ground truth bounding boxes at the lowest cost. This cost function includes regression loss, classification loss, etc. The cost c of the i-th pre-screened predicted bounding box and the j-th ground truth bounding box ij is defined as:
[0102]
[0103] where is the classification loss generated by the predicted bounding box i and the ground truth bounding box j, is the regression loss generated by the predicted bounding box and the ground truth bounding box, and σ is the proportional parameter.
[0104] The specific implementation of the classification loss and regression loss is the same as that of the YOLOV5 standard model.
[0105] Traverse all the pre-screened predicted bounding boxes and ground truth bounding boxes to calculate all c ij to obtain the cost matrix C.
[0106] The IoU matrix M described in Step 4-2 specifically includes:
[0107] Traverse all the pre-screened predicted bounding boxes and all ground truth bounding boxes to calculate the intersection over union (IoU) to obtain the IoU matrix M, and each element m ijDenote the intersection over union (IoU) between the $i$-th predicted bounding box and the $j$-th ground truth bounding box.
[0108] The intersection over union $m$ ij is calculated as follows: Find the intersection region between the $i$-th predicted bounding box and the $j$-th ground truth bounding box, which is the overlapping part of the two boxes; Calculate the intersection area: Then, calculate the area of the intersection region; Find the total area covered by the $i$-th predicted bounding box and the $j$-th ground truth bounding box, which includes the parts unique to each box plus their common intersection part; Divide the area of the intersection by the area of the union to obtain the intersection over union.
[0109] Step 4-3: Traverse each ground truth bounding box and obtain the number $k$ of corresponding positive predicted bounding boxes from the cost matrix j ;
[0110] For the $j$-th ground truth bounding box, take the top ten predicted bounding boxes with the largest intersection over union with this ground truth bounding box, accumulate their intersection over unions and round up (i.e., the minimum is 1) to get the result, which is the number $k$ of corresponding positive predicted bounding boxes for the $j$-th ground truth bounding box. j 。
[0111] Step 4-4: Reassign predicted bounding boxes for each ground truth bounding box to form the assignment matrix $P$;
[0112] Traverse all ground truth bounding boxes. For the $j$-th ground truth bounding box, query the cost matrix and assign its top $k$ j predicted bounding boxes with the smallest cost to it. Traverse all pre-screened predicted bounding boxes and all ground truth bounding boxes to check the assignment relationships to obtain the assignment matrix $P$, where each element $p$ ij indicates whether the $i$-th predicted bounding box is assigned to the $j$-th ground truth bounding box. If it is, take 1; if not, take 0.
[0113] Step 4-5: Further screen out duplicate predicted bounding boxes corresponding to different ground truth bounding boxes to ensure that one predicted bounding box corresponds to only one ground truth bounding box;
[0114] The further screening of duplicate predicted bounding boxes corresponding to different ground truth bounding boxes in Step 4-5 to ensure that one predicted bounding box corresponds to only one ground truth bounding box specifically includes:
[0115] Step 4-5-1: According to the assignment matrix $P$, screen out which predicted bounding boxes are assigned to multiple ground truth bounding boxes at the same time and remove the redundant assignment relationships. Assume that the $i$-th predicted bounding box is assigned to multiple ground truth bounding boxes. Query the cost matrix and only retain the assignment relationship with the ground truth bounding box with the smallest cost, and modify the corresponding values in the assignment matrix $P$;
[0116] Step 4-5-2: Traverse all ground truth bounding boxes and screen out which ground truth bounding boxes have the number of assigned positive samples less than the number $k$ of positive samples according to the assignment matrix $P$ j . If the number of assigned positive samples for all ground truth bounding boxes is equal to the number $k$ of positive samples j, then execute Step 4-5-5, otherwise execute Step 4-5-3;
[0117] Step 4-5-3: Traverse all the ground truth boxes to assign the missing prediction boxes. For the j-th ground truth box, assume it currently has k j ′ assigned prediction boxes, and it should have k j assigned prediction boxes, and k j > k j ′ . Query the cost matrix, and assign the k j - k j ′ prediction boxes with the smallest unassigned cost to the j-th ground truth box, and modify the corresponding values in the assignment matrix P;
[0118] Step 4-5-4: Return to Step 4-5-1;
[0119] Step 4-5-5: Output the assignment relationship between the prediction boxes and the ground truth boxes.
[0120] Step 4-6: Obtain the positive samples and their corresponding ground truth boxes from the above process, and classify the remaining prediction boxes as negative samples, thereby obtaining the loss function.
[0121] Step 4-6-1: Combine all the ground truth boxes and their corresponding prediction boxes into multiple groups of training data. Assume that the j-th ground truth box and its corresponding k j positive samples and n j negative samples form a group of training data;
[0122] Step 4-6-2: Define the loss function of the j-th ground truth box as:
[0123]
[0124] where L cls is the classification loss; L reg is the regression loss; L obj is the confidence loss; λ is the balance coefficient of the regression loss.
[0125] The specific implementation of the classification loss, regression loss, and confidence loss is the same as the standard implementation of YOLOV5.
[0126] Step 4-6-3: Define the total loss function as:
[0127]
[0128] where N gd is the number of ground truth boxes in the training data.
[0129] Step 5 For the lightweight YOLOV5 network, asFigure 5 As shown in Figure 5 , a heating-distillation cycle training process is designed to ensure that the performance of the lightweight model is similar to that of the normal (non-lightweight) model;
[0130] After the model lightweight operation, the overall number of parameters of the model is greatly reduced, which will lead to a decrease in the model's expression ability and cause performance degradation problems. To improve the model accuracy and make the performance of the lightweight model close to that of the original model, the present invention designs a heating-distillation cycle training process. By simultaneously training a high-precision original YOLOV5 model (teacher model) and a fast lightweight model (student model) on the airport scene monitoring dataset (data pictures are shown in Figure 6 ), and then using the teacher model and the real results to jointly train the student model, the specific algorithm is as follows:
[0131] Step 5-1: Train the teacher model and the student model for a certain number of rounds using the same training parameters and loss function, that is, the heating stage;
[0132] Step 5-2: Freeze the teacher model and set the learning rate of the student model to one-tenth of that in Step 5-1;
[0133] Step 5-3: For the same picture, use the teacher model to predict, obtain the prediction box, and use the real box and the teacher prediction box to calculate the loss function by weighted calculation, and fine-tune the student model for a certain number of rounds, that is, the distillation stage;
[0134] The weighted calculation of the loss function using the real box and the teacher prediction box described in Step 5-3 specifically includes:
[0135] In the distillation stage, the loss function is calculated by weighted calculation using the real box and the teacher prediction box, and the loss function is defined as follows:
[0136] L dis =αL soft +L hard (6)
[0137] Among them, L dis is the distillation loss, L soft is the soft target loss, that is, using the prediction box generated by the teacher network as the real box to calculate the loss with the prediction box generated by the student network; L hard is the hard target loss, that is, using the real box to calculate the loss with the prediction box generated by the student network. α is the proportionality coefficient.
[0138] For L soft and L hard , the improved simOTA loss function of Equation (3) is adopted, and the difference is that Lsoft Among them, the sample probability generated by the teacher model is calculated as follows:
[0139]
[0140] where p si is the probability of a certain type of sample output by the teacher model, and T is the distillation temperature. When the entropy of the sample probability distribution is relatively small, the values of negative samples are all very close to 0, and their contribution to the loss function is very small. Adding the distillation temperature T can increase the entropy of the probability distribution, amplify the negative sample information, and make the model pay more attention to negative samples.
[0141] Step 5-4: Repeat Steps 5-1 to 5-3 until the preset number of iterations is reached.
[0142] Step 6, use the trained lightweight YOLOV5 network to perform object detection on the airport scene to complete the lightweight detection of airport scene objects.
[0143] Use the lightweight improved model trained through the heating-distillation cycle to detect airport scene objects, output the coordinates and categories of the objects, and the output results are shown in Figure 7 .
[0144] The embodiment of the present invention provides a lightweight detection method for airport scene objects to improve the accuracy and speed of airport scene detection.
[0145] The method of this embodiment mainly includes the following steps:
[0146] Step 1, design a repeated cross bottleneck layer (RCBottleNeck);
[0147] Use a 3×1 convolution and a 1×3 convolution to replace the 3×3 convolution in the BottleNeck module of the standard YOLOV5 network. While directly reducing the number of parameters and time complexity, the combination of the two can still enable each pixel feature value to obtain the surrounding pixel features, and the improved model has the characteristics of lightweight. At the same time, this structure naturally constructs the relationship between feature weights and spatial positions, that is, the central feature weight > the row and column feature weight > the corner feature weight, further enhancing the attention of the neural network to the central pixels.
[0148] Step 2, design a multi-level perception convolutional layer (MPConv);
[0149] The YOLOV5 network structure includes four parts: an input end, a backbone network, a neck network, and a detection head. For a standard convolutional layer, its input channel size is the feature of CH1, and it maps it to the feature value with a channel size of CH2. Its time complexity is:
[0150] GFLOPs SC ~O(WHK1K2CH1CH2)#(1)
[0151] Among them, W and H are the sizes of the features, and K1 and K2 are the sizes of the convolutional kernels.
[0152] For the time complexity of standard convolution, it can be seen that the larger the number of output channels, the higher the time complexity. To effectively reduce the time complexity, depthwise separable convolution (DWConv) performs depthwise convolution operations on each channel of the input separately, and at the same time performs pointwise convolution operations at each position with a 1×1 convolutional kernel, and linearly combines the outputs of the depthwise convolution. However, it completely cuts off the hidden connections between each channel, resulting in a decrease in feature extraction ability and network accuracy. Group convolution (GConv) is between standard convolution and depthwise separable convolution. By performing block convolution on the features, local features of each group are obtained, and it has certain feature extraction ability and time performance.
[0153] The MPConv designed in the present invention comprehensively considers the characteristics of standard convolution, group convolution, and depthwise separable convolution, thereby constructing a multi-level perception convolution structure that simultaneously extracts pixel features, local features, and global features, and effectively reduces the time complexity at the same time. For the feature with an input channel size of CH1, first use standard convolution to map it to a global feature with a channel of and then use depthwise separable convolution to obtain a pixel feature with a channel of At the same time, perform group convolution with a group number of g to obtain a local feature with a channel of After concatenating and shuffling the three, a feature value with a channel of CH2 is obtained, and its time complexity is:
[0154]
[0155] By replacing the standard convolution with the MPConv proposed in the present invention, the time complexity can be reduced while retaining better feature extraction ability.
[0156] Step 3, replace the corresponding parts in the YOLOV5 network structure with RCBottleNeck and MPConv to form a lightweight YOLOV5 network
[0157] The present invention is based on the YOLOV5 network structure. The YOLOV5 network structure includes four parts: an input end, a backbone network, a neck network, and a detection head. The backbone network is used to extract image features, the detection head is used to output prediction boxes, that is, the results of object detection, and the neck network is a transitional part connecting the backbone network and the detection head, which is used to better utilize the features extracted by the backbone network.
[0158] Among them, there are a large number of C3 modules in the backbone network, whose function is to learn residual features. The C3 module is divided into two branches. One branch uses standard convolution operations and is stacked with BottleNeck, and the other branch only uses standard convolution operations to obtain. The finally output features are the result of concatenating the two branches. Due to the existence of a large number of BottleNeck stacks, the parameter quantity of this structure is large, resulting in a slow detection speed. At the same time, these operations will not improve the accuracy of the model on small models. Replacing the BottleNeck in the C3 module of the backbone network in the YOLOV5 network structure with RCBottleNeck directly reduces the parameter quantity and time complexity, and has the characteristics of lightweight.
[0159] The comparison of the parameter quantities between the lightweight YOLOV5 network proposed by the present invention and the standard YOLOV5 network is shown in Table 1.
[0160] Table 1
[0161]
[0162] Step 4, for the lightweight YOLOV5 network, design an improved simOTA loss function
[0163] During the training process of YOLOV5, in order to calculate the loss function, it is necessary to correspond the predicted bounding boxes with the ground truth bounding boxes one by one. The label assignment method adopted by YOLOV5 is static assignment, but the static assignment strategy does not consider the situations of the size, shape or occlusion of the ground truth bounding boxes, resulting in a decline in the model effect. To alleviate the occurrence of this situation, the present invention adopts an improved simOTA loss function, which specifically includes the following steps:
[0164] Step 4-1: Pre-screen the predicted bounding boxes;
[0165] Draw a square box of 160 pixels × 160 pixels with the center of the ground truth bounding box (i.e., the size on the feature map is 5×5, because the scaling ratio of the yolov5 feature map is 32), calculate the intersection area between this square box and the ground truth bounding box, and only the predicted bounding boxes that have an intersection with this area are the pre-screened predicted bounding boxes.
[0166] Step 4-2: Calculate the classification loss and regression loss for the pre-screened predicted bounding boxes and all ground truth bounding boxes to obtain a cost matrix C and an iou matrix M;
[0167] The cost matrix C described in Step 4-2 specifically includes:
[0168] simOTA views the strategy matching problem as an optimal transmission problem. That is, the predicted bounding boxes are assigned to the ground truth bounding boxes at the lowest cost. The cost function includes regression loss, classification loss, etc. The cost c between the i-th pre-screened predicted bounding box and the j-th ground truth bounding box is ij defined as follows:
[0169]
[0170] where is the classification loss between the predicted bounding box i and the ground truth bounding box j, is the regression loss between the predicted bounding box and the ground truth bounding box, and λ is a proportionality parameter, which is taken as 3.0 here.
[0171] The specific implementation of the classification loss and the regression loss is the same as that of the YOLOV5 standard network.
[0172] Traverse all pre-screened predicted bounding boxes and ground truth bounding boxes to calculate all c ij , and obtain the cost matrix C.
[0173] The iou matrix M described in step 4-2 specifically includes:
[0174] Traverse all pre-screened predicted bounding boxes and all ground truth bounding boxes to calculate the intersection over union (iou), and obtain the iou matrix M. Each element m ij represents the intersection over union between the i-th predicted bounding box and the j-th ground truth bounding box.
[0175] The calculation method of the intersection over union m ij is as follows: find the intersection area between the i-th predicted bounding box and the j-th ground truth bounding box, which is the overlapping part of the two boxes; calculate the intersection area: then, calculate the area of the intersection region; find the total area covered by the i-th predicted bounding box and the j-th ground truth bounding box, which includes the unique parts of the two boxes plus their common intersection part; divide the area of the intersection by the area of the union to obtain the intersection over union.
[0176] Step 4-3: Traverse each ground truth bounding box, and obtain the number k of corresponding predicted bounding box positive samples from the cost matrix j ;
[0177] For the j-th ground truth bounding box, take the top ten predicted bounding boxes with the largest intersection over union with this ground truth bounding box, and the result obtained by accumulating and rounding up (i.e., the minimum is 1) their intersection over union is the number k of predicted bounding box positive samples corresponding to the j-th ground truth bounding box j .
[0178] Step 4-4: Reassign predicted bounding boxes for each ground truth bounding box to form the assignment matrix P;
[0179] Traverse all ground truth bounding boxes. For the j-th ground truth bounding box, query the cost matrix and take the top k with the lowest costj One prediction box is assigned to it. Traverse all pre-screened prediction boxes and all ground truth boxes to check the assignment relationship, and obtain the assignment matrix P. Each element p ij indicates whether the i-th prediction box is assigned to the j-th ground truth box. If so, take 1; if not, take 0.
[0180] Step 4-5: Further filter the prediction boxes corresponding to different ground truth boxes that are duplicates, ensuring that one prediction box corresponds to only one ground truth box;
[0181] The further filtering of the prediction boxes corresponding to different ground truth boxes as described in Step 4-5 to ensure that one prediction box corresponds to only one ground truth box specifically includes:
[0182] Step 4-5-1: According to the assignment matrix P, filter out which prediction boxes are assigned to multiple ground truth boxes at the same time, and remove the redundant assignment relationships. Assume that the i-th prediction box is assigned to multiple ground truth boxes. Query the cost matrix, and only retain the assignment relationship of the ground truth box with the minimum cost, and modify the corresponding values in the assignment matrix P;
[0183] Step 4-5-2: Traverse all ground truth boxes, and filter out which ground truth boxes have the number of positive samples assigned less than the number of positive samples k according to the assignment matrix P j , if the number of positive samples assigned to all ground truth boxes is equal to the number of positive samples k j , then execute Step 4-5-5; otherwise, execute Step 4-5-3;
[0184] Step 4-5-3: Traverse the prediction boxes missing from the assignment of all ground truth boxes. For the j-th ground truth box, assume it currently has k j ′ assigned prediction boxes, and there should be k j assigned prediction boxes, and k j >k j ′ , query the cost matrix, and assign the k j -k j ′ prediction boxes with the minimum unassigned cost to the j-th ground truth box, and modify the corresponding values in the assignment matrix P;
[0185] Step 4-5-4: Return to Step 4-5-1;
[0186] Step 4-5-5: Output the assignment relationship between the prediction boxes and the ground truth boxes.
[0187] Step 4-6: Obtain the positive samples and the corresponding ground truth boxes from the above process, and classify the remaining prediction boxes as negative samples, thereby obtaining the loss function.
[0188] Step 4-6-1: Combine all the real boxes and their corresponding predicted boxes into multiple sets of training data. Let the jth real box and its corresponding k j positive samples and n j Negative samples constitute a set of training data;
[0189] Step 4-6-2: Define the loss function of the jth true box as:
[0190]
[0191] Among them, L cls is the classification loss; L reg is the regression loss; L obj is the confidence loss; λ is the balance coefficient of regression loss, which is 5.0 here.
[0192] The specific implementation of classification loss, regression loss, and confidence loss is the same as the standard implementation of YOLOV5.
[0193] Step 4-6-3: The total loss function is defined as:
[0194]
[0195] Among them, N gd is the number of ground-truth boxes in the training data.
[0196] Step 5: Design a heating-distillation cycle training process for the lightweight YOLOV5 network to ensure that the performance of the lightweight model is similar to that of the normal (non-lightweight) model;
[0197] This embodiment uses a real airport scene monitoring data to construct a data set and conduct simulation experiments. The video has 54,000 frames, each of which is manually annotated, that is, there are 54,000 annotated airport images. In this data set, 43,200 images are randomly selected as the training set, 5,400 as the validation set, and the remaining 5,400 as the test set. The image size is 1920×1080. There are four categories: plane, ladder, truck, and bus.
[0198] After the model is lightweighted, the overall number of model parameters is greatly reduced, which will lead to a decrease in the model's expressiveness and performance degradation. In order to improve the model accuracy and make the performance of the lightweight model close to the original model, the present invention designs a heating-distillation cycle training process, which simultaneously trains the high-precision original YOLOV5 model (teacher model) and the fast lightweight model (student model), and then uses the teacher model and the real results to jointly train the student model. The specific algorithm is as follows:
[0199] Step 5-1: Train the teacher model and the student model for a certain number of rounds using the same training parameters and loss function, i.e., the heating stage;
[0200] In this experiment, the settings of each hyperparameter are as follows: image_size is 640, batch_size is 16, the learning rate is 0.001. In this paper, the K-means clustering algorithm is used to cluster the target boxes of this dataset to obtain appropriate anchor sizes.
[0201] The CPU model used in the experiment is 64Intel(R)Xeon(R)Gold 6226R CPU@2.90GHz, the running memory is 251GB, the GPU is: GeForce RTX 3090, the video memory is 24GB, the Pytorch deep learning framework, and the CUDA version is 11.1.
[0202] Step 5-2: Freeze the teacher model and set the learning rate of the student model to one-tenth of that in Step 5-1;
[0203] Step 5-3: For the same image, use the teacher model to predict, obtain the prediction boxes, use the real boxes and the teacher's prediction boxes to calculate the loss function by weighting, and fine-tune the student model for a certain number of rounds, i.e., the distillation stage;
[0204] The calculation of the loss function by weighting the real boxes and the teacher's prediction boxes described in Step 5-3 specifically includes:
[0205] In the distillation stage, calculate the loss function by weighting the real boxes and the teacher's prediction boxes. The loss function is defined as follows:
[0206] L dis =αL soft +L hard #(6)
[0207] Among them, L dis is the distillation loss, L soft is the soft target loss, that is, using the prediction boxes generated by the teacher network as the real boxes to calculate the loss with the prediction boxes generated by the student network; L hard is the hard target loss, that is, using the real boxes to calculate the loss with the prediction boxes generated by the student network. α is the proportionality coefficient, taking 0.1, making the student network more inclined to learn real data.
[0208] For L soft and L hard , the improved simOTA loss function of Equation (3) is adopted. The difference is that L softAmong them, the sample probability generated by the teacher model is calculated as shown in the following formula:
[0209]
[0210] where p si is the probability of a certain type of sample output by the teacher model, and T is the distillation temperature. When the entropy of the sample probability distribution is relatively small, the values of negative samples are all very close to 0, and their contribution to the loss function is very small. Adding the distillation temperature T can increase the entropy of the probability distribution, amplify the negative sample information, and make the model pay more attention to negative samples.
[0211] Step 5-4: Repeat Steps 5-1 to 5-3 until the preset number of iterations is reached.
[0212] Step 6: Use the trained lightweight YOLOV5 network to perform target detection on the airport scene to complete the lightweight detection of airport scene targets.
[0213] Use the lightweight improved model after heating-distillation cycle training to detect airport scene targets and output the coordinates and categories of the targets.
[0214] Five network models were selected for comparative experiments. YOLOv5 is the basic model, Fast-RCNN is a two-stage model, NanoDet is a lightweight one-stage model, and YOLOv5-CBAM and MNtECA are lightweight improved models for airport target detection based on the YOLOv5 architecture. The index results obtained from the trained network models for the validation set images of the airport scene monitoring dataset are compared as shown in Table 2 below.
[0215] Table 2
[0216] algorithm precision recall MAP50 YOLOv5 0.995 0.985 0.991 FAST-RCNN 0.989 0.991 0.975 NanoDet 0.996 0.994 0.992 YOLOv5-CBAM 0.995 0.992 0.997 MNtECA 0.983 0.090 0.983 the method of the present invention 0.998 0.999 0.995
[0217] It can be found that the method of this embodiment is superior to other methods in both recall rate and precision rate, is similar to the optimal method YOLOv5-CBAM in the MAP50 index, and is superior to other comparative methods. In summary, the method of the present invention has the characteristics of lightweight and high precision, is suitable for real-time target detection in airport monitoring, and helps managers coordinate the scheduling of the airport flight area.
[0218] The present invention provides a lightweight detection method for airport scene targets. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A lightweight detection method for airport surface targets, characterized in that, It includes the following steps: Step 1: Design the Repeated Cross Bottleneck layer (RCBottleNeck); Step 2: Design the Multi-Level Perception Convolution layer (MPConv); Step 3: Replace the corresponding parts in the YOLOV5 network structure with the Repeated Cross Bottleneck layer (RCBottleNeck) and the Multi-Level Perception Convolution layer (MPConv) to form a lightweight YOLOV5 network; Step 4: For the lightweight YOLOV5 network, design an improved simplified optimal transport assignment (simOTA) loss function; Step 5: For the lightweight YOLOV5 network, design a heating and distillation cyclic training process to obtain a trained lightweight YOLOV5 network; Step 6: Use the trained lightweight YOLOV5 network to perform object detection on the airport scene to complete the lightweight detection of airport scene targets.
2. The method according to claim 1, characterized in that Step 1 includes: The Repeated Cross Bottleneck layer (RCBottleNeck) replaces all 3×3 convolutions in the Bottleneck module of the C3 structure in the BackBone network with a 3×1 convolution followed by a 1×3 convolution.
3. The method according to claim 2, characterized in that, Step 2 includes: For the feature with an input channel size of CH1, the multi-level perception convolutional layer MPConv first maps the feature to a global feature with a channel size of using a standard convolution, then uses a depthwise separable convolution to obtain a pixel feature with a channel size of . At the same time, grouped convolution with a group size of g is performed to obtain a local feature with a channel of . After concatenating the global feature, pixel feature, and local feature and performing a shuffle, a feature value with a channel of CH2 is obtained, and the time complexity GFLOPs GSConv ~O is: Where W and H are the feature width and height respectively, and K1 and K2 are the convolution kernel sizes.
4. The method according to claim 3, wherein Step 3 includes: The YOLOV5 network structure includes an input end, a backbone network (Backbone), a neck network (Neck), and a detection head (Head); The backbone network (Backbone) is used to extract image features, the detection head (Head) is used to output prediction boxes, that is, the results of object detection, and the neck network (Neck) is a transitional part connecting the backbone network (Backbone) and the detection head (Head) for better utilization of the features extracted by the backbone network; The neck network (Neck) includes a standard convolution layer and an upsampling layer; Replace the standard convolution layer with the Multi-Level Perception Convolution layer (MPConv); The backbone network (Backbone) includes a C3 module, and the C3 module is used to learn residual features. Replace the Bottleneck in the C3 module of the backbone network with the Repeated Cross Bottleneck layer (RCBottleNeck).
5. The method according to claim 4, characterized in that Step 4 includes: Step 4-1: Pre-screen the prediction boxes; Step 4-2: Calculate the classification loss and regression loss between the pre-screened prediction boxes and all ground truth boxes to obtain a cost matrix C and an intersection over union (IoU) matrix M; Step 4-3: Traverse each ground truth box and obtain the number k of corresponding predicted box positive samples from the cost matrix C j ; Step 4-4: Reassign prediction boxes for each ground truth box to form an assignment matrix P; Step 4-5: Further screen the repeated prediction boxes corresponding to different ground truth boxes to ensure that one prediction box only corresponds to one ground truth box; Step 4-6: Obtain positive samples and corresponding ground truth boxes through Steps 4-1 to 4-5, and classify the remaining prediction boxes as negative samples to obtain the loss function.
6. The method according to claim 5, wherein Step 4-1 includes: Draw a 160 pixel × 160 pixel box centered on the ground truth box, calculate the intersection area between the box and the ground truth box, and only the prediction boxes that intersect with the intersection area are the pre-screened prediction boxes.
7. The method according to claim 6, characterized in that, Step 4-2 includes: The cost c between the i-th pre-screened prediction box and the j-th ground truth box ij is defined as: wherein, is the classification loss generated by the i-th pre-screened prediction box and the j-th ground truth box, is the regression loss generated by the prediction box and the ground truth box, and σ is the proportionality parameter; The cost c calculated by traversing all pre-screened predicted bounding boxes and ground truth bounding boxes ij , to obtain the cost matrix C; Traverse the intersection over union (IoU) calculated between all pre-screened prediction boxes and all ground truth boxes to obtain the IoU matrix M. The element m at the i-th row and j-th column in the IoU matrix M ij represents the IoU between the i-th pre-screened prediction box and the j-th ground truth box.
8. The method according to claim 7, wherein Step 4-3 includes: For the j-th ground truth box, select the top X1 prediction boxes with the largest intersection over union (IoU) with the j-th ground truth box. Calculate the IoU between the X1 prediction boxes and the j-th ground truth box, accumulate the results and round up. The resulting value is the number k of positive prediction boxes corresponding to the j-th ground truth box j 。 9. The method according to claim 8, wherein Step 4-4 includes: Traverse all the ground truth boxes. For the j-th ground truth box, query the cost matrix and assign the top k prediction boxes with the minimum cost to the j-th ground truth box; j Traverse all pre-screened prediction boxes and all ground truth boxes to check the assignment relationship, and obtain the assignment matrix P. Each element p ij indicates whether the i-th pre-screened prediction box is assigned to the j-th ground truth box. If so, p ij takes 1. If not, p ij takes 0.
10. The method according to claim 9, characterized in that Step 4-5 includes: Step 4-5-1: According to the assignment matrix P, screen out which prediction boxes are assigned to more than two ground truth boxes at the same time, and remove the redundant assignment relationships: Assume that the i-th prediction box is assigned to more than two ground truth boxes. Query the cost matrix, and only retain the assignment relationship of the ground truth box with the minimum cost, and modify the corresponding value in the assignment matrix P; Step 4-5-2: Traverse all the ground truth boxes, and filter according to the assignment matrix P to check whether the number of assigned positive samples for all ground truth boxes is less than the number of positive samples k j , if so, execute Step 4-5-5, otherwise execute Step 4-5-3; Step 4-5-3: Traverse all the ground truth boxes to assign the missing prediction boxes. For the j-th ground truth box, assume that there are currently k j ′ assigned prediction boxes. There should be k j assigned prediction boxes, and k j >k j ′ . Query the cost matrix, and assign the k j -k j ′ prediction boxes with the smallest unassigned cost to the j-th ground truth box, and modify the corresponding values in the assignment matrix P; Step 4-5-4: Return to Step 4-5-1; Step 4-5-5: Output the assignment relationship between the prediction boxes and the ground truth boxes; Step 4-6 includes: Step 4-6-1: Combine all the ground truth boxes and the corresponding predicted boxes into two or more sets of training data. Let the j-th ground truth box and the corresponding k j positive samples and n j negative samples form a set of training data; Step 4-6-2: Define the loss function Loss of the j-th ground truth box j as follows: Among them, L cls is the classification loss; L reg is the regression loss; L obj is the confidence loss; λ is the balance coefficient of the regression loss; Step 4-6-3: Total loss function Loss total is defined as: Among them, N gd is the number of ground truth boxes in the training data; Step 5 includes: Step 5-1, heating stage: Train the teacher model and the student model once using the same training parameters and loss function; The teacher model adopts the YOLOV5 model, and the student model adopts the lightweight YOLOV5 network; Step 5-2, freeze the teacher model and set the learning rate of the student model; Step 5-3, distillation stage: For the same picture, use the teacher model to make predictions, obtain the predicted bounding boxes, and use the ground truth bounding boxes and the teacher's predicted bounding boxes to calculate the loss function by weighted calculation, and fine-tune the student model; in the distillation stage, use the ground truth bounding boxes and the teacher's predicted bounding boxes to calculate the loss function L by weighted calculation dis : L dis = αL soft + L hard (6) Among them, L dis is the distillation loss, L soft is the soft target loss, that is, the prediction box generated by the teacher network is used as the true box and the prediction box generated by the student network is used to calculate the loss; L hard is the hard target loss, that is, the loss is calculated using the real box and the predicted box generated by the student network; α is the proportional coefficient; For L soft and L hard , the improved simOTA loss function of Equation (3) is adopted. The difference is that in L soft , the following formula is used to calculate the sample probability generated by the teacher model: Among them, p si is the probability of a certain type of sample output by the teacher model; z i refers to the weight of the i-th type of sample output by the teacher model; T is the distillation temperature; exp is the natural exponential function; Step 5-4, repeat Step 5-1 to Step 5-3 until the preset number of iterations is reached; Step 6 includes: Use the lightweight improved model after heating and distillation cycle training to detect the airport surface targets, and output the coordinates and categories of the targets.