Unmanned aerial vehicle road bridge identification method based on lightweight deep neural network model

Through the improvement of the lightweight deep neural network model, combined with data preprocessing, GSConv convolution, bidirectional feature pyramid network and dynamic detection head, the problem of poor detection of small and medium-sized targets of drone bridge recognition is solved, and more efficient and accurate bridge recognition effect is achieved.

CN120564075APending Publication Date: 2025-08-29JIANGSU WATER CONSERVANCY SCI RES INST +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510510509.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Drones do not have the effect of identifying small targets in complex contexts, especially in multi-scale environments and occlusions, and existing methods are difficult to achieve efficient and accurate target detection.

Method used

The lightweight deep neural network model is adopted to enhance image diversity through data preprocessing, combined with the GSConv convolution module, bidirectional feature pyramid network and dynamic detection head, a Wise-IoU loss function is introduced to optimize bounding box regression and improve detection performance.

Benefits of technology

It improves the accuracy and robustness of drone bridge identification, can better cope with the needs of diverse scenarios, and adapt to complex backgrounds and small object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564075A_ABST
    Figure CN120564075A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle road bridge identification method based on a lightweight deep neural network model. The method comprises the steps of obtaining a road bridge image data set shot by an unmanned aerial vehicle; performing image enhancement on the data set through a data preprocessing module; establishing a road and bridge identification model, extracting features of the data set after image enhancement by using a backbone network, and generating a feature map; performing fusion and multi-scale feature transmission on the feature map through a neck network connected with the backbone network to generate an enhanced multi-scale feature map; target classification and bounding box regression are carried out through the detection head, and a detection result is obtained; and performing road and bridge identification on the unmanned aerial vehicle image acquired in real time by using the trained road and bridge identification model. The method provided by the invention improves the detection efficiency and accuracy, especially provides a better solution for small target recognition and complex background processing, and can better cope with diversified scene requirements under the view angle of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drone image target recognition, and more specifically, to a drone road and bridge recognition method based on a lightweight deep neural network model. Background Art

[0002] In recent years, the rapid development of drone technology has greatly promoted its widespread application in military, civilian, and commercial fields. In particular, object detection has become a core research direction in drone image processing. Traditional object detection methods rely on manual annotation and visual analysis. While effective in initial applications, they often face limitations such as high labor costs, low efficiency, and subjective interference when processing large amounts of data.

[0003] In contrast, drones equipped with high-resolution sensors can efficiently collect large amounts of high-precision image data. By combining computer vision and machine learning algorithms, they achieve automated and efficient target detection and recognition. This fusion of technologies not only significantly improves the accuracy and real-time performance of drone target detection, but also provides broad development prospects and room for technological innovation in its application and expansion into more fields.

[0004] Currently, mainstream deep learning object detection methods are mainly divided into two categories: two-stage detection algorithms and one-stage detection algorithms. Two-stage detection methods first generate multiple candidate regions from the input image, then classify and locate these candidate regions, and finally output the detection results. This type of algorithm has high accuracy, but because it requires two feature extraction steps, the detection speed is slow and it performs poorly in real-time detection tasks. One-stage methods complete the entire object detection process with a single feature extraction. Therefore, they have higher real-time performance and better accuracy in object detection tasks and are widely used in various scenarios. However, their accuracy is slightly lower than that of two-stage methods. Representative one-stage methods include YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and RetinaNet.

[0005] Over the past few years, some researchers have conducted technical research and applied it in this field, achieving some progress. However, these methods are still insufficient for detecting small objects in multi-scale environments and complex backgrounds. In particular, the model's performance remains unsatisfactory when the object is occluded.

[0006] To sum up, in the target recognition task of drones, there are problems such as the large range of scenes contained in the image, the small size of the target objects and their multi-scale distribution, the complex and changeable background, and the target being easily occluded. Systematic innovation and breakthroughs are needed from multiple levels such as data processing, network structure design, and loss function optimization. Summary of the Invention

[0007] The purpose of this invention is to address the problems existing in the existing drones' recognition of road and bridge targets, and to propose a drone road and bridge recognition method and system based on a lightweight deep neural network model. By improving the algorithm's architecture and related modules, it aims to improve detection efficiency and accuracy, especially in small target recognition and complex background processing, and provide better solutions; this can better cope with the diverse scene requirements from the drone's perspective.

[0008] The technical solution of the present invention is:

[0009] In a first aspect, the present invention provides a method for identifying roads and bridges using an unmanned aerial vehicle (UAV) based on a lightweight deep neural network model, comprising:

[0010] S1. Obtain a dataset of road and bridge images taken by a drone;

[0011] S2. performing image enhancement on the data set through a data preprocessing module;

[0012] S3. Establish a road and bridge recognition model, use a backbone network to extract features of the image enhancement dataset and generate a feature map; use a neck network connected to the backbone network to fuse the feature map and perform multi-scale feature transfer to generate an enhanced multi-scale feature map; use a detection head connected to the neck network to perform target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; calculate the error of the detection result using a loss function, and the road and bridge recognition model is trained when the error meets the standard;

[0013] S4. Use the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired UAV image; and visualize the recognition results through a result display module.

[0014] Further in S2, image enhancement is performed on the dataset by a data preprocessing module to increase the diversity of the dataset, which specifically includes one or more of the following steps:

[0015] Randomly crop the images of the dataset to generate images of different sizes;

[0016] Flipping and rotating the images of the dataset to generate images in different orientations;

[0017] Adjusting the brightness of the images in the dataset to simulate different lighting conditions;

[0018] Adding Gaussian noise to the image of the data set to generate a noisy image;

[0019] Performing contrast transformation on the images of the dataset to generate images with different contrasts.

[0020] Further in S3, the backbone network is used to extract features of the image enhanced dataset to generate a feature map, specifically including:

[0021] The GSConv convolution module is used to perform multi-layer convolution operations on the input image to gradually identify the edge, texture, and shape features of the input image; a feature map is generated based on the extracted features, the spatial resolution of the feature map is gradually reduced, the number of channels is gradually increased, deep semantic information is extracted, and a multi-scale feature map is output to provide spatial and semantic features of objects of different sizes.

[0022] Furthermore, the GSConv convolution module includes a standard convolution module and a depthwise separable convolution module; the standard convolution module and the depthwise separable convolution module are used in turn to downsample the input image to reduce the spatial resolution of the input; the feature maps output by the two convolution modules are concat- orted , where the standard convolution module is used for dense feature extraction and the depthwise separable convolution module focuses on lightweight calculation; the concatenated feature maps are subjected to a channel rearrangement shuffle operation so that the feature information from the standard convolution is adjacent to the corresponding features from the depthwise separable convolution.

[0023] Further in S3, the neck network connected to the backbone network fuses the feature maps and transfers multi-scale features to generate an enhanced multi-scale feature map, including:

[0024] Using a bidirectional feature pyramid network structure as a feature fusion mechanism, the multi-scale feature maps from the backbone network are used as input, and high-level semantic information and low-level spatial information are fused through top-down and bottom-up bidirectional cross-scale connections;

[0025] A weighted feature fusion mechanism is introduced to dynamically adjust the weights of features at different levels during the fusion process, and output the fused multi-scale feature map to improve the detection capability of targets of different scales.

[0026] Further in S3, the detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result, including:

[0027] A dynamic detection head is used to improve detection performance by introducing different self-attention mechanisms in the three dimensions of the feature tensor.

[0028] In terms of scale perception, it dynamically integrates feature information at different levels in the feature pyramid. In terms of spatial perception, it enhances sensitivity to different spatial positions in the image through the spatial attention mechanism. In terms of task perception, it applies the attention mechanism on feature channels to adaptively activate the corresponding feature channels according to the requirements of different tasks.

[0029] The multi-scale feature maps from the neck network are fed into their respective detection branches, where local information is further extracted through the convolutional layer, and features related to the target category and bounding box are generated. The classification and regression modules are used to predict the category probability and corresponding bounding box coordinates of each target, respectively, to obtain the detection results.

[0030] Furthermore, in S3, the WIoU loss function is used as the loss function for bounding box regression, and the concept of outlier is introduced to evaluate the quality of the anchor box, including;

[0031] Define outlier to describe the quality of the anchor box and assign a gradient gain to it;

[0032] Dynamically adjust the gradient gain according to the degree of outlier of the anchor box;

[0033] Among them, the gradient gain assigned to the anchor box with smaller outliers is greater than the gradient gain assigned to the anchor box with larger outliers, so that the bounding box regression focuses on the anchor box of normal quality while preventing low-quality examples from generating large harmful gradients;

[0034] Through the dynamic anchor frame quality division standard, a gradient gain allocation strategy that suits the current situation is made at each moment.

[0035] Furthermore, S3 also includes evaluating the performance of the model after optimization by the loss function through a model evaluation module, and the evaluation indicators include mAP50, mAP50-95, precision Precision and recall Recall.

[0036] Further in S4, the visual display of the recognition result by the result display module includes:

[0037] The detection results are calibrated on the image through bounding boxes, category labels are added to the detected objects, and the corresponding confidence scores are annotated for each bounding box. The image output is generated, which includes the bounding box, object category label and confidence score.

[0038] By visualizing the image, the road and bridge targets detected by the model are intuitively presented;

[0039] Generate reports for subsequent analysis by batch processing multiple images.

[0040] In a second aspect, the present invention provides a UAV road and bridge recognition system based on a lightweight deep neural network model, the system comprising:

[0041] Data acquisition module, used to obtain road and bridge image datasets taken by drones;

[0042] A data preprocessing module, configured to perform image enhancement on the data set through the data preprocessing module;

[0043] The recognition model training module is used to establish a road and bridge recognition model, using a backbone network to extract features of the image enhancement dataset to generate a feature map; a neck network connected to the backbone network fuses the feature map and performs multi-scale feature transfer to generate an enhanced multi-scale feature map; a detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; and a loss function is used to calculate the error of the detection result. When the error meets the standard, the road and bridge recognition model is trained.

[0044] The measurement display module is used to use the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired drone images; the recognition results are visualized through the result display module.

[0045] Beneficial effects of the present invention:

[0046] The present invention discloses a method and system for unmanned aerial vehicle (UAV) road and bridge recognition based on a lightweight deep neural network model. The method intelligently recognizes and analyzes a dataset of road and bridge images taken by UAVs. Data preprocessing is used to enhance images and improve data diversity. A backbone network of the GSConv convolution module is used to extract features, a bidirectional feature pyramid network is used for feature fusion, and a dynamic detection head is used to achieve target classification and bounding box regression. A loss function is introduced to dynamically adjust the gradient gain according to the quality of the anchor frame and optimize the bounding box regression. Finally, the recognition results are visualized to intuitively present the road and bridge targets.

[0047] This invention improves the accuracy and robustness of road and bridge identification through a number of innovative technologies, provides an effective solution for UAV image analysis and infrastructure monitoring, and has important practical application value.

[0048] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present invention.

[0050] Figure 1 This is a flow chart of a UAV road and bridge recognition method based on a lightweight deep neural network model of the present invention.

[0051] Figure 2 Schematic diagram of the original image, the YOLOv8 model, and the recognition effect of the recognition method of the present invention after processing the original image. DETAILED DESCRIPTION

[0052] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0053] Example 1

[0054] Figure 1 A flowchart of a method for identifying roads and bridges using a drone based on a lightweight deep neural network model is shown according to an embodiment of the present invention.

[0055] like Figure 1 As shown, the road and bridge identification method includes:

[0056] S1. Obtain a dataset of road and bridge images taken by a drone;

[0057] The image dataset was obtained using images of roads and bridges taken by drones along the banks of rivers and lakes. A total of 420 images were selected as the dataset. The roads and bridges in the images were then annotated using rotatable rectangular frames. The annotated dataset was then divided into a training set and a test set in a 4:1 ratio, that is, 336 images in the training set and 84 images in the test set.

[0058] S2. Performing image enhancement on the dataset using a data preprocessing module; specifically comprising one or more of the following steps:

[0059] Randomly crop the images of the dataset to generate images of different sizes;

[0060] Flipping and rotating the images of the dataset to generate images in different orientations;

[0061] Adjusting the brightness of the images in the dataset to simulate different lighting conditions;

[0062] Adding Gaussian noise to the image of the data set to generate a noisy image;

[0063] Performing contrast transformation on the images of the dataset to generate images with different contrasts.

[0064] After these enhancements, the total number of images in the training set increased to 2,016, providing a richer sample set for model training. These enhancements aim to address data limitations by generating more diverse training samples and improving the model's generalization capabilities, thereby effectively reducing overfitting and ensuring the model's stability and accuracy in real-world applications.

[0065] S3. Establish a road and bridge recognition model, use a backbone network to extract features of the image enhancement dataset and generate a feature map; use a neck network connected to the backbone network to fuse the feature map and perform multi-scale feature transfer to generate an enhanced multi-scale feature map; use a detection head connected to the neck network to perform target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; calculate the error of the detection result using a loss function. When the error meets the standard, the road and bridge recognition model is trained, specifically including:

[0066] S3.1. Use the GSConv convolution module to perform multi-layer convolution operations on the input image to gradually identify the edge, texture, and shape features of the input image.

[0067] A feature map is generated based on the extracted features, the spatial resolution of the feature map is gradually reduced, the number of channels is gradually increased, deep semantic information is extracted, and a multi-scale feature map is output to provide spatial and semantic features of objects of different sizes.

[0068] Among them, the GSConv convolution module includes a standard convolution module and a depth-wise separable convolution module; the standard convolution module and the depth-wise separable convolution module are used in turn to downsample the input image to reduce the spatial resolution of the input; the feature maps output by the two convolution modules are concatenated, where the standard convolution module is used for dense feature extraction and the depth-wise separable convolution module focuses on lightweight calculation; the channel shuffle operation is performed on the spliced ​​feature map so that the feature information from the standard convolution is adjacent to the corresponding features from the depth-wise separable convolution.

[0069] Specifically, the convolution layer Conv is mainly replaced by the GSConv module; as a core component, GSConv can effectively reduce the complexity of the model and maintain a high detection accuracy on this basis. Through this design, the model can better strike a balance between recognition accuracy and inference speed, and adapt to the real-time requirements of drone image processing. The working principle of the GSConv module includes the following: First, the input is downsampled through a standard convolution operation to reduce the spatial resolution of the input and reduce the amount of calculation. Then, the module further extracts features through deep convolution DWConv. The key design of the GSConv module is to concatenate the results of standard convolution SC and depthwise separable convolution DSC, that is, to combine the output feature maps of the two convolutions, where SC is responsible for dense feature extraction and DSC focuses on lightweight calculations.

[0070] In order to further enhance the utilization efficiency of feature information, GSConv introduces a shuffle operation after splicing, which rearranges the channels generated by the first two convolutions SC and DSC so that the feature information from the standard convolution is adjacent to the features from the depthwise separable convolution. The purpose of this operation is to maximize the feature quality of the DSC output and make it as close as possible to the feature representation generated by SC, thereby retaining the expressive power of the standard convolution while enjoying the computational efficiency improvement brought by DSC. This design of infiltrating the information generated by the standard convolution into the depthwise separable convolution information through shuffle ensures that the information generated by the standard convolution is fully mixed and avoids complex operations. The advantage of this method is that the fusion of information is natural and efficient, and the dense feature information from SC can be fully integrated with the output information of DSC without increasing the amount of computation, thereby maintaining a high feature expression ability and accuracy in the lightweight model.

[0071] S3.2. Using a bidirectional feature pyramid network structure as a feature fusion mechanism, we take the multi-scale feature maps from the backbone network as input and fuse high-level semantic information with low-level spatial information through top-down and bottom-up bidirectional cross-scale connections.

[0072] A weighted feature fusion mechanism is introduced to dynamically adjust the weights of features at different levels during the fusion process, output the fused multi-scale feature map, and improve the detection capability of targets of different scales.

[0073] Specifically, because drone images are typically captured from high altitudes, many objects appear as small objects in the image, posing a greater challenge to small object detection. The characteristic information of small objects primarily resides in shallow networks, but as the network depth increases, this characteristic information may be gradually lost, resulting in a decrease in small object detection accuracy.

[0074] The present invention adopts an improved bidirectional feature pyramid network BiFPN (Bidirectional FeaturePyramid Network) to enhance the generalization ability of shallow networks for spatial features. BiFPN promotes the full fusion of underlying semantic features and high-level semantic features through repeated bidirectional cross-scale connections, and introduces a weighted feature fusion mechanism. This mechanism can dynamically adjust the weights of features at different levels in the fusion process, highlighting the contribution of key features while suppressing the influence of irrelevant or secondary features. Unlike PAFPN, which only supports one-way top-down feature transfer, BiFPN supports bidirectional feature transfer, that is, it not only allows features to be transferred from shallow networks to deep networks (bottom-up), but also supports feature feedback from deep networks to shallow networks (top-down). The design of bidirectional feature transfer enables the network to more effectively utilize feature information of different scales, thereby significantly improving the detection performance of multi-scale targets, especially in drone images, the detection effect of small targets is more significant.

[0075] S3.3, using a dynamic detection head to improve detection performance by introducing different self-attention mechanisms in the three dimensions of the feature tensor;

[0076] In terms of scale perception, it dynamically integrates feature information at different levels in the feature pyramid. In terms of spatial perception, it enhances sensitivity to different spatial positions in the image through the spatial attention mechanism. In terms of task perception, it applies the attention mechanism on feature channels to adaptively activate the corresponding feature channels according to the requirements of different tasks.

[0077] The multi-scale feature maps from the neck network are fed into their respective detection branches, where local information is further extracted through the convolutional layer, and features related to the target category and bounding box are generated. The classification and regression modules are used to predict the category probability and corresponding bounding box coordinates of each target, respectively, to obtain the detection results.

[0078] Specifically, this invention utilizes a dynamic detection head. YOLOv8's existing detection head uses a fixed feature fusion method, making it unable to flexibly adjust feature extraction based on target scale differences. This performs poorly when processing objects of varying sizes, such as roads and bridges, resulting in insufficient semantic information for large objects and missing detail information for small objects. Secondly, in terms of scale perception, while feature pyramids can provide multi-scale features, they cannot effectively handle objects of extreme size differences. This is especially true when large and small objects coexist, often limiting detection performance.

[0079] The Dynamic Head, a head framework, improves detection performance by incorporating different self-attention mechanisms across the three dimensions of the feature tensor. First, in terms of scale-awareness, Dynamic Head dynamically fuses feature information from different levels of the feature pyramid to adapt to differences in target scale. For example, in the task of recognizing roads and bridges, where roads are larger and bridges are smaller, Dynamic Head effectively distinguishes objects of different scales, avoiding the semantic discrepancies caused by traditional feature scaling and thus improving detection performance. Second, in terms of spatial-awareness, Dynamic Head employs a spatial attention mechanism to enhance sensitivity to different spatial locations within the image. This allows it to effectively focus on foreground objects and mitigate the influence of background noise, particularly in complex drone imagery. This is particularly important for regular, small objects like bridges, and also facilitates the accurate localization of road objects with diverse morphologies. Finally, in terms of task-awareness, Dynamic Head applies an attention mechanism across feature channels to adaptively activate corresponding feature channels based on the requirements of different tasks (such as classification and bounding box regression). In road and bridge object detection, Dynamic Head intelligently adjusts feature usage based on task requirements, improving classification and positioning accuracy. This not only enhances the model's multi-scale processing and spatial perception capabilities, but also improves its robustness in complex scenarios, significantly improving the overall performance of road and bridge recognition tasks.

[0080] S3.4. Define the outlier degree to describe the quality of the anchor box and assign a gradient gain to it;

[0081] Dynamically adjust the gradient gain according to the degree of outlier of the anchor box;

[0082] Among them, the gradient gain assigned to the anchor box with smaller outliers is greater than the gradient gain assigned to the anchor box with larger outliers, so that the bounding box regression focuses on the anchor box of normal quality while preventing low-quality examples from generating large harmful gradients;

[0083] Through the dynamic anchor box quality division standard, a gradient gain allocation strategy that suits the current situation is made at each moment for bounding box regression.

[0084] Specifically, images captured by drones often contain a large number of small objects and low-quality samples. Existing models over-penalize the geometric properties of these samples (such as center point distance and aspect ratio). In particular, when the aspect ratio of the predicted box is similar to that of the true box, the optimization effect is often insufficient. This can lead to unstable detection of small objects by the model, and even convergence fluctuations during training. This phenomenon is particularly noticeable in drone images, because the size of the objects varies greatly and the scenes are complex. The impact of low-quality samples will be more prominent, which will weaken the generalization ability of the model and affect the overall detection performance. In order to solve these problems in drone image target recognition,

[0085] This paper adopts Wise-IoU (WIoU) as the loss function, introduces the concept of "outlier degree" to evaluate the quality of anchor boxes, and provides an intelligent gradient gain allocation strategy to more effectively deal with small target detection in drone images. The formula is defined as follows:

[0086] L WIoUv1 =R WIoU L WIoU

[0087]

[0088] L IoU =1-R IoU

[0089] Where (x,y) and (x gt ,y gt ) are the coordinates W of the center points of the anchor box and the target box respectively g and H g is the size of the minimum bounding box, R IoU It represents the intersection over union (IoU), which is a common measure of the degree of overlap between the predicted box and the true box.

[0090] L WIoUv2 Based on L WIoUv1 Constructed monotonic focusing coefficient This coefficient is introduced to better focus the model’s attention, especially on examples that are difficult to classify or perform poorly on bounding box regression.

[0091]

[0092] But during model training, the gradient gain With L IoU The decrease of L leads to a slower convergence speed in the later stage of training. IoU The mean of is used as the normalization factor:

[0093]

[0094] Among them The momentum is the sliding average of m, and the normalization factor is dynamically updated to make the gradient gain The overall performance remains at a high level, solving the problem of slow convergence in the late stages of training.

[0095] L WIoUv3 In L WIoUv1 The outlier degree is defined based on to describe the quality of the anchor box, which is defined as:

[0096]

[0097] A small outlier means a high-quality anchor box, and we assign a large gradient gain to it so that the bounding box regression focuses on anchor boxes of normal quality. Assigning a smaller gradient gain to anchor boxes with large outliers will effectively prevent low-quality examples from generating large harmful gradients. We use β to construct a non-monotonic focusing coefficient and apply it to WIoUv1:

[0098]

[0099] Among them, when β=δ, δ makes r=1. When the outlier degree of the anchor box satisfies β=C (C is a constant), the anchor box will obtain the highest gradient gain. It is dynamic, and the quality division criteria of the anchor boxes are also dynamic, which enables WIoU v3 to make the gradient gain allocation strategy that best suits the current situation at every moment.

[0100] Compared to CIOU, WIoU reduces the competition for high-quality anchor boxes while mitigating the harmful gradient effects of low-quality samples. This mechanism can focus on medium-quality anchor boxes in dense scenes such as drone imagery, especially complex multi-scale objects, thereby improving the model's performance in detecting small and edge objects, making object detection in drone imagery more accurate and stable.

[0101] S4. Using the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired UAV image; visually displaying the recognition results through a result display module;

[0102] Specifically, the detection results are calibrated on the image through bounding boxes, category labels are added to the detected objects, the corresponding confidence scores are annotated for each bounding box, and an image output containing the bounding box, object category label and confidence score is generated;

[0103] By visualizing the image, the road and bridge targets detected by the model are intuitively presented;

[0104] Generate reports for subsequent analysis by batch processing multiple images.

[0105] In one example, S3 further includes evaluating the performance of the model after optimization by the loss function through a model evaluation module, and the evaluation indicators include mAP50, mAP50-95, precision, and recall.

[0106] Setting hyperparameters is crucial during training. The learning rate chosen in this paper is 0.001. The batch size is set to 4, depending on the GPU memory capacity used. An appropriate batch size helps speed up training and improve model stability. The number of iterations is 300 epochs, which can be dynamically adjusted later based on the convergence of the loss function during training and performance on the validation set.

[0107] The experimental configuration is: Windows 10 system; GPU is NVIDIA RTX 3060 6GB; CPU is 11th Gen Intel(R) Core(TM) i7-11800H. The development environment is Python 3.10, CUDA 12.3, and PyTorch 2.3.1. The evaluation indicators are as follows:

[0108] mAP50: mAP (mean Average Precision), which represents the average precision over multiple categories. mAP50 represents the mAP value at an IoU threshold of 50%.

[0109] mAP50-95: This is a stricter evaluation metric that calculates the mAP value within the 50-95% IoU threshold range and then takes the average. This can more accurately evaluate the performance of the model at different IoU thresholds.

[0110] Precision: Precision is the proportion of positive samples that the evaluation model predicts correctly. In object detection, if the bounding box predicted by the model coincides with the true bounding box, the prediction is considered correct.

[0111] Recall: Recall is the ratio of all true positive samples that the evaluation model can find. In object detection, if the true bounding box coincides with the predicted bounding box, the sample is considered to be correctly recalled.

[0112] Where Precision = TP / (TP+FP), Recall = TP / (TP+FN)

[0113] TP (True Positive): Predicting the positive class as the positive class number is a correct prediction. If the actual number is 0, the prediction is also 0.

[0114] FN (False Negative): Predicting the positive class as the negative class is an incorrect prediction, the actual value is 0, and the prediction is 1.

[0115] FP (False Positive): Predicting a negative class as a positive class is an incorrect prediction, the actual value is 1, and the prediction is 0.

[0116] TN (True Negative): predicts the negative class as the negative class number, that is, the correct prediction, the actual is 1, and the prediction is also 1.

[0117] The following are the comparative experimental results after various model training:

[0118] Table 1. Comparison of target detection results of different algorithms on road and bridge datasets

[0119]

[0120] The experimental results in Table 1 clearly demonstrate the gradual improvement in mAP50 after the introduction of different improved modules in YOLOv8n. The base model YOLOv8n achieved an mAP50 of 0.838, which is considered average performance as a baseline model. The introduction of BiFPN significantly improved mAP50 to 0.894, a 5.6 percentage point increase, demonstrating that BiFPN significantly enhances feature extraction and multi-scale object detection. Using Wise-IoU v3 as the loss function, mAP50 increased to 0.883, a 4.5 percentage point increase, demonstrating that Wise-IoU v3 effectively optimizes object bounding box regression accuracy. Combining GSConv with the Slim-Neck architecture achieved a mAP50 of 0.890, a 5.2 percentage point improvement, demonstrating that this lightweight architecture maintains high accuracy while also balancing detection speed. After further introducing the DynamicHead architecture, mAP50 improved to 0.897, a 5.9 percentage point increase, demonstrating that the dynamic detection head can better optimize object classification and localization accuracy in object detection tasks. Finally, a comprehensive improved model integrating BiFPN, Wise-IoU v3, DynamicHead, GSConv, and Slim-Neck increased mAP50 to 0.906, a 6.8 percentage point improvement, achieving an optimal balance of overall detection performance. Compared to single improvements, the combined improvement modules enable the model to maintain high precision while improving recall and enhancing robustness to objects of varying scales, making it particularly suitable for complex road and bridge object detection tasks.

[0121] Furthermore, the data in the table shows that by further optimizing model performance through data augmentation, mAP50 reached 0.912, a 7.4 percentage point improvement. This also improved mAP50-95 and recall, demonstrating that data augmentation significantly enhances the model's generalization and adaptability to a wider range of scenarios and targets. Overall, the significant improvement in mAP50 achieved by the multi-module integrated YOLOv8n model demonstrates its enormous potential in practical applications and its ability to effectively address object detection requirements in complex environments.

[0122] Example 2:

[0123] The present invention provides a road and bridge recognition system based on drone images, the system comprising:

[0124] Data acquisition module, used to obtain road and bridge image datasets taken by drones;

[0125] A data preprocessing module, configured to perform image enhancement on the data set through the data preprocessing module;

[0126] The recognition model training module is used to establish a road and bridge recognition model, using a backbone network to extract features of the image enhancement dataset to generate a feature map; a neck network connected to the backbone network fuses the feature map and performs multi-scale feature transfer to generate an enhanced multi-scale feature map; a detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; and a loss function is used to calculate the error of the detection result. When the error meets the standard, the road and bridge recognition model is trained.

[0127] The measurement display module is used to use the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired drone images; the recognition results are visualized through the result display module.

[0128] While various embodiments of the present invention have been described above, the above description is intended to be illustrative, not exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A UAV road and bridge recognition method based on a lightweight deep neural network model, characterized in that: include: S1. Obtain a dataset of road and bridge images taken by a drone; S2. performing image enhancement on the data set through a data preprocessing module; S3. Build a road and bridge recognition model, use the backbone network to extract the features of the image-enhanced dataset, and generate a feature map. The neck network connected to the backbone network fuses the feature maps and transfers multi-scale features to generate an enhanced multi-scale feature map; the detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; the error of the detection result is calculated using a loss function, and the road and bridge recognition model is trained after the error meets the standard; S4. Use the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired UAV image; and visualize the recognition results through a result display module.

2. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S2, the data set is enhanced by a data preprocessing module to increase the diversity of the data set. Include one or more of the following steps: Randomly crop the images of the dataset to generate images of different sizes; Flipping and rotating the images of the dataset to generate images in different orientations; Adjusting the brightness of the images in the dataset to simulate different lighting conditions; Adding Gaussian noise to the image of the data set to generate a noisy image; Performing contrast transformation on the images of the dataset to generate images with different contrasts.

3. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S3, the backbone network is used to extract features of the image enhanced dataset and generate a feature map, which specifically includes: Use the GSConv convolution module to perform multi-layer convolution operations on the input image to gradually identify the edge, texture, and shape features of the input image; A feature map is generated based on the extracted features, the spatial resolution of the feature map is gradually reduced, the number of channels is gradually increased, deep semantic information is extracted, and a multi-scale feature map is output to provide spatial and semantic features of objects of different sizes.

4. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 3 is characterized in that The GSConv convolution module includes a standard convolution module and a depth-separable convolution module; The standard convolution module and the depth-wise separable convolution module are used sequentially to downsample the input image and reduce the spatial resolution of the input; The feature maps output by the two convolution modules are concatenated. The standard convolution module is used for dense feature extraction, while the depthwise separable convolution module focuses on lightweight computation. The concatenated feature maps are subjected to a channel rearrangement shuffle operation so that the feature information from the standard convolution is adjacent to the corresponding features from the depthwise separable convolution.

5. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S3, the neck network connected to the backbone network fuses the feature maps and transfers multi-scale features to generate an enhanced multi-scale feature map, including: Using a bidirectional feature pyramid network structure as a feature fusion mechanism, the multi-scale feature maps from the backbone network are used as input, and high-level semantic information and low-level spatial information are fused through top-down and bottom-up bidirectional cross-scale connections; A weighted feature fusion mechanism is introduced to dynamically adjust the weights of features at different levels during the fusion process, and output the fused multi-scale feature map to improve the detection capability of targets of different scales.

6. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S3, the detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result, including: A dynamic detection head is used to improve detection performance by introducing different self-attention mechanisms in the three dimensions of the feature tensor. In terms of scale perception, it dynamically integrates feature information at different levels in the feature pyramid. In terms of spatial perception, it enhances sensitivity to different spatial positions in the image through the spatial attention mechanism. In terms of task perception, it applies the attention mechanism on feature channels to adaptively activate the corresponding feature channels according to the requirements of different tasks. The multi-scale feature maps from the neck network are fed into their respective detection branches, where local information is further extracted through the convolutional layer, and features related to the target category and bounding box are generated. The classification and regression modules are used to predict the category probability and corresponding bounding box coordinates of each target, respectively, to obtain the detection results.

7. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S3, the WIoU loss function is used as the loss function for bounding box regression, and the concept of outlier is introduced to evaluate the quality of the anchor box, including; Define outlier to describe the quality of the anchor box and assign a gradient gain to it; Dynamically adjust the gradient gain according to the degree of outlier of the anchor box; Among them, the gradient gain assigned to the anchor box with smaller outliers is greater than the gradient gain assigned to the anchor box with larger outliers, so that the bounding box regression focuses on the anchor box of normal quality while preventing low-quality examples from generating large harmful gradients; Through the dynamic anchor frame quality division standard, a gradient gain allocation strategy that suits the current situation is made at each moment.

8. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that S3 also includes evaluating the performance of the model after optimization by the loss function through a model evaluation module, and the evaluation indicators include mAP50, mAP50-95, precision, and recall.

9. The UAV road and bridge recognition method based on a lightweight deep neural network model as claimed in claim 1 is characterized in that In S4, visually displaying the recognition result by a result display module includes: The detection results are calibrated on the image through bounding boxes, category labels are added to the detected objects, and the corresponding confidence scores are annotated for each bounding box. The image output is generated, which includes the bounding box, object category label and confidence score. By visualizing the image, the road and bridge targets detected by the model are intuitively presented; Generate reports for subsequent analysis by batch processing multiple images.

10. A UAV road and bridge recognition system based on a lightweight deep neural network model, characterized by The system includes: Data acquisition module, used to obtain road and bridge image datasets taken by drones; A data preprocessing module, configured to perform image enhancement on the data set through the data preprocessing module; The recognition model training module is used to establish a road and bridge recognition model, using a backbone network to extract features of the image enhancement dataset to generate a feature map; a neck network connected to the backbone network fuses the feature map and performs multi-scale feature transfer to generate an enhanced multi-scale feature map; a detection head connected to the neck network performs target classification and bounding box regression on the enhanced multi-scale feature map to obtain a detection result; and a loss function is used to calculate the error of the detection result. When the error meets the standard, the road and bridge recognition model is trained. The measurement display module is used to use the trained road and bridge recognition model to perform road and bridge recognition on the real-time acquired drone images; the recognition results are visualized through the result display module.

Citation Information

Cited By

  • Bridge vibration displacement detection method and system based on image enhancement and space-time tracking

    CN121639806A