Dual-mode-based all-time unmanned aerial vehicle traffic infrastructure inspection system and method
By combining dual-modal sensors and the YOLO11 algorithm, efficient detection of traffic infrastructure by drones is achieved around the clock, solving the problems of insufficient light at night and excessive consumption of computing resources, and improving the accuracy of drone inspections and computer performance.
Patent Information
- Application Number
- CN202510990352.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-28
AI Technical Summary
Existing drone inspection technology mainly relies on single-modal imaging equipment. The detection accuracy decreases when there is insufficient light at night, and it consumes too much local computing resources, affecting the efficiency of inspections around the clock and computer performance.
By employing dual-modal fusion infrared and visible light sensors, combined with the YOLO11 algorithm model, all-weather detection of traffic infrastructure is achieved. The local computing pressure is alleviated by using a cloud server, and Websocket and FFmpeg technologies are used to switch between real-time and offline detection and to transmit data.
It enables efficient inspection of drones under all-day conditions, improves detection accuracy, alleviates the computing and storage pressure on local computers, supports real-time and offline detection modes, and enhances the stand-alone working capability of drones.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of drone identification, and mainly to an all-weather drone-based traffic infrastructure inspection system and method based on dual-modality operation. Background Technology
[0002] Transportation infrastructure is the fundamental framework for a city's economic activities. As a key hub connecting industrial chains, efficient transportation infrastructure can significantly reduce logistics costs and improve the efficiency of economic activities. However, over time, due to a lack of corresponding maintenance and repair measures, as well as the impact of prolonged weather conditions, some transportation infrastructure begins to age rapidly, posing safety hazards that endanger personal safety and even urban economic development. For example, ground signs become damaged or blurred due to prolonged friction between the road and vehicles, as well as erosion from rainwater and sand. If not detected and addressed in time, this can lead to driver misjudgment and unnecessary traffic accidents. Secondly, prolonged immersion in water on bridges causes cracks in the bridge supports, threatening the stability of the bridges. Furthermore, poorly maintained road surfaces crack and subside due to prolonged traffic pressure and weather conditions, forming potholes. Slope protection nets corrode and break down after heavy rains and other severe weather, losing their protective function and allowing rocks to fall onto the roads. If these transportation infrastructure issues are not detected and maintained in a timely manner, they will seriously threaten personal safety and even lead to urban economic losses. Therefore, timely screening, detection, and maintenance of these potentially hazardous transportation infrastructure facilities are of great importance.
[0003] With the development of the low-altitude economy, many scholars have used drones as the main tool for carrying out research technologies and expanding the application fields of drones. Among them, many domestic and foreign researchers focus on combining target detection technology with drone development platforms, using monocular cameras as the data receiving entry point. However, at night, due to insufficient light, the image generated by a single imaging device cannot fully represent all the information of the detected target. Therefore, dual-modal feature images formed by the fusion of infrared and visible light have become an important research technology for all-weather detection.
[0004] Most existing solutions for drone inspection technology are based on single-modal automated inspection. Specifically, visible light images are used as input data, and a deep learning strategy is used to train a model capable of recognizing specific scenes. This model is then integrated with the drone, which uses a monocular camera to capture scene data along its flight path during flight. The captured video data is transmitted in real time to a local computer, which then projects the results generated by the trained model onto a monitor for researchers to observe specific targets. However, single-modal automated inspection technology still has certain limitations. For example, this strategy is only suitable for daytime scenes with good lighting. At night, due to insufficient light, the monocular camera cannot capture enough feature information about the target, leading to a significant decrease in the model's recognition accuracy. Therefore, implementing an all-day strategy can better improve the utilization efficiency of drones. Secondly, existing technologies typically store the training model on a local computer. Model inference and result storage consume significant computing resources and storage space, increasing the storage burden on the local computer. Summary of the Invention
[0005] This invention aims to improve the all-day operational capability of drones for detecting traffic infrastructure, while simultaneously alleviating the computational and storage pressure on local computers and further enhancing the stand-alone operational capability of drones.
[0006] To address the aforementioned issues, this invention provides a dual-modal, all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system. This system comprises: a data acquisition module, a target detection module, a data processing module, a transmission module, and local and cloud modules.
[0007] The data acquisition module is a device in which a drone flies to a designated area and uses infrared and visible light sensors to collect data from the same area.
[0008] The target detection module uses the YOLO11 algorithm model to detect targets in the video or image of the traffic infrastructure to be detected. This algorithm model is integrated into the UAV.
[0009] The data processing module transforms the detected target data to improve its readability.
[0010] The transmission module is used to transmit the detection results and data processing content of the target detection module to the ground station or cloud server, as well as to transmit the instructions issued by the ground station.
[0011] The local and cloud modules are used to save or stream the detection content transmitted by the transmission module, and to send commands to the drone to achieve different detection tasks.
[0012] The data acquisition module requires that the images or videos of traffic infrastructure collected by the UAV be converted into image data by frame extraction, and that the images be manually labeled using the LabelImg tool to construct a dataset. The dataset is then divided into training set, validation set and test set in a ratio of 8:1:1.
[0013] The target detection module uses YOLO11 as the algorithm network and YOLO11n as the pre-trained weights to train the model on the dataset from the data acquisition module, generating a detection model file. The drone carries the model file to perform real-time or offline detection on the video or image data to be detected. Real-time detection involves detecting images or videos of traffic infrastructure data acquired in real-time by the drone during flight. Offline detection involves detecting locally stored videos or images by the drone when it is not in flight.
[0014] The data processing involves extracting the target information output by the target detection module after real-time or offline detection, converting the extracted target category information into readable text information, and saving the target's coordinate information and confidence level in text format.
[0015] The transmission module pushes the real-time detected video stream to a public address, and the ground station obtains the real-time detection results by pulling the streaming media. The detected images are then directly uploaded to the cloud server. For offline detected images or video streams, the storage address of the video or image is obtained to acquire the resources, and the results are uploaded to the cloud server after offline detection.
[0016] The drone commands mentioned above are commands for switching between different training models and for offline or real-time detection.
[0017] A dual-modal, all-weather unmanned aerial vehicle (UAV) system and method for inspecting traffic infrastructure, wherein the system is a dual-modal, all-weather UAV system for inspecting traffic infrastructure, and the method is implemented through the following steps:
[0018] S1, collects data from a specified area and constructs a training set through the data acquisition module;
[0019] S2, the initialization parameters for training the YOLO11 algorithm model are set, the completed dataset is input into the module for training, and a low-altitude traffic infrastructure inspection model with detection capabilities is obtained and mounted on a drone.
[0020] S3 receives task instructions sent by the ground station and switches between real-time and offline detection processes based on the detection model.
[0021] S4, real-time detection transmits the drone inspection video stream or image to the algorithm module, while offline detection downloads the video or image according to the instruction address and transmits the video or image to the algorithm module.
[0022] S5, process the detected data to form text data;
[0023] S6 pushes the real-time detection video stream to a public address and uploads the detection results and data processing text to a cloud server. Offline detection only uploads the detection results and data processing text to the cloud server.
[0024] In S2, the initialization parameters are based on YOLO11 pre-trained weights. The YOLO11 algorithm network is used to train the input dataset. The network mainly consists of an infrared and visible light feature extraction module, a feature fusion module, and a head network.
[0025] Furthermore, the infrared and visible light feature extraction modules consist of a convolution module and a residual module to extract shallow and deep feature information from infrared and visible light images, respectively. The convolution module is composed of 3x3 convolution kernels and is used to extract shallow features of the image. The residual module is a C3k2 module used to extract deep features of the image.
[0026] Furthermore, the feature fusion module fuses shallow feature information from infrared images with shallow feature information from visible light, and fuses deep feature information from infrared images with deep feature information from visible light. Finally, the shallow fusion information and the deep fusion information are combined to generate the final fused image, thereby improving the UAV's detection capabilities in nighttime scenes.
[0027] Furthermore, the head network is responsible for performing the final detection on the feature map after feature fusion.
[0028] The instructions described in S3 establish a bidirectional communication system using WebSocket.
[0029] S6 describes a streaming channel constructed using the FFmpeg streaming algorithm.
[0030] The beneficial effects of this invention include:
[0031] 1. The drone is equipped with a dual-modal detection system, which improves the drone's ability to inspect traffic infrastructure around the clock and reduces the computing pressure on the local computer.
[0032] 2. Drones can detect traffic infrastructure inspection results in real time, extract key information from detected targets and return it to the local computer for storage, which alleviates the local computing pressure and makes it easier for researchers to observe.
[0033] 3. Local commands enable the switching of different models for different detection content, improving the accuracy of UAVs in detecting different inspection tasks. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the overall process framework of the present invention.
[0036] Figure 2 This is a flowchart illustrating the training process of the dual-modal model adopted in this invention.
[0037] Figure 3 This is a flowchart of the core algorithm of the present invention. Detailed Implementation
[0039] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0040] like Figure 1 As shown, this embodiment provides a dual-modal all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system and method. The system includes: a data acquisition module, a target detection module, a data processing module, a transmission module, and local and cloud modules.
[0041] S1: Construct a dataset for traffic infrastructure inspection.
[0042] Specifically, the drone flies to a designated area and uses infrared and visible light sensors to collect data on the same area at different times. The collected image data is manually labeled using LabelImg to construct a bimodal training dataset, which is then divided into training, testing, and validation sets in an 8:1:1 ratio.
[0043] S2: Training a traffic infrastructure inspection model.
[0044] Specifically, YOLO11 is the training algorithm, YOLO11n.pt contains the pre-trained weights, and training parameters are set, for example, batch size is 8, input image size is 640*640, training iterations are 100, SGD gradient learning strategy is used, and the initial learning rate is 0.001. The model training process is as follows: Figure 2 As shown, this step includes feature extraction, feature fusion, and a head network. The feature extraction process includes the following steps:
[0045] Step 1: Input image; The input to the training model is image information with 6 channels, where the first 3 channels represent visible light images and the last 3 channels represent infrared images.
[0046] Step 2: Convolutional Layer; A convolutional layer with a kernel of 3 and a stride of 2 is used to process the infrared and visible light images in parallel, outputting infrared and visible light feature maps with shallow features. The feature information is processed through an activation function and a batch normalization layer to give the features non-linear characteristics and prevent gradient vanishing from causing overfitting. The stride of 2 gradually reduces the size of the output feature map, preparing for subsequent extraction of deeper features from the image.
[0047] Step 3: C3k2 layer; The C3k2 layer is mainly a bottleneck layer connected in a residual manner. It can obtain features at different levels of infrared and visible light images through multiple convolution operations to form deep features. and The residual is structurally formed by a concat operation between the first and last convolutional layers.
[0048]
[0049]
[0050] in, Represents an infrared image. Represents a visible light image. Represents infrared feature map, This represents a visible light characteristic map.
[0051] Step 4: Feature Fusion; The feature fusion process involves performing a concat operation on features at different levels after feature extraction to form fused features at different levels. The calculation formula is as follows:
[0052]
[0053]
[0054]
[0055] in, These represent feature maps at different levels.
[0056] Step 5: Head Network; The head network is responsible for the final detection on the feature map after feature fusion, including classification and regression tasks, and generating the final bounding boxes and class confidence scores.
[0057] After the iterative training is completed, the corresponding scene model best.pt file can be generated.
[0058] Step 6: Algorithm transfer; The trained model is transferred to the UAV algorithm system.
[0059] S3: Drone Inspection; The drone inspection process mainly includes modules such as drone communication, model inference, data processing, and data transmission, and its flow is as follows: Figure 3 As shown. The specific steps are as follows:
[0060] Step 1: Communication and interaction between the UAV and the ground station; the UAV is the object being accessed, and the ground station accesses it via WebSocket. The WebSocket is constructed using the WebSockets module of the Python library for real-time asynchronous communication with the locally accessed server. Heartbeat packets are periodically sent to the locally accessed server to prevent the WebSocket module from losing connection with the server, and an anomaly detection mechanism is implemented to ensure reconnection after communication failure.
[0061] Step 2: The ground station sends task instructions, which mainly include whether to perform offline or real-time detection, and the model selection for the detection task. The instructions for offline detection include the address of the detection resource.
[0062] Step 3: Set up the drone storage and detection resource files; before model inference, first build a folder to store the model inference results and the resources to be detected. This process is created using the os library in Python.
[0063] Step 4: Recognize Instructions; The drone receives task instructions from the ground station via WebSocket, such as: receiving RTSP keywords from the backend to enable real-time drone detection, or to perform offline static image or video stream detection. The key corresponding to the model is used as the keyword for model switching, thus controlling the selection of the operating model.
[0064] Step 5: Offline detection and resource retrieval; First, the images or video streams to be detected need to be downloaded locally using a multi-process approach, i.e., by sending a POST request to save the corresponding resources locally by name, and using the time module to record the download time.
[0065] Step 6: Model inference. For both real-time and offline detection, the acquired video streams, images, or downloaded resources are transferred to the OS via command line to enable the model to infer the corresponding resources.
[0066] Step 7: Data processing; In the reasoning process in step 6, firstly generate the category information, location information and confidence level of the target in the image to be reasoned, traverse the category information of the target and convert it into the corresponding category text, and record all the converted information in text format.
[0067] Step 8: Data Transmission; For real-time detection, during model inference, FFmpeg's streaming technology is used to stream each frame of the detected image to a designated address in 640*480 format, achieving the broadcast characteristic of real-time detection. For offline or real-time detection images, after detection, a transmission module is built using the Python library oss2 to store the text file generated in Step 7 and the complete video information after detection. The addresses of the text and video information are placed in a Bucket, enabling the drone to upload the information to the cloud server.
[0068] S9: Ground station streaming; During real-time monitoring, the ground station uses streaming media software to stream the address broadcast by the drone so that local researchers can view the drone's inspection content in real time.
Claims
1. A dual-modal, all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system, characterized in that, The system includes a data acquisition module, a target detection module, a data processing module, a transmission module, and local and cloud modules. The data acquisition module involves the UAV flying to a designated area and using infrared and visible light sensors to collect data from the same area. The target detection module uses the YOLOv11 algorithm model to detect targets in videos or images of traffic infrastructure; this algorithm model is integrated into the UAV. The data processing module transforms the detected targets to improve data readability. The transmission module transmits the detection results and processed data from the target detection module to a ground station or cloud server, as well as transmitting commands issued by the ground station. The local and cloud modules store or stream the detection content transmitted by the transmission module and send commands to the UAV to achieve different detection tasks.
2. The all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system based on dual-modality as described in claim 1, characterized in that, The data acquisition module requires that images or videos of traffic infrastructure collected by UAVs be converted into image data by frame extraction, and that the images be manually labeled using the LabelImg tool to construct a dataset. The dataset is then divided into training, validation and test sets in an 8:1:1 ratio.
3. The all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system based on dual-modality as described in claim 2, characterized in that, The target detection module uses YOLO11 as the algorithm network and YOLO11n as the pre-trained weights to train the model on the dataset from the data acquisition module, generating a detection model file. The UAV carries the model file to perform real-time or offline detection on the video or image data to be detected. Real-time detection refers to the detection of images or videos of traffic infrastructure data acquired in real-time by the UAV during flight. Offline detection refers to the detection of locally stored videos or images by the UAV when it is not in flight.
4. The all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system based on dual-modality as described in claim 3, characterized in that, The data processing involves extracting the target information output by the target detection module after real-time or offline detection, converting the extracted target category information into readable text information, and saving the target's coordinate information and confidence level in text format.
5. The all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system based on dual-modality as described in claim 4, characterized in that, The transmission module pushes the real-time detected video stream to a public address, and the ground station obtains the real-time detection results by pulling the streaming media. The real-time detected images are then directly uploaded to the cloud server after detection. For offline detected images or video streams, the storage address of the video or image is obtained to acquire the resources, and the results are uploaded to the cloud server after offline detection.
6. The all-weather unmanned aerial vehicle (UAV) traffic infrastructure inspection system based on dual-modality as described in claim 5, characterized in that, The aforementioned drone commands include switching between different training models and offline or real-time detection commands.
7. A method for all-weather unmanned aerial vehicle (UAV) inspection of transportation infrastructure based on dual-modality operation, employing the all-weather unmanned aerial vehicle (UAV) inspection system based on dual-modality operation as described in claim 6, characterized in that... The process includes the following steps: S1, collecting data from a designated area using the data acquisition module and constructing a training set; S2, setting initialization parameters for the YOLO11 algorithm model, inputting the constructed dataset into the module for training, obtaining a low-altitude traffic infrastructure inspection model with detection capabilities, and mounting it on the UAV; S3, receiving task instructions from the ground station, and switching between real-time and offline detection processes based on the detection model; S4, for real-time detection, transmitting the UAV inspection video stream or images to the algorithm module, and for offline detection, downloading the video or images from the instruction address and transmitting them to the algorithm module; S5, processing the detected data to form text data; S6, pushing the real-time detection video stream to a public address, and uploading the detection results and data processing text to the cloud server, while for offline detection, only uploading the detection results and data processing text to the cloud server.
8. The all-weather unmanned aerial vehicle (UAV) inspection method for transportation infrastructure based on dual-modality as described in claim 7, characterized in that, The initialization parameters in S2 use YOLO11 pre-trained weights as initial weights. The YOLO11 algorithm network is used to train the input dataset. The network mainly consists of an infrared and visible light feature extraction module, a feature fusion module, and a head network (S2.1). The infrared and visible light feature extraction module is composed of a convolutional module and a residual module to extract shallow and deep feature information from infrared and visible light images, respectively. The convolutional module uses 3x3 kernels to extract shallow features. The residual module is a C3k2 module used to extract deep features. (S2.2) The feature fusion module fuses the shallow feature information of the infrared image with the shallow feature information of the visible light image, and fuses the deep feature information of the infrared image with the deep feature information of the visible light image. Finally, the shallow and deep fused information are combined to generate the final fused image, which improves the UAV's detection capability in nighttime scenes. (S2.3) The head network is responsible for the final detection on the feature map after feature fusion.
9. The all-weather unmanned aerial vehicle (UAV) inspection method for transportation infrastructure based on dual-modality as described in claim 7, characterized in that, The instructions described in S3 establish a bidirectional communication system using WebSocket.
10. The all-weather unmanned aerial vehicle (UAV) inspection method for transportation infrastructure based on dual-modality as described in claim 7, characterized in that, The streaming described in S6 uses the FFmpeg streaming algorithm to construct the streaming channel.