Large model-based full-modal unmanned aerial vehicle post-flood disaster intelligent patrol system and method
By integrating multimodal sensors and large-scale multimodal models, the UAV system solves the problems of single sensor modes and reliance on manual data processing in existing technologies. It realizes multi-dimensional information collection and real-time comprehensive analysis, improves the efficiency of post-disaster patrols and the comprehensiveness of information acquisition, and provides real-time disaster assessment and intelligent guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-29
AI Technical Summary
Existing drone-based disaster relief systems suffer from several problems in flood relief efforts, including limited sensor modalities or low integration, reliance on manual data processing, restricted inspection efficiency, lack of real-time comprehensive analysis capabilities, lack of real-time interaction and intelligent guidance, and fragmented information presentation and report generation. These issues make it difficult to achieve efficient and comprehensive disaster relief and assessment.
The system employs a large-scale model-based full-modal unmanned aerial vehicle (UAV) system, integrating a high-definition visible light camera, an infrared thermal imager, a lidar, and a microphone array. It combines a large multimodal model (MLLM) to perform multi-source data fusion analysis, enabling real-time data processing and comprehensive evaluation. Furthermore, it achieves human-machine collaboration through voice interaction and automatically generates structured inspection reports.
It enables multi-dimensional information collection under various weather conditions and day and night, reducing reliance on manual interpretation, improving inspection efficiency and the comprehensiveness and accuracy of information acquisition, providing real-time disaster assessment and intelligent guidance, and generating rapid comprehensive inspection reports.
Smart Images

Figure CN122111065A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent inspection technology, specifically relating to a full-modal unmanned aerial vehicle (UAV) intelligent inspection system and method for flood disasters based on a large model. Background Technology
[0002] Floods are among the most severe and frequent natural disasters globally, often causing significant loss of life and property and damage to infrastructure in a short period. Rapid and accurate post-disaster assessment is crucial for emergency response, relief deployment, and recovery.
[0003] Traditional methods of post-flood inspection have many limitations. Ground-based field surveys are the most direct method, but in disaster areas where flooding has occurred, roads are blocked, and the environment is hazardous, this method is inefficient, time-consuming, has limited coverage, and poses significant safety risks to investigators. While satellite remote sensing technology can provide large-scale observational information, its application is constrained by various factors. For example, satellite revisit cycles are long, making it difficult to meet the real-time needs of dynamic disaster monitoring; optical satellites are susceptible to adverse weather conditions such as clouds and rain, making all-weather observation impossible; satellite imagery has relatively low spatial resolution, making it difficult to meet the needs of detailed damage assessments of buildings, roads, and other infrastructure; furthermore, existing remote sensing-based emergency response mechanisms suffer from insufficient data sharing, complex activation procedures, and low response efficiency. While aerial remote sensing using manned aircraft offers higher resolution, it is costly, difficult to schedule, and carries extremely high risks during flights in extreme post-disaster weather conditions. Therefore, there is an urgent need for a more efficient, flexible, safe, and comprehensive method for post-disaster inspection and assessment.
[0004] In recent years, unmanned aerial vehicle (UAV) technology has emerged as a key player in disaster emergency response due to its advantages such as maneuverability, rapid deployment, relatively low cost, ability to acquire high-resolution data, and improved worker safety. Current applications of UAVs in disasters (including floods, earthquakes, and mining accidents) mainly include: using onboard high-definition cameras for disaster area reconnaissance, mapping the scope of disaster impact, and assessing damage to buildings and other infrastructure; using infrared thermal imaging cameras to detect signs of life or abnormal heat sources; and using lidar (LiDAR) for terrain mapping or obstacle detection. Some research has begun to explore the use of artificial intelligence (AI) technologies, such as machine learning or deep learning models, to automate the analysis of image data collected by UAVs, such as classifying building damage.
[0005] However, existing technologies still have significant shortcomings when applied to post-flood inspections: 1. Single sensor modality or low integration: Most systems rely primarily on optical visible light cameras, which have limited ability to acquire information in complex environments such as nighttime, dense smoke, fog, or vegetation obstruction. While some systems may be equipped with infrared thermal imagers or lidar, they often lack the ability to effectively fuse and collaboratively analyze data from multiple sensors. For example, although synthetic aperture radar (SAR) has all-weather observation capabilities, its data processing is usually independent of optical data. Emerging technologies using drones equipped with microphone arrays for sound source localization to search for trapped personnel are mostly in the experimental stage or used as standalone systems, failing to achieve deep integration with visual, thermal imaging, and other information. This reliance on a single modality or the fragmentation of multimodal information limits the comprehensive and accurate perception of complex disaster situations.
[0006] 2. Data processing and analysis rely on manual labor: Although drones can quickly collect large amounts of data, subsequent data processing and information extraction often require significant manual interpretation and analysis. For example, thousands of aerial photographs need to be manually examined to identify damaged buildings or locate signs of trapped individuals. This not only consumes a large amount of time and manpower but also severely delays the speed at which information needed for emergency response decisions can be obtained. While there have been attempts to apply AI to specific tasks (such as image classification), these are usually limited in function and lack the ability to comprehensively understand and intelligently analyze multi-source, multi-modal data.
[0007] 3. Limited Inspection Efficiency: While commonly used multi-rotor drones offer flexible takeoff and landing, their limited endurance and range make it difficult to efficiently cover vast flood-affected areas. Traditional fixed-wing drones, despite their long endurance, require runways for takeoff and landing, making deployment in disaster areas difficult.
[0008] 4. Lack of real-time comprehensive analysis capabilities: Existing systems mostly process and analyze data offline after the flight mission. Although some systems emphasize real-time data transmission, this is usually limited to transmitting raw data or providing basic situational awareness, lacking the ability to perform real-time fusion and in-depth analysis of multimodal data during flight and generate comprehensive disaster assessment results. In particular, the ability to correlate information from different sensors in real time (such as correlating sound signals with thermal imaging anomalies at specific locations) is very limited.
[0009] 5. Lack of Real-Time Interaction and Intelligent Guidance Capabilities: Current UAV systems are typically one-way data acquisition tools, with ground control personnel acting as passive information receivers. The systems lack the ability to interact intelligently with ground command centers or rescue personnel in real time. For example, ground personnel cannot use natural language to query the specific situation "seen" or "heard" by the UAV in real time, nor can they obtain automated emergency suggestions or risk warnings based on real-time multimodal analysis results. Existing communication recovery technologies primarily focus on establishing communication links themselves, rather than achieving intelligent interactive information services.
[0010] 6. Fragmented Information Presentation and Report Generation: Data and analysis results from different sensors are often presented in separate forms, requiring manual summarization and organization to form a comprehensive disaster report. There is a lack of functionality to automatically integrate multimodal analysis results, key image / video clips, geographic location information, risk assessments, and other content, and to quickly generate comprehensive inspection reports in a structured and visualized manner.
[0011] In summary, while existing technologies have made some progress in post-flood patrol applications, significant shortcomings remain in areas such as patrol efficiency, the comprehensiveness and depth of information acquisition, the intelligence and automation of data analysis, real-time interaction and guidance capabilities, and the efficiency of comprehensive report generation. In particular, there is a lack of a holistic solution that organically combines a high-efficiency flight platform, full-modal perception capabilities, a powerful unified AI analysis core, and real-time interactive functions. This technological gap limits the full realization of the potential of drones in post-disaster emergency response, highlighting the need for a more intelligent, automated, and comprehensive post-disaster patrol system. Summary of the Invention
[0012] To address the shortcomings and problems of existing drone inspection systems in post-flood inspection applications, this invention provides a full-modal drone intelligent post-flood inspection system and method based on a large model.
[0013] A large-scale model-based full-modal UAV intelligent patrol system for flood disasters includes a UAV, as well as an integrated multimodal sensor module, a data transmission and communication module, a data processing module, and a ground control station; The integrated multimodal sensor module is installed on the UAV and is used to collect multimodal data information from the disaster area synchronously or asynchronously. The data transmission and communication module is used to transmit multimodal data collected by the integrated multimodal sensor module to the data processing module in real time; The data processing module is used to receive the collected multimodal data, analyze and process the acquired multimodal data, and generate an inspection report based on the processing results. The ground control station includes a mission planning unit, used to plan UAV inspection routes and mission parameters based on disaster information and inspection needs, and to deploy UAVs for takeoff; a flight monitoring unit, used to monitor the UAV flight status display unit, used to display the UAV's real-time video, thermal imaging, point cloud, and map overlay analysis results; a report receiving and display unit, used to receive and display reports generated by the data processing module; and a real-time voice interaction unit, used by ground control personnel to issue query commands to the system via voice and receive the returned voice responses.
[0014] The aforementioned intelligent flood patrol system based on a large model and multimodal unmanned aerial vehicle (UAV) includes a high-definition visible light camera, an infrared thermal imager, a lidar, and a microphone array.
[0015] The aforementioned intelligent flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) also includes a voice recognition module at its ground control station. This voice recognition module is connected to the data processing module and is used to receive voice commands and transmit the received voice commands to the data processing module.
[0016] The aforementioned intelligent flood post-disaster patrol system based on a large model and using unmanned aerial vehicles (UAVs) further includes the following data processing unit: It receives visible light images and LiDAR point cloud data synchronously collected by a multimodal sensor array carried by a drone, uses the TMRoPE position coding algorithm to perform feature projection and time alignment of the visible light images and LiDAR point clouds using a multimodal data fusion architecture, and performs feature fusion through a multi-head attention mechanism. Semantic understanding and structural analysis were performed on the fused data to identify bridge structural anomalies; Mark the bridge status on the map interface and generate an initial text alert, which is then pushed to the ground control station.
[0017] The aforementioned intelligent flood post-disaster patrol system based on a large model and using unmanned aerial vehicles (UAVs) further includes the following data processing unit: It receives environmental sound data and thermal signal data synchronously collected by the microphone array and infrared thermal imager in the multimodal sensor array carried by the UAV, and transmits them to the data processing module in real time; The system performs sound source localization on audio data and estimates the coordinates of the signal source; simultaneously, it performs anomaly detection on thermal imaging data and identifies weak but persistent thermal signals. The TMRoPE algorithm was used to time-align audio and thermal imaging data, and the probability of life signs at each angle was calculated. Bayesian inference was then used to comprehensively determine the presence of life signs at the signal source. The target location is marked as high priority, an inspection report is generated, and the inspection report is sent to the ground control station via voice synthesis as an alarm.
[0018] The aforementioned intelligent flood post-disaster patrol system based on a large model and using unmanned aerial vehicles (UAVs) further includes the following data processing unit: It receives voice commands from ground control station operators, performs natural language parsing on the received voice commands, and extracts the core intent. Based on the core intent, commands are issued to the drone through the flight control system to control the drone to adjust its flight attitude and gimbal angle, perform 360-degree circling flight around high-priority targets, and simultaneously adjust the focal length and focus of the high-definition visible light camera and infrared thermal imager to obtain clearer, multi-angle images. During the orbital flight, it receives real-time visible light and thermal imaging video streams collected by a multimodal sensor array, and continuously analyzes the real-time video streams to understand the content of the images; Based on the combined video analysis results, target geographical location information, and initial commands, a voice response is generated using speech synthesis technology. This voice response, along with the real-time video feed, is then pushed to the ground control station operator interface, completing a full real-time interactive loop.
[0019] The aforementioned intelligent flood post-disaster patrol system based on a large model and using unmanned aerial vehicles (UAVs) further includes the following data processing unit: Receive RTSP video stream from the drone gimbal, periodically capture keyframe images, and perform preprocessing; Semantic segmentation algorithms are used to identify water areas in images, while target detection and classification algorithms are used to identify and classify the damage to buildings. The system automatically generates a flood inundation map based on the segmentation results, and produces a list of damaged buildings and analysis results, including damage type, location and severity. All analysis results are automatically integrated into the map interface and displayed in a visual format, supporting real-time viewing and interaction by ground personnel.
[0020] This invention also provides a method for intelligent post-flood patrol using a full-modal unmanned aerial vehicle (UAV) based on a large model, comprising the following steps: Step 1, Task Planning and Deployment: Based on disaster information and inspection needs, plan the inspection route and mission parameters of the UAV at the ground control station; deploy the UAV to the take-off and landing point and perform vertical take-off; Step 2, Autonomous Inspection and Data Acquisition: The UAV flies autonomously according to the preset route or real-time commands, and at the same time activates the multimodal sensor array to continuously collect visible light images / videos, infrared thermal imaging data, LiDAR point cloud data and environmental sound data of the disaster area; Step 3: Real-time data transmission and processing: The acquired multimodal data is processed in real-time or near real-time via a communication link through a data processing module. Step 4: Intelligent Analysis: The data processing module receives and fuses the received multimodal data streams based on a large-scale multimodal model, and executes the set analysis tasks; Step 5: Report Generation and Mission Completion: The data processing module automatically generates a structured inspection report containing all findings of the current mission based on the inspection results, and sends the inspection report to the ground control station; after the UAV completes the inspection of all predetermined areas, it returns to the take-off and landing point and lands vertically according to instructions or automatically planned paths; the system generates the final inspection report and archives all raw data and reports.
[0021] Compared with the prior art, the beneficial effects of the present invention are: This invention utilizes a VTOL fixed-wing UAV platform, integrating multiple sensors such as visible light, infrared thermal imaging, LiDAR, and microphone arrays. This enables comprehensive, multi-dimensional information collection of the disaster area environment, overcoming the limitations of single-modal sensors and acquiring richer and more reliable data under various weather conditions and obstructions, including day and night. Through an advanced large-scale multimodal model (MLLM), the system can perform unified and in-depth fusion analysis and understanding of multi-source heterogeneous data, automatically completing complex tasks such as flood extent mapping, infrastructure damage assessment, and, in particular, detecting trapped personnel using visual and auditory cues, significantly reducing reliance on manual interpretation. Leveraging the natural language processing and speech generation capabilities of MLLM, unprecedented real-time two-way voice interaction between the system and ground personnel is achieved. Operators can conveniently query information, issue commands, and receive automated early warnings and emergency guidance based on real-time analysis, enhancing human-machine collaboration efficiency and decision support levels.
[0022] The system of this invention can automatically integrate MLLM analysis results, key multimodal evidence (images, videos, point clouds, audio information), geographic location data and risk assessments to quickly generate structured and visualized comprehensive inspection reports, ensuring the timely and accurate transmission of disaster information. Attached Figure Description
[0023] Figure 1 This is a diagram showing the data interaction relationship between the ground station, the UAV, and the data processing module of this invention.
[0024] Figure 2 This is a flowchart of the intelligent flood post-disaster patrol method based on a large model using a full-modal UAV. Detailed Implementation
[0025] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0026] Example 1: This example provides a large-model-based, multi-modal UAV-based intelligent flood disaster patrol system. This system combines the efficient inspection capabilities of a vertical takeoff and landing (VTOL) UAV, the omnidirectional perception capabilities of a multi-modal sensor array, and the powerful analysis and generation capabilities of a large-scale multi-modal model. It includes... The drone is a vertical take-off and landing fixed-wing drone, which has the characteristics of vertical take-off and landing capability, efficient fixed-wing cruise capability, long endurance, high payload, and good wind resistance. It is suitable for deployment and large-scale inspection missions in the complex environment of disaster areas.
[0027] An integrated multimodal sensor module, installed on a drone, includes at least a high-definition visible light camera, an infrared thermal imager, a lidar (LiDAR) system, and a microphone array, for synchronous or asynchronous acquisition of multi-dimensional information about the disaster area.
[0028] The data transmission and communication module employs technologies such as 4G / 5G, satellite communication, or high-bandwidth radio links to ensure stable data transmission between the UAV and the ground control station, as well as with the cloud server (if cloud processing is used).
[0029] The data processing module can be an airborne or cloud-based processing unit. It can be deployed locally on the drone (edge computing), or connected to a cloud server via a communication link, or a hybrid deployment method can be adopted. It includes computing hardware and software, which are used to run large-scale multimodal models to receive and uniformly analyze and process the collected multimodal data, generate inspection reports, and complete voice interaction tasks. The ground control station includes a mission planning unit, which is used to plan the drone inspection route and mission parameters based on disaster information and inspection needs, and to deploy the drone to take off. Flight monitoring unit is used to monitor the flight status of the drone; The display unit is used to display the real-time video, thermal imaging, point cloud, and map overlay analysis results of the UAV; The report receiving and display unit is used to receive and display reports generated by the data processing module; The real-time voice interaction unit is used by ground control personnel to issue query commands to the system via voice and receive voice responses.
[0030] Example 2: Based on the patrol system of Example 1, this example provides a smart patrol method for post-flood disaster using a full-modal UAV based on a large-scale multimodal model, including the following steps: Step 1, Task Planning and Deployment: Based on disaster information and inspection needs, plan the inspection route and mission parameters of the UAV at the ground control station; deploy the UAV to the take-off and landing point and perform vertical take-off.
[0031] Step 2, Autonomous Inspection and Data Acquisition: The UAV flies autonomously according to the preset route or real-time commands, and at the same time activates the multimodal sensor array to continuously collect visible light images / videos, infrared thermal imaging data, LiDAR point cloud data and environmental sound data of the disaster area.
[0032] Step 3: Real-time data transmission and processing: The collected multimodal data is processed in real-time or near real-time through a communication link by a data processing module.
[0033] Step 4: Intelligent Analysis: The data processing module receives and fuses the received multimodal data streams based on a large-scale multimodal model, performing one or more analysis tasks, as follows: (1) Building damage identification and marking S411, Multimodal Data Fusion Processing: The large multimodal model receives visible light images and LiDAR point cloud data synchronously collected by the multimodal sensor array carried by the UAV. It uses the TMRoPE position coding algorithm to perform feature projection and time alignment of the visible light images and LiDAR point clouds using the multimodal data fusion architecture, and performs feature fusion through a multi-head attention mechanism. S412. Intelligent Analysis and Recognition: The large-scale multimodal model performs semantic understanding and structural analysis on the fused data to identify bridge structural anomalies and determine that they are in an "interrupted" state. S413. Result Output and Alarm Generation: The system automatically marks the bridge as "interrupted" on the map interface and generates a preliminary alarm in text form, which is then pushed to the ground control station (GCS).
[0034] (2) Signal detection and verification of trapped personnel S421, Audio and Thermal Imaging Data Acquisition: The large multimodal model receives environmental sound data and thermal signal data synchronously collected by the microphone array and infrared thermal imager in the multimodal sensor array carried by the UAV, and transmits them to the data processing module in real time. S422. Sound source localization and thermal signal analysis: A large-scale multimodal model is used to localize the sound source of audio data and estimate the coordinates of the signal source; at the same time, anomaly detection is performed on thermal imaging data to identify weak but persistent thermal signals. S423, Multimodal Cross-validation: Large-scale multimodal models use the TMRoPE algorithm to time-align audio and thermal imaging data, calculate the corresponding probability of life signs at each angle, and use Bayesian inference to comprehensively determine the presence of life signs at the signal source. S424, High-priority marking and voice alarm: The data processing module marks the target location as high priority, generates an inspection report, and sends an alarm to the ground control station through voice synthesis; for example, "High-priority alarm! Suspected distress knocking sound and life heat source signal detected at coordinates [specific coordinates], it is recommended to verify immediately!"
[0035] (3) Real-time voice interaction and target verification S431. Voice command reception and parsing: The ground control station operator issues commands through the voice interaction page. Its built-in voice recognition module receives the commands and transmits them to the data processing module. The large multimodal model in the data processing module performs natural language parsing on the received voice commands to extract the core intent.
[0036] S432, Command Execution and Sensor Control: The large multimodal model issues commands to the UAV through the flight control system, controls the UAV to adjust its flight attitude and gimbal angle, performs 360-degree circling flight around high-priority targets (such as buildings suspected of having trapped people), and simultaneously adjusts the focal length and focus of the high-definition visible light camera and infrared thermal imager to obtain clearer, multi-angle images.
[0037] S433, Real-time Data Stream Processing and Feedback: During orbital flight, the real-time visible light and thermal imaging video streams acquired by the multimodal sensor array are transmitted back to the data processing module through the data transmission and communication module. The large multimodal model in the data processing module continuously analyzes the real-time video streams to understand the content of the images.
[0038] S434. Multimodal Information Synthesis and Response: The large-scale multimodal model integrates video analysis results, target geographical location information, and initial commands to generate a voice response (e.g., "Performing surround observation, target building structure damaged, signs of activity at second-floor window") using text-to-speech (TTS) technology. This voice response, along with the real-time video feed, is then pushed to the ground control station operator interface, completing a full real-time interactive loop.
[0039] (4) Flood inundation map and building damage analysis S441, Video Stream Data Reception and Processing: The data processing module receives the RTSP video stream from the drone gimbal, periodically captures key frame images, and performs preprocessing. S442, Semantic Segmentation and Classification: The large-scale multimodal model of the data processing module uses semantic segmentation algorithms to identify water areas in images, and at the same time uses target detection and classification algorithms to identify and classify the damage to buildings; S443. Flood Map Drawing and Damage Statistics: The data processing module automatically draws a flood flood map based on the segmentation results, generates a list of damaged buildings, and produces analysis results, including damage type, location, and severity.
[0040] S444. Results Integration and Visualization Output: All analysis results are automatically integrated into the map interface and displayed in a visual format, supporting real-time viewing and interaction by ground personnel.
[0041] Step 5: Report Generation and Mission Completion: When the above tasks are completed or more than half completed, the data processing module automatically generates a structured inspection report containing all findings of the current task based on the inspection situation, and sends the inspection report to the ground control station. Subsequently, after the UAV completes the inspection of all predetermined areas, it returns to the take-off and landing point and lands vertically according to instructions or automatically planned paths; the system generates a final inspection report and archives all raw data and reports. Based on the accurate and timely information provided by the system, the emergency command center quickly dispatches assault boats and rescue teams to the marked distress points and plans safe rescue routes.
[0042] As an example, the system constructed by this invention is applied to a complete post-flood inspection task process in a specific scenario: a river in a certain area breaches its banks due to continuous heavy rainfall, a town downstream is surrounded by floodwaters, some houses collapse, communications are interrupted, and there are reports of people missing.
[0043] Specifically as follows: (1) Task planning and deployment Operators at the ground control station plan key inspection areas, including flooded town centers, main roads, bridges, and several areas where people have reported missing persons. The system automatically generates optimized flight routes based on the area scope and priority.
[0044] The drone was assembled and its systems checked in a relatively open and safe area on the outskirts of the disaster area. After confirming that everything was in order, it took off vertically and climbed to the predetermined altitude.
[0045] (2) Autonomous inspection and real-time analysis The drone switched to fixed-wing cruise mode and began autonomous flight according to the planned route.
[0046] The multimodal sensor array begins operation, acquiring visible light video, infrared thermal imaging, LiDAR point clouds, and ambient sound in real time.
[0047] Data is transmitted in real time to the data processing module deployed in the cloud via the 5G / SatCom link. The large-scale multimodal model begins to process the data stream in real time and completes the identification and marking of building damage, detection and verification of trapped personnel signals, real-time voice interaction and target verification, flood inundation map and building damage analysis respectively.
[0048] During the process, the flight monitoring unit of the ground control station monitors the UAV's flight status in real time. Operators can observe the real-time video, thermal imaging, point cloud, and map overlay analysis results transmitted back by the UAV through the display unit. Ground control personnel can issue query commands to the system via voice (such as "Check if there are signs of life in area A" or "Display the latest LiDAR scan results of bridge B"). The MLLM understands the commands, retrieves the analysis results or controls the sensors, and generates a response through voice synthesis. The system can also proactively issue voice alarms or provide action suggestions to ground personnel based on emergency situations detected in real-time analysis (such as detecting distress signals).
[0049] (3) Report generation and task completion When the inspection task is more than halfway completed, the operator instructs the system to generate an interim report. Within minutes, the system automatically generates a structured report containing all current findings (flood map, damage list, bridge interruptions, location and evidence of high-priority distress signals) and sends it to the emergency command center.
[0050] After completing the inspection of all designated areas, the drone returns to the take-off and landing point according to instructions or an automatically planned path and performs a vertical landing.
[0051] The system generates a final inspection report and archives all raw data and reports. Based on the accurate and timely information provided by the system, the emergency command center quickly dispatches inflatable boats and rescue teams to the marked distress points and plans safe rescue routes.
[0052] The above description is only a preferred embodiment of the present invention and does not limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A full-modal unmanned aerial vehicle (UAV) intelligent patrol system for post-flood disasters based on a large model, comprising UAVs, characterized in that: It also includes an integrated multimodal sensor module, a data transmission and communication module, a data processing module, and a ground control station; The integrated multimodal sensor module is installed on the UAV and is used to collect multimodal data information from the disaster area synchronously or asynchronously. The data transmission and communication module is used to transmit multimodal data collected by the integrated multimodal sensor module to the data processing module in real time; The data processing module is used to receive the collected multimodal data, analyze and process the acquired multimodal data, and generate an inspection report based on the processing results. The ground control station includes a mission planning unit, which is used to plan the drone inspection route and mission parameters based on disaster information and inspection needs, and to deploy the drone to take off. Flight monitoring unit is used to monitor the flight status of the drone; The display unit is used to display the real-time video, thermal imaging, point cloud, and map overlay analysis results of the UAV; The report receiving and display unit is used to receive and display reports generated by the data processing module. The real-time voice interaction unit is used by ground control personnel to issue query commands to the system via voice and receive voice responses.
2. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The integrated multimodal sensor module includes a high-definition visible light camera, an infrared thermal imager, a lidar, and a microphone array.
3. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The ground control station also includes a voice recognition module, which is connected to the data processing module and is used to receive voice commands and transmit the received voice commands to the data processing module.
4. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The data processing unit further includes: It receives visible light images and LiDAR point cloud data synchronously collected by a multimodal sensor array carried by a drone, uses the TMRoPE position coding algorithm to perform feature projection and time alignment of the visible light images and LiDAR point clouds using a multimodal data fusion architecture, and performs feature fusion through a multi-head attention mechanism. Semantic understanding and structural analysis were performed on the fused data to identify bridge structural anomalies; Mark the bridge as "interrupted" on the map interface and generate a preliminary text alert, which is then pushed to the ground control station.
5. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The data processing unit further includes: It receives environmental sound data and thermal signal data synchronously collected by the microphone array and infrared thermal imager in the multimodal sensor array carried by the UAV, and transmits them to the data processing module in real time; The system performs sound source localization on audio data and estimates the coordinates of the signal source; simultaneously, it performs anomaly detection on thermal imaging data and identifies weak but persistent thermal signals. The TMRoPE algorithm was used to time-align audio and thermal imaging data, and the probability of life signs at each angle was calculated. Bayesian inference was then used to comprehensively determine the presence of life signs at the signal source. The target location is marked as high priority, an inspection report is generated, and the inspection report is sent to the ground control station via voice synthesis as an alarm.
6. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The data processing unit further includes: It receives voice commands from ground control station operators, performs natural language parsing on the received voice commands, and extracts the core intent. Based on the core intent, commands are issued to the drone through the flight control system to control the drone to adjust its flight attitude and gimbal angle, perform 360-degree circling flight around high-priority targets, and simultaneously adjust the focal length and focus of the high-definition visible light camera and infrared thermal imager to obtain clearer, multi-angle images. During the orbital flight, it receives real-time visible light and thermal imaging video streams collected by a multimodal sensor array, and continuously analyzes the real-time video streams to understand the content of the images; Based on the combined video analysis results, target geographical location information, and initial commands, a voice response is generated using speech synthesis technology. This voice response, along with the real-time video feed, is then pushed to the ground control station operator interface, completing a full real-time interactive loop.
7. The intelligent post-flood patrol system based on a large model and full-modal unmanned aerial vehicle (UAV) as described in claim 1, characterized in that: The data processing unit further includes: Receive RTSP video stream from the drone gimbal, periodically capture keyframe images, and perform preprocessing; Semantic segmentation algorithms are used to identify water areas in images, while target detection and classification algorithms are used to identify and classify the damage to buildings. The system automatically generates a flood inundation map based on the segmentation results, and produces a list of damaged buildings and analysis results, including damage type, location and severity. All analysis results are automatically integrated into the map interface and displayed in a visual format, supporting real-time viewing and interaction by ground personnel.
8. A method for intelligent post-flood patrol using a full-modal unmanned aerial vehicle (UAV) based on a large model, characterized in that: Includes the following steps: Step 1, Task Planning and Deployment: Based on disaster information and inspection needs, plan the inspection route and mission parameters of the UAV at the ground control station; deploy the UAV to the take-off and landing point and perform vertical take-off; Step 2, Autonomous Inspection and Data Acquisition: The UAV flies autonomously according to the preset route or real-time commands, and at the same time activates the multimodal sensor array to continuously collect visible light images / videos, infrared thermal imaging data, LiDAR point cloud data and environmental sound data of the disaster area; Step 3: Real-time data transmission and processing: The acquired multimodal data is processed in real-time or near real-time through a communication link by a data processing module. Step 4: Intelligent Analysis: The data processing module receives and fuses the received multimodal data streams based on a large-scale multimodal model, and executes the set analysis tasks; Step 5: Report Generation and Mission Completion: The data processing module automatically generates a structured inspection report containing all findings of the current mission based on the inspection results, and sends the inspection report to the ground control station; after the UAV completes the inspection of all predetermined areas, it returns to the take-off and landing point and lands vertically according to instructions or automatically planned paths; the system generates the final inspection report and archives all raw data and reports.