Disaster site analysis and decision method for unmanned aerial vehicle multi-modal perception data fusion
By fusing multimodal data from UAVs and using edge computing, an integrated three-dimensional situation map is generated, which solves the problems of fragmented information at disaster sites and reliance on experience for decision-making, and enables real-time, scientific and efficient decision support at disaster sites.
Patent Information
- Application Number
- CN202610995042.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-25
AI Technical Summary
At disaster sites, existing technologies suffer from problems such as fragmented perception dimensions, reliance on experience for decision-making, and a lack of edge intelligence leading to response delays. This results in non-real-time information processing, unscientific decision-making, and low efficiency.
By fusing multimodal perception data from UAVs, using edge computing devices for spatiotemporal benchmark alignment and cross-modal fusion analysis, an integrated three-dimensional situation map is generated. Based on artificial intelligence, optimal decision-making suggestions are generated, achieving automated processing from multidimensional perception to deep cognition.
It enhances the dimensionality and penetration of situational awareness at disaster sites, enables computational and optimized decision-making, strengthens the safety of rescue operations, and ensures the immediacy of emergency response.
Smart Images

Figure CN122637263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of emergency command and artificial intelligence technology, and more specifically to a method for using an unmanned aerial vehicle (UAV) platform to perform real-time fusion and intelligent analysis of multimodal sensing data such as visible light, infrared thermal imaging, and gas concentration at disaster sites to generate auxiliary decision-making information. Background Technology
[0002] In emergency response to sudden disasters such as fires, hazardous chemical leaks, and earthquakes, the core challenge for on-site commanders is to make rapid and accurate decisions under conditions of high information uncertainty, rapidly changing environment, and extremely tight time constraints.
[0003] Existing emergency command technology solutions have the following fundamental technical defects in practical applications: ① The fragmentation of perception and the cognitive gap: Current technologies, even in scenarios using drones, essentially present data from different sensors (e.g., visible light, infrared) as independent, heterogeneous information sources to the commander. Commanders must rely on their personal professional experience to subjectively integrate two-dimensional video footage, abstract temperature readings, and fragmented voice reports within a short period, associating them with an imprecise three-dimensional spatial concept. This process is not only cognitively demanding and inefficient, but also highly susceptible to misjudgments due to incorrect information association under emergency pressure.
[0004] ② Decision-making is experience-dependent rather than computationally driven: Due to the lack of effective tools for uniformly quantifying and calculating multi-dimensional on-site information, commanders' decisions (e.g., the selection of attack routes, the setting of assembly points) largely rely on personal experience and intuitive judgment. This decision-making model cannot guarantee its optimality under complex disaster situations, nor can it cope with emergencies beyond the scope of experience. For example, two seemingly similar rescue routes may have orders of magnitude differences in the comprehensive risks they conceal (e.g., temperature, toxic gas, structural instability), and the human brain cannot quickly and accurately quantify and compare these.
[0005] ③ The lack of edge intelligence leads to inherent delays in response: Disaster sites are typically areas with limited or disrupted communication. Technical solutions relying on public networks or back-end cloud computing centers for data analysis often fail in practice due to network outages or congestion. This results in the massive amounts of data collected by drones not being processed in real time, and consequently, not being able to be transformed into real-time situational information and decision-making support. This limits the use of drones to a single-function aerial sensing platform, causing them to miss critical decision-making windows measured in seconds during emergency response.
[0006] Therefore, there is an urgent need in this field for a new technological approach that can automatically complete the entire process from multi-dimensional perception to deep cognition and then to quantitative decision-making at the edge computing end of a disaster site, thereby improving the scientific nature and efficiency of emergency command. Summary of the Invention
[0007] To overcome the shortcomings of existing technologies, this invention discloses a disaster site analysis and decision-making method based on UAV multimodal perception data fusion. The purpose of this invention is to address technical problems in existing technologies, such as incomplete disaster site situational information acquisition, difficulty in effectively fusing multi-source heterogeneous data, low intelligence in decision support, and poor real-time response. This invention collects multimodal data streams from disaster sites via UAV flight, transmits these streams to edge computing devices at the disaster site for spatiotemporal benchmark alignment, performs cross-modal fusion analysis on the aligned multimodal data streams, identifies targets and generates hazardous area boundaries, and constructs an integrated three-dimensional situational map using the targets and hazardous area boundaries. Finally, it uses this integrated three-dimensional situational map to generate decision recommendations including optimal safe paths and optimal rescue force deployment points, and visualizes these recommendations. This invention can automatically complete the entire process from multi-dimensional perception to deep cognition and then to quantitative decision-making at the edge computing end of the disaster site, thereby improving the scientific nature and efficiency of emergency command.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A disaster site analysis and decision-making method based on UAV multimodal perception data fusion includes the following steps: I. Multimodal Data Acquisition and Spatiotemporal Alignment S100: The drone flies and collects multimodal data streams from the disaster site, and transmits the multimodal data streams to computing devices at the edge of the disaster site. The computing devices perform spatiotemporal reference alignment on the multimodal data streams. Preferably, step S100 includes: Concurrent data acquisition: Controlling the drone to fly, the drone uses multiple sensors mounted on it to concurrently acquire multimodal data streams from the disaster site, and transmits the multimodal data streams to computing devices at the edge of the disaster site; Spatiotemporal reference alignment: In the computing device, using a unified high-precision clock as a reference, the UAV pose data stream in the multimodal data stream and the pre-calibration parameters of each sensor are used to register all heterogeneous data streams in the multimodal data stream in time and space.
[0009] Preferably, in step S100, the multimodal data stream includes: visible light image data stream, infrared thermal imaging data stream, gas concentration data stream containing three-dimensional spatial coordinates and timestamps, and high-precision pose data stream of the UAV itself.
[0010] II. Cross-modal fusion analysis and dynamic 3D situation modeling The S200 and computing devices perform cross-modal fusion analysis on the aligned multimodal data streams, identify targets and generate hazardous area boundaries, and use the targets and hazardous area boundaries to construct an integrated three-dimensional situation map; 2.1 Fusion-enhanced target recognition Preferably, in step S200, identifying the target includes: The infrared thermal imaging data stream in the multimodal data stream is processed to identify high-temperature areas that are significantly higher than the ambient background temperature, which are then identified as potential target areas. Within the visible light image region corresponding to the potential target region in the multimodal data stream, an artificial intelligence target detection model is invoked to identify the target region in the image. A weighted fusion decision function based on dynamic weights is used to mark targets in potential target regions and image target regions.
[0011] Preferably, in step S200, the weighted fusion decision function based on dynamic weights for:
[0012] in, For the confidence level of the visual model, To normalize the infrared signal intensity, These are environmental quality parameters assessed in real time based on visible light images; Weighted fusion decision function according to The value is adjusted adaptively. and Contribution to the final result; in When the value is below the threshold, the function output is more biased towards The judgment result is the same as the previous one; otherwise, it is more inclined to be the opposite. The judgment result.
[0013] 2.2 Dynamic Hazard Area Generation Preferably, in step S200, generating the boundary of the hazardous area includes: Using a spatial interpolation algorithm, the continuous distribution field of gas concentration in the entire three-dimensional space is estimated based on the gas concentration data stream in the multimodal data stream. Based on a preset safety threshold, the continuous distribution field of the gas concentration is segmented, and one or more three-dimensional danger zone boundaries are automatically generated and updated in real time.
[0014] 2.3 Constructing an integrated three-dimensional situation map Preferably, in step S200, constructing the integrated three-dimensional situation map includes: Use the on-site lidar point cloud or oblique photogrammetry 3D model as the basic geographic scene. The identified targets and generated danger zone boundaries are overlaid as independent information layers onto the basic geographic scene to obtain the integrated three-dimensional situation map.
[0015] Preferably, in step S200, the integrated three-dimensional situation map is dynamically maintained using a hybrid update mechanism that combines periodic and event-driven updates: routine information is refreshed periodically at low frequency, while key information is updated in real time at high frequency based on its rate of change or whether an alarm is triggered.
[0016] III. Generation of AI-based Assisted Decision-Making Suggestions The S300 computing device uses an integrated 3D situation map to generate decision recommendations that include the optimal safe path and the optimal deployment points for rescue forces, and then visualizes these recommendations.
[0017] 3.1 Intelligent planning of rescue / evacuation routes Preferably, in step S300, generating the optimal safe path includes: Abstract the integrated 3D situation map into a weighted 3D map; Each edge in the weighted 3D graph is assigned a continuously varying passage cost based on fuzzy logic, which is calculated by a comprehensive cost function; The optimal path search algorithm is invoked to search for the path with the lowest total cost in the weighted 3D graph, which is taken as the optimal safe path; wherein the total cost is obtained by adding the travel costs of multiple edges on the path.
[0018] Preferably, in step S300, the passage cost is:
[0019] in, For the cost of passage, Let be the geometric length of the side. The first edge in the region The intensity of the risk factors For this nonlinear penalty function, As weight, The types of risk factors.
[0020] 3.2 Recommended Deployment Points for Rescue Forces Preferably, in step S300, generating the optimal deployment point for rescue forces includes: based on an integrated three-dimensional situation map, using a multi-objective optimization algorithm to comprehensively consider multiple decision factors, calculating and recommending the optimal deployment point for rescue forces; wherein, the multiple decision factors include upwind direction, safe distance, accessibility, and rescue efficiency.
[0021] 3.3 Visualization of Decision-Making Information Preferably, in step S300, the visualization presentation includes: overlaying the generated optimal safety path and optimal rescue force deployment points onto an integrated three-dimensional situation map in a visual manner, and presenting it to the commander through a user interface.
[0022] Preferably, in step S300, the user interface includes: The real-time video and data panel on the left: its upper half includes a visible light real-time video window and an infrared thermal imaging real-time video window, and its lower half is a key information panel, which displays key quantitative information in real time in the form of a list or dashboard. The integrated 3D situation map window on the right: This window contains a 3D disaster site model that users can rotate and zoom by touching or using a mouse. Targets on the 3D disaster site model are marked with eye-catching, standardized icons, and the danger zone is represented by semi-transparent 3D geometric shapes. The optimal safe path is drawn on the 3D disaster site model as a highlighted, flashing green dynamic line, clearly indicating the route forward. The optimal deployment point of rescue forces is marked on the ground with an icon containing a letter.
[0023] The beneficial effects of this invention are: 1. Enhancing the Dimensionality and Penetration of Situational Awareness: By fusing non-visible light information such as infrared and gas sensors, this method can penetrate physical obstructions such as smoke, dust, and darkness to perceive invisible chemical hazards. In particular, the dynamic weight-based fusion recognition algorithm allows the system to adaptively prioritize more reliable infrared signals under adverse visual conditions, ensuring robustness of perception and extending situational awareness from two-dimensional vision to a three-dimensional, multi-physical-quantity spatial environment, achieving in-depth understanding of disaster sites. This method can adaptively adjust the confidence weights of different sensor information according to real-time environmental changes to solve the problem of accurate target identification in complex and variable environments (e.g., smoke, dust, low light).
[0024] 2. Achieving Computational and Optimal Decision-Making: This method automatically processes complex, multi-source, and heterogeneous information into a structured three-dimensional situation map and actionable decision recommendations. Based on a fuzzy logic-based path cost function, the decision-making process is elevated from qualitative "obstacle avoidance" to quantitative "risk-benefit" trade-offs, and from simple "obstacle avoidance" to global optimization calculations that "maximize benefits and minimize harm," resulting in paths that are closer to reality and have greater operational value. This significantly reduces the cognitive load on commanders, ensuring the objectivity, scientific rigor, and optimality of decisions.
[0025] 3. Enhancing the safety of rescue operations: By dynamically generating hazardous area boundaries and planning optimal safe routes, this method provides rescue personnel with clear risk avoidance guidance. The continuously changing travel cost function allows route planning to be refined to avoid "moderate danger" rather than just "absolute danger," thereby seeking maximum efficiency while ensuring safety and providing technical support for protecting the lives of frontline personnel.
[0026] 4. Ensuring the Immediacy of Emergency Response: The core calculations of this method can all be completed on computing devices at the edge of the disaster site, without relying on backend cloud or public network communication. The hybrid model update mechanism, combining periodic and event-driven updates, ensures the overall timeliness of the situation map while enabling instantaneous responses to critical emergencies, achieving intelligent allocation of computing resources. This feature ensures that the system can still operate independently in extreme disaster scenarios where communication is interrupted, achieving "second-level" situational analysis and decision support. Attached Figure Description
[0027] Figure 1 This is an overall flowchart of the method of the present invention; Figure 2 This is a schematic diagram illustrating the principle of multimodal information fusion in this invention. Figure 3 This is a schematic diagram of the path planning algorithm of the present invention; Figure 4 This is a schematic diagram of the user interface for assisting decision-making in this invention. Detailed Implementation
[0028] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with the embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention.
[0029] A disaster site analysis and decision-making method based on UAV multimodal perception data fusion is proposed. This method is executed by computing devices deployed on the UAV platform or its ground control system (e.g., an onboard edge computing unit integrated into an all-terrain vehicle). The specific steps of this method are as follows: Figure 1 As shown, it includes: Step S100: Multimodal data acquisition and spatiotemporal alignment S110: Concurrent Data Acquisition: Controls the UAV platform to concurrently acquire multimodal data streams from the disaster site through multiple sensors mounted on it. The data streams include at least: visible light image data streams, infrared thermal imaging data streams, gas concentration data containing three-dimensional spatial coordinates and timestamps, and the UAV's own high-precision pose data stream.
[0030] S120: Spatiotemporal reference alignment: In the computing device, using a unified high-precision clock as a reference, and utilizing the UAV pose data and the pre-calibration parameters of each sensor, all heterogeneous data streams are precisely registered in time and space to ensure that in subsequent processing, any point in the three-dimensional space has multi-dimensional physical attributes (e.g., visual texture, temperature, gas concentration) that can be queried simultaneously.
[0031] Step S200: Cross-modal fusion analysis and dynamic 3D situation modeling This step is the core of the algorithm of this invention, aiming to transform the aligned raw data into structured, computable situational information. Specifically, as follows... Figure 2 As shown, it includes: S210: Fusion-enhanced target recognition: The infrared thermal imaging data stream is processed to identify high-temperature areas that are significantly higher than the ambient background temperature, which are then designated as "potential target areas".
[0032] Within the visible light image area corresponding to the "potential target area", an artificial intelligence target detection model is invoked for identification.
[0033] A weighted fusion decision function based on dynamic weights is used to determine the final target confidence level, as follows:
[0034] in, For the confidence level of the visual model, To normalize the infrared signal intensity, It is an environmental quality parameter (e.g., smoke concentration, image sharpness) that is evaluated in real time based on visible light images. Weighted fusion decision function according to The value is adjusted adaptively. and Contribution to the final result; in When the value is low, the function output is more biased towards The judgment result is the same as the previous one; otherwise, it is more inclined to be the opposite. The judgment result is determined by this adaptive fusion strategy, which aims to ensure the robustness and accuracy of target recognition in changing environments.
[0035] S220: Dynamic Hazard Area Generation: A spatial interpolation algorithm is used to estimate the continuous distribution field of gas concentration in the entire three-dimensional space based on discrete gas concentration data points.
[0036] Based on a preset safety threshold, the concentration field is segmented, and one or more three-dimensional "hazardous area boundaries" (i.e. hazardous gas envelopes) are automatically generated and updated in real time.
[0037] S230: Constructing an integrated three-dimensional situation map: Use the on-site lidar point cloud or oblique photogrammetry 3D model as the basic geographic scene.
[0038] The targets identified in step S210 and the boundaries of the danger zones generated in step S220 are used as independent information layers and precisely overlaid onto the basic geographic scene.
[0039] The situation map is dynamically maintained using a hybrid update mechanism that combines periodic and event-driven updates: routine information is refreshed periodically with low frequency, while critical information is updated frequently and in real-time based on its rate of change or whether an alarm is triggered. This mechanism aims to ensure the timeliness of the situation while optimizing the allocation of edge computing resources.
[0040] Step S300: Generation of AI-based Assisted Decision-Making Suggestions This step aims to translate situational awareness into actionable steps, specifically as follows: Figure 3 As shown, it includes: S310: Intelligent planning of rescue / evacuation routes: The integrated 3D situation map is abstracted into a weighted 3D graph.
[0041] Each edge in the graph is assigned a continuously varying "travel cost" based on fuzzy logic. The travel cost C (edge) is calculated using a comprehensive cost function, as follows:
[0042] in, For the cost of passage, Let be the geometric length of the side. The first edge in the region The intensity of each hazard factor (e.g., temperature, gas concentration). For this nonlinear penalty function, As weight, This design quantifies different levels and types of hazards into continuous travel costs, enabling route planning to quantify the trade-off between risk and efficiency across all potential routes.
[0043] Invoke an optimal path search algorithm (e.g., A* algorithm) to search for the path with the lowest total cost in the weighted 3D graph, which is designated as the "optimal safe path".
[0044] S320: Recommended Deployment Points for Rescue Forces: Based on an integrated 3D situation map, a multi-objective optimization algorithm is used to comprehensively consider multiple decision factors such as upwind direction, safe distance, accessibility, and rescue efficiency to calculate and recommend the optimal deployment points for rescue forces.
[0045] S330: Visualization of Decision-Making Information The generated "optimal safe path" and "suggested deployment point" are displayed visually overlaid on an integrated 3D situation map, and accessible through a user interface (such as...). Figure 4 (As shown) is presented to the commander.
[0046] This invention can be applied to emergency response to fires in high-rise buildings in cities. Specifically, this embodiment uses a fire in a high-rise residential building in a city as an example to illustrate the specific application of the method of this invention, as follows: S100 Phase: Upon receiving the instruction, the drone operation vehicle (edge computing unit carrier) equipped with the system of this invention quickly arrives at the safe area on site. After takeoff, the drone first performs a surround oblique photography flight, acquiring multi-angle high-definition images of the building within minutes, and the edge computing unit quickly generates a basic 3D scene model with realistic textures. Subsequently, the drone hovers at a key altitude and uses its zoom camera, infrared thermal imager, and (optional) air quality sensor to scan the building layer by layer and window by window, simultaneously collecting visible light, infrared, and gas / particulate matter concentration data streams, and performing real-time spatiotemporal alignment.
[0047] S200 phase: S210: The system analyzes the infrared data stream and quickly locates significant high-temperature anomalies on the 15th and 16th floors. Simultaneously, due to heavy smoke at the scene, the visible light image quality assessment parameter Q_env for the 15th-floor window is low. The system automatically increases the infrared signal weight β, so even though it is not visually clearly identifiable, based on its strong humanoid thermal signal, the system still marks it as "high-confidence trapped person 1". For the 16th floor, there is an open flame in its window, resulting in an extremely high P_thermal, and the system marks it as the "primary fire source".
[0048] S220 & S230: The system accurately labels the identified "Trapped Person 1" and "Main Fire Source" icons onto the corresponding window positions on the 15th and 16th floors of the basic 3D model. Simultaneously, based on the direction of smoke diffusion, it renders a dynamic smoke impact range on the model. The situation map, through an event-driven mechanism, updates in real-time the spread of the fire source and the movement of trapped personnel (if any).
[0049] S300 Phase: S310: The commander sets the rescue entrance as "1st Floor Lobby" and the target as "Trapped Person 1" on the interface. The system aligns the building's internal structural diagram (which can be preset or imported from BIM data) with the 3D model and abstracts it into a weighted diagram. In the diagram, the access cost for floors 16 and above is set to infinite, the access cost for floor 15 is set to high, and other floors are set to normal. The A* algorithm searches the diagram and ultimately generates an "optimal safe path" from the 1st Floor Lobby, through the relatively safe fire escape in Building B to the 14th Floor, then laterally into Building A, and finally up the stairs to the 15th Floor. This path avoids the main staircase in Building A, where the fire is most intense.
[0050] S320 & S330: The system comprehensively considers factors such as wind direction and fire truck operating radius, and recommends the optimal landing point for the aerial ladder truck and the deployment point for the ground rescue air cushion on the 3D model. All information, including the 3D model, target annotations, optimal path, and deployment point suggestions, is available in... Figure 4 The interface shown is presented in a unified manner.
[0051] This invention can be applied to the handling of hazardous chemical leaks in chemical industrial parks. Specifically, this embodiment takes a scenario of a toxic and flammable liquid leak from a storage tank in a chemical industrial park as an example.
[0052] Phase S100: The drone takes off from the upwind safety boundary, carrying a visible light camera, an infrared thermal imager, and gas sensors for specific chemicals (such as chlorine and methane). The drone first conducts a high-altitude wide-area reconnaissance of the accident area, then descends to a lower altitude and performs a detailed scan of the area around the leak point along a pre-set grid or zigzag path, collecting multimodal data and completing spatiotemporal alignment.
[0053] S200 phase: S210: The infrared thermal imager detected a low-temperature zone formed by the leaked liquid on the ground, while simultaneously detecting an abnormally high temperature at a nearby pump, posing a risk of secondary explosion. Visible light imaging clearly showed the extent and direction of the leak's spread. The system labeled these targets as "leak source," "low-temperature coverage area," and "secondary hazard point," respectively.
[0054] S220: Based on gas concentration readings collected by UAVs at different altitudes and locations, a spatial interpolation algorithm generates and updates a three-dimensional "high-concentration hazardous gas envelope" and a broader "low-concentration diffusion zone envelope" in real time on the edge computing unit.
[0055] S230: The system overlays all the aforementioned targets and the three-dimensional gas envelope onto the basic three-dimensional scene model of the park, forming a dynamic, integrated situational map. Model updates are event-driven; for example, when the wind direction changes or the leakage rate increases, the gas envelope will instantly refresh its shape and extent.
[0056] S300 Phase: S310: The commander sets the task as "closing the upstream valve of the leak source" and specifies the locations of the inlet and target valves. In the weighted map abstracting the park area, the system sets the passage cost within the gas envelope to extremely high or infinite. The A* algorithm then calculates an optimal safe approach path that is entirely upwind and bypasses all gas diffusion zones.
[0057] S320 & S330: The system recommends optimal upwind locations for emergency command posts, decontamination stations, and medical aid stations. Simultaneously, based on gas diffusion model predictions, the system dynamically delineates suggested downstream evacuation zones on the interface. All information is available on... Figure 4 Presented on a similar interface, it provides comprehensive decision support for on-site response and regional joint defense.
[0058] Figure 1 This is a flowchart of the overall process of the method of the present invention.
[0059] This diagram aims to provide a macroscopic view of the complete logical chain and core processing stages of the method of this invention.
[0060] The diagram structure adopts a standard top-down flowchart layout, consisting of three core processing modules (rectangles) and data flow directions (arrows).
[0061] Module 1: Multimodal Data Acquisition and Spatiotemporal Alignment (S100). This module is located at the top of the flowchart and is the starting point for all processing. It contains two sub-step boxes: S110: Concurrent Data Acquisition, which schematically draws the drone icon and its various onboard sensors (e.g., camera, thermal imager, gas sensor) to represent the diversity of data sources; S120: Spatiotemporal Reference Alignment, which schematically draws the time axis and three-dimensional coordinate system icons to represent that all input data are unified under the same spatiotemporal reference.
[0062] Module 2: Cross-modal fusion analysis and dynamic 3D situation modeling (S200). This module, located at the center of the flowchart, is the core technology of this invention. It receives aligned data from S100. This module contains three parallel sub-processes, respectively pointing to: S210: fusion-enhanced target identification; S220: dynamic hazard area generation; and S230: constructing an integrated 3D situation map. The outputs of these three sub-processes ultimately converge to reflect the information fusion process.
[0063] Module 3: AI-Based Assisted Decision-Making Recommendation Generation (S300). This module, located at the bottom of the flowchart, represents the value output of this invention. It receives the situation map from S200. This module contains two core functional blocks: S310: Intelligent Planning of Rescue / Evacuation Routes and S320: Recommendation of Rescue Force Deployment Points. The outputs of these two functional blocks are ultimately delivered to the user through the S330: Decision Information Visualization Module.
[0064] Figure 2 This is a schematic diagram illustrating the principle of multimodal information fusion.
[0065] This diagram aims to illustrate in detail how heterogeneous data is intelligently integrated into a unified 3D scene in step S200.
[0066] The diagram structure uses a layered, superimposed schematic diagram structure.
[0067] Bottom layer: Basic 3D scene model. At the bottom of the image is a semi-transparent, meshed 3D building or terrain model, labeled "Basic 3D Scene (Generated by LiDAR / Oblique Photogrammetry)," serving as the geographic substrate for all information.
[0068] Middle layer: Multimodal data streams and interpretation. In the upper part of the figure, there are three parallel "data stream" inputs with arrows, labeled "Visible light image stream", "Infrared thermal imaging stream" and "Gas sensor data stream".
[0069] The visible light and infrared data streams point to a rectangular box labeled "Dynamic Weighted Fusion Judgment Module". Inside this box, a small module labeled "Environmental Quality Assessment (e.g., smoke concentration)" outputs two weighting coefficients, α and β, indicated by dashed arrows, illustrating the logic of dynamically adjusting weights based on the environment. The final output of this module is a "Target Information Layer" labeled with icons such as "Fire Source" and "High-Confidence Personnel".
[0070] The gas sensor data stream points to a rectangle labeled "Spatial Interpolation and Threshold Segmentation Module," whose output is a semi-transparent, irregular three-dimensional geometry labeled "Hazardous Gas Envelope Layer."
[0071] Top layer: Information overlay. The vertically downward dashed arrow in the diagram indicates that the generated "target information layer" and "hazardous gas envelope layer" are precisely overlaid on the "basic 3D scene model" below, ultimately forming a highly information-rich, integrated 3D situation map.
[0072] Figure 3 This is a schematic diagram of a path planning algorithm.
[0073] This diagram aims to intuitively explain the core principle of the path planning algorithm in step S310 through a specific example.
[0074] The diagram's structure: A simplified disaster scene layout using a two-dimensional overhead view or a 2.5D oblique view.
[0075] Scene Elements: The image includes a starting point labeled "Rescue Entrance (Point A)" and an ending point labeled "Trapped Personnel (Point B)". Different types of danger zones are distributed throughout the scene: a dark red area labeled "High-Temperature Core Zone (Cost = ∞)"; surrounding it is a light red area labeled "High-Temperature Affected Zone (Cost = High)". Additionally, there is a dark yellow semi-transparent area labeled "High-Concentration Toxic Gas Zone (Cost = ∞)"; surrounding it is a light yellow semi-transparent area labeled "Low-Concentration Diffusion Zone (Cost = Medium)".
[0076] Path comparison: At least two paths should be drawn in the diagram: A "shortest geometric path" represented by a thin gray dashed line, which may pass directly through the dark red "high-temperature core region".
[0077] Another “optimal safe path”, represented by a thick green solid line, completely avoids all dark (cost ∞) areas but may partially cross the light yellow (cost medium) areas because it has a lower total weighted cost than another longer path (which can also be schematically drawn) that completely bypasses all danger zones.
[0078] Logical explanation: This diagram, through a visual comparison of paths, clearly illustrates that the path planning of this invention is not simply obstacle avoidance, but rather a "risk-benefit trade-off" model that quantifies different levels of danger into different passage costs, thereby finding the optimal solution that balances safety and efficiency.
[0079] Figure 4 This is a schematic diagram of a user interface for assisting decision-making.
[0080] This figure is intended to illustrate the final product form and human-computer interaction interface delivered to the user by the method of the present invention.
[0081] The diagram's structure simulates the display interface of a tablet computer or a large screen in a command center, employing a typical column layout.
[0082] Left side panel: Real-time video and data panel. The upper half of this area contains two side-by-side video windows, labeled "Real-time Visible Light Video" and "Real-time Infrared Thermal Imaging Video" respectively. The lower half is the "Key Information Panel," which displays key quantitative information in real time, such as "Detected Gas: CO, Concentration: 50ppm" and "Number of Targets Identified: 3," in a list or dashboard format.
[0083] The main area on the right: an integrated 3D situation map window. This is the core of the interface. This window contains a 3D disaster scene model that can be rotated and zoomed by the user via touch or mouse. On the model, prominent, standardized icons (e.g., flame icons, human figures) mark the locations of fire sources and personnel identified in step S210. Semi-transparent 3D geometries (such as the envelope generated in step S220) represent the range of hazardous gases. Most importantly, the "optimal safe path" calculated in step S310 is drawn on the 3D model as a highlighted, flashing green dynamic line, clearly indicating the route forward. The "suggested deployment points" recommended in step S320 are marked on the ground with a special icon bearing the letter "P".
[0084] Interactivity: The diagram illustrates the interactive effect of a detailed information box popping up after clicking the path or target point, demonstrating the system's usability. This diagram fully demonstrates how this invention transforms complex background algorithm analysis results into intuitive, interactive, and actionable decision support information for commanders.
[0085] The embodiments of the present invention have been described in detail above, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalents or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A disaster site analysis and decision-making method based on UAV multimodal perception data fusion, characterized in that, Includes the following steps: S100: The drone flies and collects multimodal data streams from the disaster site, and transmits the multimodal data streams to computing devices at the edge of the disaster site. The computing devices perform spatiotemporal reference alignment on the multimodal data streams. The S200 and computing devices perform cross-modal fusion analysis on the aligned multimodal data streams, identify targets and generate hazardous area boundaries, and use the targets and hazardous area boundaries to construct an integrated three-dimensional situation map; The S300 computing device uses an integrated 3D situation map to generate decision recommendations that include the optimal safe path and the optimal deployment points for rescue forces, and then visualizes these recommendations.
2. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, The S100 steps include: Concurrent data acquisition: Controlling the drone to fly, the drone uses multiple sensors mounted on it to concurrently acquire multimodal data streams from the disaster site, and transmits the multimodal data streams to computing devices at the edge of the disaster site; Spatiotemporal reference alignment: In the computing device, using a unified high-precision clock as a reference, the UAV pose data stream in the multimodal data stream and the pre-calibration parameters of each sensor are used to register all heterogeneous data streams in the multimodal data stream in time and space. In step S100, the multimodal data stream includes: visible light image data stream, infrared thermal imaging data stream, gas concentration data stream containing three-dimensional spatial coordinates and timestamps, and high-precision pose data stream of the UAV itself.
3. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S200, identifying the target includes: The infrared thermal imaging data stream in the multimodal data stream is processed to identify high-temperature areas that are significantly higher than the ambient background temperature, which are then identified as potential target areas. Within the visible light image region corresponding to the potential target region in the multimodal data stream, an artificial intelligence target detection model is invoked to identify the target region in the image. A weighted fusion decision function based on dynamic weights is used to mark targets in potential target regions and image target regions.
4. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 3, characterized in that, In step S200, the weighted fusion decision function based on dynamic weights for: in, For the confidence level of the visual model, To normalize the infrared signal intensity, These are environmental quality parameters assessed in real time based on visible light images; Weighted fusion decision function according to The value is adjusted adaptively. and Contribution to the final result; When the value is below the threshold, the function output is more biased towards The judgment result is the same as the previous one; otherwise, it is more inclined to be the opposite. The judgment result.
5. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S200, generating the boundary of the hazardous area includes: Using a spatial interpolation algorithm, the continuous distribution field of gas concentration in the entire three-dimensional space is estimated based on the gas concentration data stream in the multimodal data stream. Based on a preset safety threshold, the continuous distribution field of the gas concentration is segmented, and one or more three-dimensional danger zone boundaries are automatically generated and updated in real time.
6. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S200, the construction of the integrated three-dimensional situation map includes: Use the on-site lidar point cloud or oblique photogrammetry 3D model as the basic geographic scene. The identified targets and generated danger zone boundaries are overlaid as independent information layers onto the basic geographic scene to obtain the integrated three-dimensional situation map. In step S200, the integrated three-dimensional situation map is dynamically maintained using a hybrid update mechanism that combines periodic and event-driven updates: routine information is refreshed periodically at low frequency, while key information is updated in real time at high frequency based on its rate of change or whether an alarm is triggered.
7. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S300, generating the optimal safe path includes: Abstract the integrated 3D situation map into a weighted 3D map; Each edge in the weighted 3D graph is assigned a continuously varying passage cost based on fuzzy logic, which is calculated by a comprehensive cost function; The optimal path search algorithm is invoked to search for the path with the lowest total cost in the weighted 3D graph, which is taken as the optimal safe path; wherein the total cost is obtained by adding the travel costs of multiple edges on the path.
8. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 7, characterized in that, In step S300, the passage cost is: in, For the cost of passage, Let be the geometric length of the side. The first edge in the region The intensity of the risk factors For this risk factor, a nonlinear penalty function is used. As weight, The types of risk factors.
9. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S300, generating the optimal deployment point for rescue forces includes: based on an integrated three-dimensional situation map, using a multi-objective optimization algorithm to comprehensively consider multiple decision factors, calculating and recommending the optimal deployment point for rescue forces; wherein, the multiple decision factors include upwind direction, safe distance, accessibility, and rescue efficiency.
10. The disaster site analysis and decision-making method based on UAV multimodal perception data fusion as described in claim 1, characterized in that, In step S300, the visualization presentation includes: visually overlaying the generated optimal safety path and optimal rescue force deployment points onto an integrated three-dimensional situation map, and presenting it to the commander through a user interface; The user interface includes: The real-time video and data panel on the left: its upper half includes a visible light real-time video window and an infrared thermal imaging real-time video window, and its lower half is a key information panel, which displays key quantitative information in real time in the form of a list or dashboard. The integrated 3D situation map window on the right: This window contains a 3D disaster site model that users can rotate and zoom by touching or using a mouse. Targets on the 3D disaster site model are marked with eye-catching, standardized icons, and the danger zone is represented by semi-transparent 3D geometric shapes. The optimal safe path is drawn on the 3D disaster site model as a highlighted, flashing green dynamic line, clearly indicating the route forward. The optimal deployment point of rescue forces is marked on the ground with an icon containing a letter.