Emergency situation understanding system and method based on multi-modal large model and storage medium

Through multimodal large models and edge computing technology, the problems of the emergency management system in terms of single data source, insufficient intelligent decision-making capabilities, and difficulty in balancing privacy and real-time performance have been solved. Multi-dimensional situational awareness, real-time decision support, and cross-departmental collaboration have been achieved, thereby improving the efficiency and safety of emergency response.

CN120707067APending Publication Date: 2025-09-26THE FIRST RES INST OF TELECOMMTECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510786940.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing emergency management system has significant technical bottlenecks in terms of single data source, insufficient intelligent decision-making capabilities, difficulty in balancing privacy and real-time performance, and weak cross-domain collaboration capabilities. It is difficult to meet the needs of modern emergency scenarios in terms of multimodal data fusion, real-time and privacy security, and cross-departmental collaboration.

Method used

It adopts multimodal large models, edge computing architecture and reinforcement learning technology, through multimodal data collection, preprocessing, multimodal large model deployment, situation understanding and question-answer generation modules, combined with incremental learning and lightweight models, to achieve localized data processing and dynamic optimization, and support cross-domain collaboration.

Benefits of technology

It improves the comprehensiveness and accuracy of emergency situation understanding, enhances real-time response capabilities and privacy security, dynamically adapts to complex environments, provides intelligent decision-making support and cross-domain collaboration capabilities, reduces operation and maintenance costs, and improves system sustainability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707067A_ABST
    Figure CN120707067A_ABST
Patent Text Reader

Abstract

The invention provides an emergency situation understanding system based on a multi-modal large model. The emergency situation understanding system comprises a multi-modal data acquisition module, a data preprocessing module, a multi-modal large model deployment module, a situation understanding and question and answer generation module and a tool registration and management module. Through four core innovations of multi-modal fusion, edge intelligence, dynamic optimization and cross-domain collaboration, the bottleneck problems of a traditional emergency system in the aspects of data utilization, real-time performance, safety, intelligence and the like are solved, and the response speed, decision-making precision and collaboration efficiency of emergency management are remarkably improved; and efficient and reliable technical support is provided for the fields of public safety, disaster rescue, military emergency and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emergency management technology, and in particular to an emergency situation understanding system, method and storage medium based on a multimodal large model. Background Art

[0002] With society's growing demand for public safety and emergency management, quickly and accurately understanding emergency situations in complex environments has become a core challenge in improving emergency response efficiency. Traditional emergency management systems typically rely on a single data source (such as sensors or video surveillance) for situation analysis. However, in modern emergency scenarios, with diverse data sources, strict real-time requirements, and sensitive privacy and security, these systems are gradually exposing significant technical bottlenecks. The limitations of existing technologies are mainly reflected in the following aspects: 1) Single data source and fragmented information: Existing systems often make decisions based on single-modal data (e.g., sensor values ​​or images alone), making it difficult to fully capture the multi-dimensional characteristics of complex events such as fires, earthquakes, and terrorist attacks. For example, while sensor data can reflect changes in environmental parameters, it lacks the integration of visual information, resulting in blind spots in situational understanding. 2) Insufficient intelligent decision-making capabilities: Traditional systems often use rule engines or shallow models for data analysis, which cannot effectively handle the nonlinear correlations of multimodal data and lack the ability to predict future trends. For example, early warning mechanisms that rely on fixed threshold triggers are difficult to adapt to dynamically changing emergency scenarios. 3) It's difficult to balance privacy and real-time performance: Existing technologies often rely on centralized cloud-based processing, requiring sensitive data (such as surveillance video from military facilities) to be transmitted remotely. This not only results in high latency but also poses a risk of leakage. For example, high-definition video from a disaster site can be transmitted to the cloud with latency ranging from seconds to minutes, making it impossible to meet real-time response requirements. 4) Weak cross-disciplinary collaboration: Emergency management involves collaboration among multiple departments, including firefighting, medical care, and public security. However, the existing system lacks a unified tool integration and scheduling mechanism, resulting in inefficient resource allocation. For example, incompatible data interfaces across departments make rapid coordinated response difficult.

[0003] Although deep learning technology has provided new insights for multimodal data fusion in recent years, existing approaches still have significant flaws. For example, some studies use convolutional neural networks (CNNs) to fuse image and sensor data, but these models are typically deployed in the cloud, making them incapable of addressing real-time and privacy issues. Other approaches use static model parameters, making them difficult to adapt to dynamic environmental changes (such as sudden weather changes that cause fluctuations in data quality). Furthermore, existing technologies lack integration of key technologies such as incremental learning and edge computing, resulting in poor system flexibility and difficulty supporting long-term evolution.

[0004] To address these challenges, there is an urgent need for an emergency situation understanding system that can achieve efficient multimodal data integration, localized intelligent processing, dynamic self-optimization, and support cross-domain collaboration. By introducing a multimodal large model, edge computing architecture, and reinforcement learning technology, this invention aims to overcome the limitations of traditional systems and provide real-time, accurate, and secure decision support for emergency management, thereby significantly improving emergency response capabilities in complex environments. Summary of the Invention

[0005] The present invention proposes an emergency situation understanding system, method, and device based on a multimodal large model, which solves the problem of insufficient comprehensiveness and accuracy of situation understanding in the existing technology. The technical solution of the present invention is achieved as follows: An emergency situation understanding system based on a multimodal large model, characterized by comprising: A multimodal data acquisition module, which is used to acquire image data, video data, and sensor data through multiple devices and dynamically adjust the parameters of the acquisition devices based on reinforcement learning algorithms to optimize data quality; Data preprocessing module, used to perform denoising, format conversion and standardization on the collected multimodal data; A multimodal large model deployment module, which includes selecting a deep learning model that is suitable for multimodal data and deploying it on edge computing nodes for localized data processing; The situation understanding and question-answering generation module is used to generate real-time situation analysis reports through multimodal feature fusion and provide question-answering support based on a generative pre-trained model; Tool registration and management module, used to register and dynamically schedule various emergency response tools and API interfaces; The system realizes data localization processing through edge computing nodes, dynamically optimizes the model with incremental learning technology, and predicts future trends through LSTM networks.

[0006] As a further solution, the multimodal data acquisition module includes: Adaptive data acquisition unit, which uses reinforcement learning algorithms to optimize sensor frequency, camera resolution, and drone flight parameters in real time to adapt to environmental changes; The collaborative acquisition control unit is used to coordinate the acquisition timing and coverage of different devices to ensure the temporal and spatial consistency of multi-source data.

[0007] As a further solution, the data preprocessing module includes: Denoising unit, which uses a U-Net-based convolutional neural network to remove noise from image and video data; Data alignment unit, which achieves spatiotemporal alignment of multimodal data through timestamp synchronization and spatial coordinate mapping; Standardized units convert data from different sources into a unified format and value range.

[0008] As a further solution, in the multimodal large model deployment module, the model selection criteria include: Multimodal input compatibility for images, videos, and sensor data; Inference speed and resource utilization on edge computing nodes; Supports incremental learning and online parameter updates.

[0009] As a further solution, the situation understanding and question-answer generation module includes: Multimodal feature fusion unit, which performs cross-modal correlation between image features, video temporal features, and sensor numerical features; The question-answer generation unit, based on the BERT or GPT model, generates natural language answers based on real-time situation analysis results; The situation prediction unit analyzes historical and real-time data through the LSTM network and outputs the prediction results of future situations.

[0010] As a further solution, the tool registration and management module includes: Tool registration interface, supporting registration and permission configuration of user-defined tools and APIs; The dynamic scheduling engine automatically matches and calls the appropriate tools or APIs based on emergency task requirements, and executes after the input parameters are verified.

[0011] As a further solution, the edge computing nodes are deployed as follows: Distribute large multimodal models on edge servers or terminal devices close to data sources; Lightweight model compression technology is used to reduce computing resource consumption and ensure real-time processing capabilities.

[0012] As a further solution, the system's dynamic model optimization method includes: Incremental learning based on real-time data feedback to update model parameters to adapt to new scenarios; The model weights are adjusted through online learning algorithms to optimize the accuracy and robustness of situation analysis.

[0013] An emergency situation understanding method based on an emergency situation understanding system comprises the following steps: Step S1: Acquire field data through a multimodal data acquisition module and dynamically adjust device parameters; Step S2: pre-process the data and input it into the multimodal large model of the edge computing node; Step S3: Generate a real-time situation analysis report and provide decision support through the question-answer generation module; Step S4: Call the registration tool to perform emergency tasks and perform resource scheduling based on the LSTM prediction results.

[0014] A non-temporary storage medium stores a computer program, which, when executed by a processor, implements the steps of the emergency situation understanding method.

[0015] Compared with the existing technology, this solution has the following beneficial effects: (1) Improve the comprehensiveness and accuracy of situational understanding. By integrating multi-source heterogeneous data such as images, videos, and sensors, we can break through the limitations of traditional single data sources and achieve multi-dimensional perception and cross-validation of emergency scenarios. For example, in a fire scene, the fusion of sensor data (temperature, smoke concentration) and video data (flame shape, personnel location) can more accurately assess the fire spread trend and personnel evacuation path. By utilizing the deep learning capabilities of multimodal large models, we can explore the implicit associations between different data modalities (such as abnormal objects in images and abnormal fluctuations in sensor data), significantly reduce the misjudgment rate, and improve the reliability of situation analysis; (2) Enhance real-time response capabilities and privacy security. Deploy large multimodal models on edge nodes close to data sources to avoid remote data transmission to the cloud, reduce data processing delays to milliseconds, and meet the real-time decision-making needs of emergency scenarios. Localized data processing: Sensitive data (such as military facility monitoring and personal privacy information) does not need to leave the collection device or local server and is directly processed by edge nodes, greatly reducing the risk of data leakage and meeting the application requirements of high-security scenarios; (3) Dynamically adapt to complex environments and scene changes. Through reinforcement learning algorithms, the parameters of the acquisition equipment are optimized in real time (such as adjusting the flight altitude of the drone to avoid obstacles and dynamically improving the camera resolution to cope with low-light environments), ensuring the high quality and stability of data acquisition under different environmental conditions. The model can dynamically update parameters based on real-time data feedback (such as adding the ability to identify disaster types) without the need to retrain the entire model, significantly improving the system's adaptability to unknown scenarios; (4) Provide intelligent decision support and cross-domain collaboration capabilities. The question-answering function based on generative pre-trained models (such as GPT) can quickly respond to user queries (such as "Is the current safety exit unobstructed?"), and combined with the time series prediction capabilities of the LSTM network, it can provide early warning suggestions for future situations (such as predicting the spread path of floods) to assist decision makers in formulating forward-looking strategies. Through a unified tool registration and management module, it integrates cross-domain tools and APIs such as firefighting, medical care, and public security, supports task-driven dynamic scheduling (such as automatically calling drones to perform search and rescue missions), and realizes the rapid deployment of multi-department collaborative response; (5) Reduce operation and maintenance costs and improve system sustainability. Lightweight model compression technology: When deployed on edge nodes, model pruning and quantization technologies are used to reduce computing resource usage, allowing the system to run on low-power devices and reducing hardware investment costs. At the same time, the system's modular design and open interface support the flexible integration of new tools and new models, facilitating future functional expansion (such as adding an earthquake early warning module) and extending the technology lifecycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is a structural block diagram of an emergency situation understanding system based on a multimodal large model of the present invention.

[0018] Figure 2 This is a structural block diagram of an emergency situation understanding system for natural disaster rescue according to the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] Reference Figure 1 , this paper proposes an emergency situation understanding system based on a multimodal large model. The specific system method is as follows

[0021] 1. Multimodal data acquisition

[0022] 1.1 Data Acquisition

[0023] This invention uses a variety of devices (drones, sensors, video surveillance, etc.) to collect multi-source data on site. These data sources include image data, video data, sensor data, etc., and the collection equipment and methods are dynamically selected according to site needs. The data acquisition process is achieved through the collaborative work between devices to ensure that the collected data covers as many information sources as possible. The formula is: D raw ={D image , D video , D sensor}, Draw represents the original multimodal data set, Dimage represents the image data collected from the camera. Dvideo represents the video data obtained from the video surveillance equipment. Dsenso represents real-time data collected from various sensors (such as temperature, humidity, pressure and other environmental data).

[0024] 1.2 Data Preprocessing

[0025] In order to improve the quality and usability of data, a series of preprocessing is required for the collected multimodal data, including denoising, format conversion, data alignment, etc. Especially in the processing of image and video data, deep learning-based denoising algorithms (such as U-Net or other convolutional neural networks) are used to remove environmental noise, and data standardization and formatting steps are used to ensure data uniformity. The formula is: , D preprocessed =f prepprocess (D raw ), Dpreprocessed represents the preprocessed multimodal data. fprepprocess represents the data preprocessing function, which includes operations such as denoising, format conversion, and data standardization.

[0026] 1.3 Adaptive Data Collection

[0027] The system can adaptively adjust the configuration of the acquisition equipment (such as sampling frequency, resolution, etc.) according to different environmental conditions (such as weather changes, lighting intensity, data accuracy requirements, etc.), ensuring that the multimodal data collected in various environments has the best quality and accuracy. The system uses reinforcement learning algorithms to optimize the working parameters of the equipment in real time to maximize the effectiveness of the collected data. The formula is: ControlParameters=RL adapt (Env state ,Device type ), ControlParameters represents the control parameters that need to be adjusted (such as sensor frequency, camera exposure, etc.), RLadapt is a reinforcement learning model used to adaptively adjust device settings based on the state of the environment. Envstate represents the environmental state, such as light, weather, geographical location and other factors. Devicetype represents the device type, such as camera, drone, sensor, etc.

[0028] 2. Multimodal Large Model Deployment

[0029] 2.1 Model Selection

[0030] Select the most suitable multimodal large model (such as Qwen2-VL-7B) in the system, which can process inputs from different data sources and perform feature fusion to achieve comprehensive situation understanding. The criteria for model selection include the model's training data set, algorithm applicability, and response speed. The formula is: M=ModelSelect(D preprocessed ), M represents the optimal multimodal large model selected, ModelSelect is a model selection function that selects the most appropriate model based on the characteristics of the input data (such as the complexity and dimension of the data). Dpreprocessed is the preprocessed multimodal data.

[0031] 2.2 Model Deployment

[0032] The multimodal large model in this invention will be deployed to the edge computing node to ensure the privacy and controllability of the data. Through edge computing, the model can be processed near the source of the collected data, reducing data transmission delays and improving response speed. The formula is: , Represents a large multimodal model that has been deployed to edge nodes. Represents the deployment operation of the model, Represents edge computing nodes where the model will perform calculations and inferences.

[0033] 2.3 Dynamic Model Update and Optimization

[0034] This invention uses online learning and incremental learning technology to enable the model to self-optimize based on newly collected data after deployment, continuously improving its processing accuracy and adaptability. The model will dynamically adjust its internal parameters based on feedback from historical data and real-time data. The formula is: M updated =M current +ΔM, Mupdated represents the optimized model. Mcurrent represents the currently used model, ΔM represents the update parameter of incremental learning.

[0035] 3. Situation Understanding and Question and Answer Generation

[0036] 3.1 Situation Analysis

[0037] Based on a large multimodal model, the system can analyze the collected data in real time, identify key events and changes, and generate analysis reports on emergency situations. After the data input is processed, the model can form a comprehensive understanding of the situation through multimodal feature fusion. The formula is: , S represents the generated situation understanding result, which is usually a high-level information structure describing the current emergency situation. It is the processing function of the large model, which generates conclusions on situation understanding by processing and fusing multimodal data. It is the preprocessed multimodal data input.

[0038] 3.2 Question and Answer Generation

[0039] Based on the situation analysis results, the system can provide decision support through the question-answer generation mechanism. Generative pre-trained models (such as BERT or GPT) are used to generate relevant answers based on the input query and the current situation analysis results. The formula is: y=fQ&A(S,Q), y represents the generated answer, fQ&A is the question-answer generation function. S is the situation analysis result output by the multimodal large model, Q represents the user's query.

[0040] 3.3 Multi-level situation reasoning and prediction

[0041] In addition to real-time situation analysis, the system can also predict future situations by combining historical data with real-time data, providing forward-looking support for decision makers. The prediction model is based on time series models such as long short-term memory networks (LSTMs) and can handle long-term dependencies. The formula is: S forecast =LSTM(D historical ,D real-time ), Sforecas represents the predicted future situation, LSTM is a long short-term memory network, which is used to process time series data. Dhistorical represents historical data, usually including past trends, event data, etc. Dreal-time represents data collected in real time and is used to predict future trends in combination with historical data.

[0042] 4. Tool registration and management

[0043] 4.1 Tool Registration

[0044] The present invention provides a tool registration module that allows users to register various tools and APIs in the system and quickly call and manage tools through a unified management platform. The system can dynamically select and schedule appropriate tools according to task requirements. The formula is: T=ToolRegister(API list ), T represents a registered tool set, ToolRegister is the tool registration function, APIlist represents a list of tools and APIs.

[0045] 4.2 Tool Calling

[0046] In actual emergency response, the system will automatically select the appropriate tool and execute it according to the task requirements. The tool is called through the API interface, and the input parameters are verified and the corresponding functions are executed, and the processing results are output. The formula is: y=f tool (x), y represents the result of tool execution, ftool stands for tool execution function, x represents input parameters (such as sensor data, commands, etc.).

[0047] Example: For details, see Figure 2 This application takes an emergency situation understanding system in natural disaster rescue as an example to further illustrate: for natural disaster rescue scenarios (such as earthquakes, floods, and typhoons), 1. The system architecture is as follows: 1) Front-end data acquisition layer: drone swarms (equipped with high-definition cameras, infrared imagers, and environmental sensors), fixed ground monitoring points, and mobile sensing units to collect images, videos, and environmental data (temperature, humidity, water level, gas concentration, etc.) in real time; 2) Edge computing nodes: Deploy portable edge computing devices (such as the NVIDIA Jetson AGX Xavier platform) around the disaster area to perform data preprocessing, preliminary feature extraction, and local inference; 3) Central Intelligent Processing Platform: A large multimodal model (such as Qwen2-VL-7B) is deployed in the cloud to integrate various multimodal data to form a unified understanding of the disaster situation and present it to rescue commanders through the command center interface; 4) Reinforcement learning adaptive control system: manages the drone's flight path and dynamically adjusts sensor parameters to improve the efficiency and accuracy of data collection in key areas; 5) Decision-making support and question-answering system: Generates rescue priority recommendations based on the current disaster situation and responds to natural language questions from commanders.

[0048] 2. Data Preprocessing Algorithm

[0049] For natural disaster rescue scenarios, the data preprocessing module includes: 2.1 Image / Video Preprocessing Noise suppression: Use the U-Net denoising network to sharpen low-light and heavily noisy images. Super-resolution reconstruction: For long-distance aerial images, ESRGAN (Enhanced Super-ResolutionGAN) is applied to improve resolution details. Image / video key frame extraction: Use optical flow algorithms (such as PWC-Net) to filter image frames with the most information and reduce redundancy.

[0050] 2.2 Sensor Data Preprocessing Outlier detection and correction: Detect sensor data anomalies based on Isolation Forest. Time synchronization and spatial alignment: All data is synchronized based on GPS timestamps, and spatial data is aligned through geocoding (GeoCoding).

[0051] 2.3 Cross-modal alignment

[0052] Extract the time and location information from the image, match it with the sensor data, and form a preliminary cross-modal mapping relationship.

[0053] 3. Reinforcement Learning Module (RL Adaptive Acquisition Control)

[0054] In order to improve the data collection quality in key areas, the system introduces a reinforcement learning (RL) control mechanism.

[0055] 3.1 Adopted RL Framework

[0056] PPO (Proximal Policy Optimization) has high stability and fast convergence, making it suitable for actual deployment on edge devices.

[0057] 3.2 State Space Definition Current drone location (latitude, longitude, altitude), Current camera field of view, Image clarity rating (based on the quality of real-time captured images), Sensor data completeness score, Weather conditions (wind speed, visibility), Mission priority maps (e.g., known casualty concentration areas, flooded areas).

[0058] 3.3 Action Space Definition Adjust flight direction (north, east, south, west, ascend, descend), Adjust the flight speed (fast, medium, slow), Adjust the shooting parameters (exposure, resolution), Dynamically activate or deactivate some sensors.

[0059] 3.4 Reward Function Design

[0060] The reward function RR takes into account the quality of collected data, the importance of the collection area, and energy consumption: , in: DataQualityScore: Based on image clarity and sensor data stability score CoverageScore: Coverage score of high-risk disaster areas EnergyConsumption: Power consumption (negative) Weight , , Can be adjusted according to actual scenarios.

[0061] 4. Data fusion method for multimodal large models

[0062] For natural disaster relief scenarios, data fusion in multimodal large models (such as Qwen2-VL-7B) uses the following key technologies: 4.1 Attention Mechanism A self-attention mechanism is introduced to image, video frame, and sensor data features to dynamically adjust the importance weight of each modality feature. For example, when smoke images are detected, the weight of temperature and gas sensor data is increased.

[0063] 4.2 Cross-modal alignment

[0064] Use cross-modal Transformer to encode image, video, and sensor data features separately and align them in the middle layer. A contrastive learning strategy is introduced to bring feature vectors of different modalities but describing the same event closer together. For example: , Through this method, the model can learn the consistency between "severe damage in the image" and "strong sensor vibration" at the disaster scene.

[0065] 4.3 Multi-layer Fusion

[0066] The model employs mid-level feature fusion, fusing information from different modalities within each Transformer layer, rather than just at the input or output stages. This ensures fine-grained interaction of features at each layer, improving the accuracy and detail of disaster understanding.

[0067] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An emergency situation understanding system based on a multimodal large model, characterized by: include: A multimodal data acquisition module, which is used to acquire image data, video data, and sensor data through multiple devices and dynamically adjust the parameters of the acquisition devices based on reinforcement learning algorithms to optimize data quality; Data preprocessing module, used to perform denoising, format conversion and standardization on the collected multimodal data; A multimodal large model deployment module, which includes selecting a deep learning model that is suitable for multimodal data and deploying it on edge computing nodes for localized data processing; The situation understanding and question-answering generation module is used to generate real-time situation analysis reports through multimodal feature fusion and provide question-answering support based on a generative pre-trained model; Tool registration and management module, used to register and dynamically schedule various emergency response tools and API interfaces; The system realizes data localization processing through edge computing nodes, dynamically optimizes the model with incremental learning technology, and predicts future trends through LSTM networks.

2. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The multimodal data acquisition module includes: Adaptive data acquisition unit, which uses reinforcement learning algorithms to optimize sensor frequency, camera resolution, and drone flight parameters in real time to adapt to environmental changes; The collaborative acquisition control unit is used to coordinate the acquisition timing and coverage of different devices to ensure the temporal and spatial consistency of multi-source data.

3. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The data preprocessing module includes: Denoising unit, which uses a U-Net-based convolutional neural network to remove noise from image and video data; Data alignment unit, which achieves spatiotemporal alignment of multimodal data through timestamp synchronization and spatial coordinate mapping; Standardized units convert data from different sources into a unified format and value range.

4. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: In the multimodal large model deployment module, the model selection criteria include: Multimodal input compatibility for images, videos, and sensor data; Inference speed and resource utilization on edge computing nodes; Supports incremental learning and online parameter updates.

5. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The situation understanding and question-answer generation module includes: Multimodal feature fusion unit, which performs cross-modal correlation between image features, video temporal features, and sensor numerical features; The question-answer generation unit, based on the BERT or GPT model, generates natural language answers based on real-time situation analysis results; The situation prediction unit analyzes historical and real-time data through the LSTM network and outputs the prediction results of future situations.

6. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The tool registration and management module includes: Tool registration interface, supporting registration and permission configuration of user-defined tools and APIs; The dynamic scheduling engine automatically matches and calls the appropriate tools or APIs based on emergency task requirements, and executes after the input parameters are verified.

7. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The edge computing nodes are deployed as follows: Distribute large multimodal models on edge servers or terminal devices close to data sources; Lightweight model compression technology is used to reduce computing resource consumption and ensure real-time processing capabilities.

8. The emergency situation understanding system based on a multimodal large model according to claim 1, characterized in that: The dynamic model optimization method of the system includes: Incremental learning based on real-time data feedback to update model parameters to adapt to new scenarios; The model weights are adjusted through online learning algorithms to optimize the accuracy and robustness of situation analysis.

9. A method for understanding an emergency situation based on the emergency situation understanding system according to any one of claims 1 to 8, characterized in that: The following steps are involved: Step S1: Acquire field data through a multimodal data acquisition module and dynamically adjust device parameters; Step S2: pre-process the data and input it into the multimodal large model of the edge computing node; Step S3: Generate a real-time situation analysis report and provide decision support through the question-answer generation module; Step S4: Call the registration tool to perform emergency tasks and perform resource scheduling based on the LSTM prediction results.

10. A non-transitory storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the emergency situation understanding method according to claim 9 are implemented.

Citation Information

Cited By

  • Dynamic priority scheduling method for multi-disaster emergency coordination

    CN122022412A

  • Dynamic priority scheduling method for multi-hazard emergency response coordination

    CN122022412B