Intelligent fire-fighting management method and system based on multi-modal AI large model
By using multimodal large model and digital twin space technology, unified semantic recognition and adaptive optimization of response strategies for multimodal data have been achieved, which solves the shortcomings of fire identification and response in existing smart fire protection systems and improves fire response efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-13
AI Technical Summary
Existing smart fire protection systems lack a unified semantic understanding mechanism when processing multimodal data, resulting in insufficient accuracy and timeliness in fire identification, fragmented response mechanisms, and an inability to dynamically generate optimal response solutions, thus affecting fire response efficiency and personnel safety.
A multimodal large model is used for data fusion and semantic recognition. Combined with a digital twin spatial model, virtual fire situation mapping and response strategy simulation are performed to generate adaptive scheduling instructions, enabling multi-department collaborative response and closed-loop optimization.
It improved the accuracy of fire identification and the robustness of the system, enhanced the efficiency of fire response and inter-departmental collaboration, and ensured the timeliness and optimization of the response.
Smart Images

Figure CN121660464A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fire management technology, specifically a multimodal approach. A large-scale intelligent fire management method and system. Background Technology
[0002] In the field of modern fire management, the rapid pace of urbanization and increasing building density have made the effective prevention and management of fire risks a crucial task. Traditional fire management methods primarily rely on manual inspections and basic fire protection facilities such as smoke detectors and sprinkler systems. While these methods can address fire incidents to some extent, they have significant shortcomings in risk prediction, rapid response, and optimal resource allocation.
[0003] For example, Chinese invention application CN117910811B discloses a multimodal-based... A large-scale intelligent fire management method and system. The method is based on multimodal... The intelligent fire management method based on the large model includes: collecting first data related to a first fire monitoring area and second data related to a second fire monitoring area; inputting the first data into a fire risk assessment model to obtain a fire risk prediction value; and correcting the second data based on the fire risk prediction value to obtain corrected second data. This invention significantly improves the intelligence level of fire safety monitoring and management by integrating AI models and big data processing technology, and effectively enhances the technological content and future adaptability of fire safety management.
[0004] For example, Chinese invention application CN119992805A discloses a smart fire protection remote management system, including a server and a user terminal. The server includes a multi-dimensional data acquisition module, a judgment module, a processing module, a warning level determination module, an alarm push module, and a fire extinguishing control module. The processing module is used to determine the fire hazard level of the detected area at the current moment based on the real-time collected temperature data, smoke data, and flame data of the detected area when the judgment result indicates that a fire has occurred in the detected area at the current moment. The warning level determination module is used to determine the fire warning level of the detected area at the current moment.
[0005] The shortcomings of the above-mentioned patents:
[0006] Current smart fire protection systems typically connect to data sources from cameras, temperature and humidity sensors, smoke detectors, and combustible gas monitoring devices. These data sources vary in type and modality, and the systems process them independently, lacking a unified semantic understanding mechanism. When one type of data is interfered with or becomes invalid, the system cannot make a reasonable judgment by combining it with other data, which can easily lead to false alarms or missed fires, affecting the accuracy and timeliness of fire identification.
[0007] Existing intelligent fire protection systems, after identifying potential fires, rely heavily on fixed rules or manual judgment for subsequent responses, failing to dynamically generate response plans based on actual conditions. This is particularly problematic when multiple departments, including fire protection, security, and property management, are collaborating. Inconsistent information flow and fragmented response mechanisms can lead to delays or errors that cause missed opportunities for optimal action, impacting fire response efficiency and personnel safety.
[0008] Therefore, this invention proposes a method based on multimodal... A large-scale intelligent fire management method and system is proposed to address the aforementioned issues. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention provides a multimodal-based... A large-scale intelligent fire management method and system are proposed to address the problems mentioned in the background.
[0010] To achieve the above objectives, the present invention provides the following technical solution: a multimodal-based... Large-scale intelligent fire management methods include:
[0011] Step 1: Collect multimodal perception data from the field environment and transmit the multimodal perception data as input data to the multimodal semantic recognition module;
[0012] Step 2: Use the multimodal semantic recognition module to perform modal parsing and semantic fusion on the input data to generate a unified semantic description containing fire status elements;
[0013] Step 3: Input the generated unified semantic description into the dynamic risk judgment module, and combine it with the current time information, spatial location parameters of the collection point and population density data to output the fire risk level that reflects the on-site risk status.
[0014] Step 4: Based on the obtained fire risk level, call the pre-built digital twin space model to realize the synchronous mapping of the fire status in the virtual space;
[0015] Step 5: Perform virtual drills of various response strategies in the constructed digital twin space model, and generate evaluation results of the response effects of each strategy.
[0016] Step 6: Select the optimal response strategy based on the evaluation results, and generate scheduling instructions through the adaptive linkage instruction generation module;
[0017] Step 7: Send the scheduling instructions to the target execution terminal device and management system to achieve automated response actions;
[0018] Step 8: Collect feedback data after the response action is executed and input it into the system for analysis. Based on the feedback results, dynamically optimize the subsequent response strategy to achieve closed-loop control.
[0019] Preferably, step 1 further includes:
[0020] Sub-step 1.1 involves constructing image acquisition channels, temperature acquisition channels, gas concentration acquisition channels, and audio acquisition channels by deploying multiple types of sensor nodes within the target building or location to acquire video frame sequences. Temperature sequence Gas concentration sequence and sound amplitude sequence ,in:
[0021] For time sampling points, for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time;
[0022] The sampling frequency of each channel must meet the condition of uniform timestamp alignment:
[0023] ,
[0024] in, , Maximum tolerance;
[0025] If the data acquisition delay of any channel exceeds the tolerance range, the synchronous re-acquisition mechanism will be triggered;
[0026] Sub-step 1.2 involves performing validity and continuity checks on the multimodal data sequence obtained in sub-step 1.1, and determining whether it meets the basic conditions for being sent to the multimodal semantic recognition module based on the following rules:
[0027] Image frame sharpness Greater than the set threshold Image sharpness can be defined using the Laplacian operator variance method:
[0028] ,in, For the Laplace operator, It is the variance function;
[0029] The change in the maximum value of three consecutive frames in the temperature sequence satisfies:
[0030] ,in, The threshold for temperature abrupt change. for Temperature value at any given time;
[0031] The rate of change in gas concentration is greater than the upper limit of background fluctuation:
[0032] ,in, The threshold for the rate of change in gas concentration. It is the unit of time differentiation;
[0033] The proportion of spectral energy in the sound signal within the alarm frequency band is greater than a preset ratio:
[0034] ,
[0035] in, , This is the frequency band for fire alarm sounds. , For the recording system frequency band, For sound spectral power density, The threshold for the ratio of audio signal energy. It is the unit of frequency differentiation;
[0036] When all four conditions above are met, a valid input window is formed. :
[0037] ;
[0038] Sub-step 1.3 involves processing the valid multimodal data filtered through sub-step 1.2. According to spatial location information Reorganize to construct a unified modal input tensor. The structure is as follows:
[0039] ,
[0040] in, The feature tensor after image modality encoding, , , Numerical encoding for temperature, gas, and audio modes. This is the spatial location information vector corresponding to the sensor;
[0041] Then tensor The input is fed into the encoding network of the multimodal semantic recognition module, preparing it for the semantic fusion process in step 2.
[0042] Preferably, step 2 further includes:
[0043] Sub-step 2.1: The input tensor constructed in step 1... The input is fed into a modality-specific encoding network to generate modality feature representation vectors. , , , ,Right now:
[0044] ,
[0045] in, For visual modal encoders, It is a temperature mode encoder. It is a gas mode encoder. For audio modal encoders,
[0046] for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time;
[0047] The output feature vectors are all mapped to a uniform dimensional space;
[0048] Sub-step 2.2 involves performing relevance matching on the modal features obtained in sub-step 2.1, and constructing a weighted fusion vector based on the following fusion weight calculation strategy. :
[0049] ,
[0050] in, , , , The fusion weight coefficients for visual, temperature, gas, and audio modalities are listed in order.
[0051] The weighting coefficients are determined through a soft attention mechanism:
[0052] ,
[0053] in, For learnable weight vectors, It is the transpose of the weight vector. , The feature vector of the mode;
[0054] If any of the fusion weights , To minimize the effective fusion weight threshold, the system determines that the modality information has a high noise ratio and performs feature suppression operation to retain the main modality information for the next step;
[0055] Sub-step 2.3 involves fusing the feature vectors. The input is fed into the semantic generation module to generate a structured semantic description vector. ,in:
[0056] The results of fire source location. According to the fire development level, This is an abnormal phenomenon type;
[0057] Among them, the fire development level The judgment is made based on the following logic:
[0058] ,
[0059] in, This is the upper limit of normal temperature. The threshold for the rate of temperature rise. For image clarity, The amplitude value of the sound. This is the critical value of the temperature gradient. This represents the critical lower limit for image sharpness. This is the minimum loudness threshold for alarm sounds;
[0060] Each level forms a fire severity decision chain based on a joint threshold logic;
[0061] The final output semantic description vector This will be used as input for step 3, participating in the further calculation of the on-site fire risk level.
[0062] Preferably, step 3 further includes:
[0063] Sub-step 3.1: The unified semantic description vector generated in step 2... As input, extract the location of the fire source. Fire development level and types of phenomena At the same time, combined with the spatial location vector obtained in step 1 and current time information Integrate and construct scene analysis input vectors The structure is as follows:
[0064] ,
[0065] in, The coordinates of the fire source location. Fire intensity level This is a label for the phenomenon type. The spatial coordinates of the sampling point This is the current timestamp. Human flow density;
[0066] vector This provides input data for the next sub-step, risk factor calculation.
[0067] Sub-step 3.2 involves processing the input vector obtained in sub-step 3.1. Input the risk factor calculation module to calculate the following risk factors in sequence: Fire risk factor Population exposure risk factors Spatiotemporal coupling risk factors ,in:
[0068] Fire risk factors Defined as:
[0069] ,
[0070] in, Assigning weights based on fire intensity level Weights for phenomenon types. A unique code for the phenomenon label;
[0071] Population exposure risk factors Defined as:
[0072] ,
[0073] in, For population density risk coefficient, For human traffic density, The spatial decay factor, The distance between the fire source and the sampling point is expressed as the Euclidean distance.
[0074] Spatiotemporal coupling risk factor Defined as:
[0075] ,
[0076] in, For spatiotemporal fluctuation coefficient, For time frequency, This is the current timestamp;
[0077] like and The situation was determined to be high-risk for fire, and the next step of the risk level assessment process was initiated.
[0078] Sub-step 3.3: Calculate the result obtained in sub-step 3.2. , , The judgment result is input into the risk level mapping function, and the fire risk level is output. The judgment logic is as follows:
[0079] ,
[0080] in, The fire risk threshold For the population exposure risk threshold, This is the threshold for spatiotemporal coupling risk.
[0081] Ultimately, the fire risk level will be determined. As input to step 4, it is used to invoke the digital twin space model and trigger the corresponding response strategy.
[0082] Preferably, step 4 further includes:
[0083] Sub-step 4.1: Based on the fire risk level generated in step 3. and the spatial location of the fire source Select the corresponding building instance from the pre-built digital twin model database. And load the current scene;
[0084] Initialize the fire situation field tensor It is used to simulate the thermal diffusion state of a fire source and is defined as follows:
[0085] ,
[0086] in, For the spatial coordinates of the fire source, This is the ignition source intensity coefficient. The heat decay coefficient, This refers to the current time point;
[0087] If no instance matching the fire source location can be found in the model database. The system stops synchronization mapping and generates a fault indication event for feedback analysis in step 8;
[0088] Sub-step 4.2, based on the multimodal raw data tensors obtained in steps 1 and 2... and fused feature vectors The temperature field, concentration field, image thermal area, and sound source direction are sequentially mapped onto the virtual space scene to construct a multimodal visualization field. ;
[0089] The definition is as follows:
[0090] Temperature distribution field :
[0091] ,
[0092] Combustible gas concentration field :
[0093] ,
[0094] Image thermal projection map Image mode encoder The output features are obtained by reverse mapping.
[0095] Sound source direction cone Calculated by inversion of spatial directivity in audio features:
[0096] ,
[0097] in, The angle of the main direction of the sound source. For audio signals at frequencies Power spectral density on;
[0098] All mapping fields together form a three-dimensional multimodal visual space. ;
[0099] If any mode is missing during the mapping process, automatic interpolation is performed at that location. The interpolation algorithm uses three-dimensional Gaussian kernel interpolation.
[0100] ,
[0101] in, For three-dimensional space points Interpolated modal values at the location, The spatial smoothing coefficient is... Original observation point Modal data values at the location, For the first Spatial coordinates of the original modal sampling points, It is an exponential function;
[0102] Sub-step 4.3 involves constructing the multimodal field of view in sub-step 4.2. Fire situation tensor The digital twin dynamic driving engine is jointly input to update the virtual fire situation status at the current moment;
[0103] Define the virtual fire scene state frame for:
[0104] ,
[0105] in, For twin state generation functions, For twin space structure model, Given the current state of fire heat spread, This is a modal data field.
[0106] status frame After outputting the results, proceed to step 5, the virtual simulation phase of the response strategy.
[0107] If the state frame update fails, the system records the current tensor and configuration state, and proceeds to the feedback analysis and processing module in step 8 for model correction.
[0108] Preferably, step 5 further includes:
[0109] Sub-step 5.1, based on the virtual fire status frame generated in step 4 and fire risk level The system selects a set of multi-response strategies. Conduct virtual drills;
[0110] Response strategies All contain the set of operation instructions to be executed. According to the location of the fire source and risk level Choose the most suitable set of strategies Input into the digital twin engine for simulation;
[0111] During this phase, all response strategies will proceed according to the predetermined simulation timeframe. The simulation is conducted internally, and the simulation process is based on the time schedule. Update fire status and operational feedback;
[0112] Sub-step 5.2, each response strategy When executed in a digital twin environment, the following performance metrics are recorded in real time:
[0113] Fire extinguishing time : Representation strategy The time required for the fire to be completely extinguished after execution is calculated using the following formula:
[0114] ,
[0115] in, The initial time when the response strategy begins to be executed. Time for extinguishing the fire This represents the thermal diffusion state of the ignition source.
[0116] Personnel evacuation time : Representation strategy The formula for calculating the time required for the evacuation of personnel involved in the incident is as follows:
[0117] ,
[0118] in, In order to evacuate the number of people, For personnel evacuation distance For strategy The average evacuation speed of people being evacuated in the middle of the evacuation process;
[0119] Device response time : Representation strategy The response time of the device is calculated using the following formula:
[0120] ,
[0121] in, For the number of devices involved, For equipment Distance for mission execution For device response speed;
[0122] Resource consumption : Representation strategy The resource consumption required for execution is calculated using the following formula:
[0123] ,
[0124] in, The number of resource types required. For resources Consumption coefficient, For strategy China's resources Demand;
[0125] Sub-step 5.3: Based on the various effect indicators recorded in sub-step 5.2, construct a comprehensive effect evaluation index for the response strategy. This is used to compare the merits of different strategies, and the evaluation formula is:
[0126] ,
[0127] in, , , , The corresponding weighting coefficients for each indicator, , In order of priority, these are fire extinguishing time, evacuation time, equipment response time, and resource consumption.
[0128] Based on the evaluation results The system selects the response strategy with the best overall effect. Proceed to the next step: generating scheduling instructions.
[0129] Preferably, step 6 further includes:
[0130] Sub-step 6.1: Based on the response strategy evaluation index calculated in step 5.3. Select the strategy with the best overall effect. It meets the following conditions:
[0131] ,
[0132] in, For the set of candidate response strategies, For the first The overall evaluation value of the strategy, The optimal response strategy;
[0133] If multiple strategies exist to satisfy Further compare the priority of key indicators in the strategy and judge them in order:
[0134] Prioritize the time for fire extinguishing The shortest;
[0135] If they are still the same, compare the evacuation times. ;
[0136] If you still cannot distinguish, select resource consumption. The lowest;
[0137] Sub-step 6.2: Based on the optimal response strategy selected in sub-step 6.1 and its operation instruction set Call the adaptive linkage instruction generation function Generate a scheduling instruction set for multiple execution terminals. ;
[0138] The logic for generating scheduling instructions is defined as follows:
[0139] ,
[0140] in, The first in the optimal response strategy Item operation, For the first The current physical location of the execution terminal. To determine the current load status of the execution terminal, This is the current system timestamp. Fire risk level, For scheduling instructions;
[0141] Simultaneously, an execution feasibility judgment function is introduced. If an instruction does not meet one of the following conditions, the scheduling instruction will be rejected:
[0142] The current load must not exceed the maximum load capacity.
[0143] The distance from the fire source must not exceed the executable range;
[0144] Command response time must be within the allowed timeframe;
[0145] in, This is the maximum load threshold for the terminal. It is a spatial distance function. The maximum executable space distance threshold. To estimate command transmission and response latency, This refers to the latest execution time of the operation instruction;
[0146] Sub-step 6.3 involves processing the scheduling instruction set generated in sub-step 6.2. Group by terminal device category and construct control distribution table ,each It contains all the instructions that similar devices need to execute;
[0147] Define the scheduling control logic for each group as follows:
[0148] ,
[0149] in, For scheduling instructions The type code of the device being pointed to. To distribute to the first The instruction set for this type of device;
[0150] Ultimately through the control interface module Push the scheduling control table to the terminal device system and upper-level management platform To ensure the automatic execution of response actions:
[0151] ,
[0152] If the device reports a status of instruction rejection, the system will transfer the instruction to the alternative execution device.
[0153] Preferably, step 7, sending the scheduling instruction to the target execution terminal device and management system to achieve automated response actions, further includes:
[0154] Sub-step 7.1, based on the scheduling control table constructed in step 6.3 Call the control interface module Send corresponding instruction sets to various terminal devices;
[0155] Define the success rate of sending each command as follows:
[0156] ,
[0157] in, For scheduling instructions, For the initial transmission success rate of the interface, For transmission attenuation coefficient, Send the path distance for the command;
[0158] like ,in, To minimize the acceptable probability of successful communication, the transmission is aborted, the command is marked as needing retransmission, and added to the retransmission queue. ;
[0159] Sub-step 7.2: All successfully issued instructions Upon reaching the target device, the terminal device activates the automatic response logic module according to the instructions. Execute the specified action, and define the probability of the terminal completing the execution as follows:
[0160] ,
[0161] in, This is the terminal baseline execution capability coefficient. This indicates the current load status of the device. This is the maximum load that the equipment can bear. The device status availability function is defined as follows:
[0162] ,
[0163] like The instruction is transferred to the backup device and the device abnormality information is recorded and transmitted to the feedback analysis module in step 8.
[0164] The actions performed include physical operations and logical control, and a response status code is actively returned after the operation. The possible values are as follows:
[0165] Response successful;
[0166] Response failed;
[0167] The terminal refused to execute.
[0168] Terminal disconnected;
[0169] Sub-step 7.3: After the device performs the action, it calls the status reporting module. , will include equipment Operation type, execution timestamp Status codes Execution feedback data Execution status data packet Report to the upper management platform and control center ;
[0170] Define state synchronization delay for:
[0171] ,
[0172] in, The timestamp of the W2 device completing the operation. The time when the W2 management platform receives the status packet;
[0173] like If the delay exceeds the maximum tolerance value for state synchronization, the state is marked as out of sync, triggering the synchronization retransmission logic. ;
[0174] After all received execution statuses are summarized, the management platform will write them into the response log, which will serve as input data for feedback data analysis in subsequent step 8.
[0175] Preferably, step 8 further includes:
[0176] Sub-step 8.1: Based on the device execution status data packet received by the management platform in step 7.3. Data collection includes terminal devices Response status codes Feedback data Execution timestamp Synchronization delay The original feedback data set, including ;
[0177] The feedback data is standardized to construct analysis vectors in a unified format. :
[0178] ,
[0179] in, , To provide the mean and standard deviation of the feedback values, , The mean and standard deviation of the synchronization delay. In response to the status code, This is the normalized value of the equipment load. This is the maximum load that the equipment can withstand.
[0180] like If any standardized indicator in any dimension exceeds the preset anomaly range, the data is marked as an anomaly and added to the anomaly feedback set. ;
[0181] Sub-step 8.2 involves standardizing the feedback data vector set from step 8.1. Perform cluster analysis and multidimensional evaluation, and define a scoring function for the effectiveness of each strategy in the current fire instance. :
[0182] ,
[0183] in, For the success rate,
[0184] This represents the average load percentage. For average synchronization delay, For strategy execution evaluation weighting coefficients, For indicator functions;
[0185] like ,in, The lower limit of acceptable performance is used to trigger the policy adjustment logic.
[0186] Sub-step 8.3 addresses the response strategies identified as inefficient in step 8.2. According to the abnormal feedback set Based on the error distribution characteristics, optimize the response strategy parameters and generate a new set of candidate strategies. ;
[0187] The specific adjustment methods are as follows:
[0188] If device response failures are concentrated in a specific type of terminal, reduce the call weight of that type of device in the policy;
[0189] For situations where the failure rate is high in high-load areas, increase the resource reservation factor for devices in that area;
[0190] If the synchronization delay is severe, terminal equipment that is geographically closer to the control center should be given priority.
[0191] Define the strategy parameter adjustment function It acts on the original policy parameter vector. Generate new parameters :
[0192] ,
[0193] in, For learning rate, The loss function is constructed based on abnormal feedback samples. The gradient of the policy parameters with respect to the loss function;
[0194] After optimization, a new set of candidate strategies will be created. The results are then resubmitted to the virtual exercise module in step 5.1 for effect evaluation, thereby achieving cyclical self-optimization of the strategy.
[0195] A multimodal-based A large-scale intelligent fire management system, comprising:
[0196] The multimodal semantic recognition module is used to fuse and process video, temperature, gas concentration and voice information from the scene and output a unified fire description;
[0197] The dynamic risk assessment module is used to generate fire risk scores based on semantic descriptions and environmental context factors.
[0198] The digital twin response module is used to build a virtual space model and perform response strategy rehearsals.
[0199] The adaptive linkage instruction generation module is used to generate scheduling instructions based on the exercise results and control the relevant equipment or systems to execute them;
[0200] The feedback analysis module is used to collect response result data and dynamically optimize subsequent response strategies.
[0201] This invention provides a multimodal-based A large-scale intelligent fire management method and system. It offers the following benefits:
[0202] 1. This invention employs a multimodal sensing data semantic fusion mechanism and multimodal... The large-model-driven fire identification framework constructs a unified semantic description of the fire status by inputting data from different modalities such as cameras, temperature and humidity sensors, smoke detectors, and combustible gas monitoring devices into a unified semantic recognition module. This achieves the technical effect of multi-source data complementarity and semantic-level collaborative identification of the fire status. Compared with existing technologies that process various sensor data independently and lack semantic fusion capabilities, this framework solves the problem that the system cannot comprehensively judge the fire status when a certain type of sensor data is interfered with or fails, thereby improving the accuracy of fire identification and the robustness of the system.
[0203] 2. This invention adopts an adaptive response decision-making mechanism based on strategy evaluation results and a multi-departmental collaborative scheduling instruction generation system. By performing multi-strategy response simulation in a digital twin space, and dynamically adjusting the response strategy based on a feedback closed-loop optimization mechanism, it achieves the technical effect of dynamically selecting the optimal response strategy according to the fire risk level and realizing collaborative execution between multiple terminal devices and the upper-level platform. Compared with the existing technical solutions based on fixed response rules or manual judgment, this invention solves the problems of rigid response schemes, dispersed linkage mechanisms, and delayed responses in complex fire scenarios, thereby significantly improving the efficiency of fire response and departmental collaboration capabilities. Attached Figure Description
[0204] Figure 1 This is a flowchart of the present invention;
[0205] Figure 2 This is a system diagram of the present invention. Detailed Implementation
[0206] To enable those skilled in the art to understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.
[0207] The present invention will now be described in detail with reference to the accompanying drawings:
[0208] Example:
[0209] Please see the appendix Figure 1 This invention provides a multimodal-based approach. Large-scale intelligent fire management methods include:
[0210] Step 1: Collect multimodal perception data from the field environment and transmit the multimodal perception data as input data to the multimodal semantic recognition module;
[0211] Sub-step 1.1 involves constructing image acquisition channels, temperature acquisition channels, gas concentration acquisition channels, and audio acquisition channels by deploying multiple types of sensor nodes within the target building or location to acquire video frame sequences. Temperature sequence Gas concentration sequence and sound amplitude sequence ,in:
[0212] For time sampling points, for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time;
[0213] The sampling frequency of each channel must meet the condition of uniform timestamp alignment:
[0214] ,
[0215] in, , Maximum tolerance;
[0216] If the data acquisition delay of any channel exceeds the tolerance range, the synchronous re-acquisition mechanism will be triggered;
[0217] Sub-step 1.2 involves performing validity and continuity checks on the multimodal data sequence obtained in sub-step 1.1, and determining whether it meets the basic conditions for being sent to the multimodal semantic recognition module based on the following rules:
[0218] Image frame sharpness Greater than the set threshold Image sharpness can be defined using the Laplacian operator variance method:
[0219] ,in, For the Laplace operator, It is the variance function;
[0220] The change in the maximum value of three consecutive frames in the temperature sequence satisfies:
[0221] ,in, The threshold for temperature abrupt change. for Temperature value at any given time;
[0222] The rate of change in gas concentration is greater than the upper limit of background fluctuation:
[0223] ,in, The threshold for the rate of change in gas concentration. It is the unit of time differentiation;
[0224] The proportion of spectral energy in the sound signal within the alarm frequency band is greater than a preset ratio:
[0225] ,
[0226] in, , This is the frequency band for fire alarm sounds. , For the recording system frequency band, For sound spectral power density, The threshold for the ratio of audio signal energy. It is the unit of frequency differentiation;
[0227] When all four conditions above are met, a valid input window is formed. :
[0228] ;
[0229] Sub-step 1.3 involves processing the valid multimodal data filtered through sub-step 1.2. According to spatial location information Reorganize to construct a unified modal input tensor. The structure is as follows:
[0230] ,
[0231] in, The feature tensor after image modality encoding, , , Numerical encoding for temperature, gas, and audio modes. This is the spatial location information vector corresponding to the sensor;
[0232] Then tensor The input is fed into the encoding network of the multimodal semantic recognition module, preparing it for the semantic fusion process in step 2;
[0233] Step 2: Use the multimodal semantic recognition module to perform modal parsing and semantic fusion on the input data to generate a unified semantic description containing fire status elements;
[0234] Sub-step 2.1: The input tensor constructed in step 1... The input is fed into a modality-specific encoding network to generate modality feature representation vectors. , , , ,Right now:
[0235] ,
[0236] in, For visual modal encoders, It is a temperature mode encoder. It is a gas mode encoder. For audio modal encoders,
[0237] for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time;
[0238] The output feature vectors are all mapped to a uniform dimensional space;
[0239] Sub-step 2.2 involves performing relevance matching on the modal features obtained in sub-step 2.1, and constructing a weighted fusion vector based on the following fusion weight calculation strategy. :
[0240] ,
[0241] in, , , , The fusion weight coefficients for visual, temperature, gas, and audio modalities are listed in order.
[0242] The weighting coefficients are determined through a soft attention mechanism:
[0243] ,
[0244] in, For learnable weight vectors, It is the transpose of the weight vector. , The feature vector of the mode;
[0245] If any of the fusion weights , To minimize the effective fusion weight threshold, the system determines that the modality information has a high noise ratio and performs feature suppression operation to retain the main modality information for the next step;
[0246] Sub-step 2.3 involves fusing the feature vectors. The input is fed into the semantic generation module to generate a structured semantic description vector. ,in:
[0247] The results of fire source location. According to the fire development level, This is an abnormal phenomenon type;
[0248] Among them, the fire development level The judgment is made based on the following logic:
[0249] ,
[0250] in, This is the upper limit of normal temperature. The threshold for the rate of temperature rise. For image clarity, The amplitude value of the sound. This is the critical value of the temperature gradient. This represents the critical lower limit for image sharpness. This is the minimum loudness threshold for alarm sounds;
[0251] Each level forms a fire severity decision chain based on a joint threshold logic;
[0252] The final output semantic description vector This will be used as input for step 3 to further calculate the on-site fire risk level;
[0253] Step 3: Input the generated unified semantic description into the dynamic risk judgment module, and combine it with the current time information, spatial location parameters of the collection point and population density data to output the fire risk level that reflects the on-site risk status.
[0254] Sub-step 3.1: The unified semantic description vector generated in step 2... As input, extract the location of the fire source. Fire development level and types of phenomena At the same time, combined with the spatial location vector obtained in step 1 and current time information Integrate and construct scene analysis input vectors The structure is as follows:
[0255] ,
[0256] in, The coordinates of the fire source location. Fire intensity level This is a label for the phenomenon type. The spatial coordinates of the sampling point This is the current timestamp. Human flow density;
[0257] vector This provides input data for the next sub-step, risk factor calculation.
[0258] Sub-step 3.2 involves processing the input vector obtained in sub-step 3.1. Input the risk factor calculation module to calculate the following risk factors in sequence: Fire risk factor Population exposure risk factors Spatiotemporal coupling risk factors ,in:
[0259] Fire risk factors Defined as:
[0260] ,
[0261] in, Assigning weights based on fire intensity level Weights for phenomenon types. A unique code for the phenomenon label;
[0262] Population exposure risk factors Defined as:
[0263] ,
[0264] in, For population density risk coefficient, For human traffic density, The spatial decay factor, The distance between the fire source and the sampling point is expressed as the Euclidean distance.
[0265] Spatiotemporal coupling risk factor Defined as:
[0266] ,
[0267] in, For spatiotemporal fluctuation coefficient, For time frequency, This is the current timestamp;
[0268] like and The situation was determined to be high-risk for fire, and the next step of the risk level assessment process was initiated.
[0269] Sub-step 3.3: Calculate the result obtained in sub-step 3.2. , , The judgment result is input into the risk level mapping function, and the fire risk level is output. The judgment logic is as follows:
[0270] ,
[0271] in, The fire risk threshold For the population exposure risk threshold, This is the threshold for spatiotemporal coupling risk.
[0272] Ultimately, the fire risk level will be determined. As input to step 4, it is used to invoke the digital twin space model and trigger the corresponding response strategy;
[0273] Step 4: Based on the obtained fire risk level, call the pre-built digital twin space model to realize the synchronous mapping of the fire status in the virtual space;
[0274] Sub-step 4.1: Based on the fire risk level generated in step 3. and the spatial location of the fire source Select the corresponding building instance from the pre-built digital twin model database. And load the current scene;
[0275] Initialize the fire situation field tensor It is used to simulate the thermal diffusion state of a fire source and is defined as follows:
[0276] ,
[0277] in, For the spatial coordinates of the fire source, This is the ignition source intensity coefficient. The heat decay coefficient, This refers to the current time point;
[0278] If no instance matching the fire source location can be found in the model database. The system stops synchronization mapping and generates a fault indication event for feedback analysis in step 8;
[0279] Sub-step 4.2, based on the multimodal raw data tensors obtained in steps 1 and 2... and fused feature vectors The temperature field, concentration field, image thermal area, and sound source direction are sequentially mapped onto the virtual space scene to construct a multimodal visualization field. ;
[0280] The definition is as follows:
[0281] Temperature distribution field :
[0282] ,
[0283] Combustible gas concentration field :
[0284] ,
[0285] Image thermal projection map Image mode encoder The output features are obtained by reverse mapping.
[0286] Sound source direction cone Calculated by inversion of spatial directivity in audio features:
[0287] ,
[0288] in, The angle of the main direction of the sound source. For audio signals at frequencies Power spectral density on;
[0289] All mapping fields together form a three-dimensional multimodal visual space. ;
[0290] If any mode is missing during the mapping process, automatic interpolation is performed at that location. The interpolation algorithm uses three-dimensional Gaussian kernel interpolation.
[0291] ,
[0292] in, For three-dimensional space points Interpolated modal values at the location, The spatial smoothing coefficient is... Original observation point Modal data values at the location, For the first Spatial coordinates of the original modal sampling points, It is an exponential function;
[0293] Sub-step 4.3 involves constructing the multimodal field of view in sub-step 4.2. Fire situation tensor The digital twin dynamic driving engine is jointly input to update the virtual fire situation status at the current moment;
[0294] Define the virtual fire scene state frame for:
[0295] ,
[0296] in, For twin state generation functions, For twin space structure model, Given the current state of fire heat spread, This is a modal data field.
[0297] status frame After outputting the results, proceed to step 5, the virtual simulation phase of the response strategy.
[0298] If the state frame update fails, the system records the current tensor and configuration state, and enters the feedback analysis and processing module in step 8 for model correction.
[0299] Step 5: Perform virtual drills of various response strategies in the constructed digital twin space model, and generate evaluation results of the response effects of each strategy.
[0300] Sub-step 5.1, based on the virtual fire status frame generated in step 4 and fire risk level The system selects a set of multi-response strategies. Conduct virtual drills;
[0301] Response strategies All contain the set of operation instructions to be executed. According to the location of the fire source and risk level Choose the most suitable set of strategies Input into the digital twin engine for simulation;
[0302] During this phase, all response strategies will proceed according to the predetermined simulation timeframe. The simulation is conducted internally, and the simulation process is based on the time schedule. Update fire status and operational feedback;
[0303] Sub-step 5.2, each response strategy When executed in a digital twin environment, the following performance metrics are recorded in real time:
[0304] Fire extinguishing time : Representation strategy The time required for the fire to be completely extinguished after execution is calculated using the following formula:
[0305] ,
[0306] in, The initial time when the response strategy begins to be executed. Time for extinguishing the fire This represents the thermal diffusion state of the ignition source.
[0307] Personnel evacuation time : Representation strategy The formula for calculating the time required for the evacuation of personnel involved in the incident is as follows:
[0308] ,
[0309] in, In order to evacuate the number of people, For personnel evacuation distance For strategy The average evacuation speed of people being evacuated in the middle of the evacuation process;
[0310] Device response time : Representation strategy The response time of the device is calculated using the following formula:
[0311] ,
[0312] in, For the number of devices involved, For equipment Distance for mission execution For device response speed;
[0313] Resource consumption : Representation strategy The resource consumption required for execution is calculated using the following formula:
[0314] ,
[0315] in, The number of resource types required. For resources Consumption coefficient, For strategy China's resources Demand;
[0316] Sub-step 5.3: Based on the various effect indicators recorded in sub-step 5.2, construct a comprehensive effect evaluation index for the response strategy. This is used to compare the merits of different strategies, and the evaluation formula is:
[0317] ,
[0318] in, , , , The corresponding weighting coefficients for each indicator, , In order of priority, these are fire extinguishing time, evacuation time, equipment response time, and resource consumption.
[0319] Based on the evaluation results The system selects the response strategy with the best overall effect. Proceed to the next step: generating scheduling instructions;
[0320] Step 6: Select the optimal response strategy based on the evaluation results, and generate scheduling instructions through the adaptive linkage instruction generation module;
[0321] Sub-step 6.1: Based on the response strategy evaluation index calculated in step 5.3. Select the strategy with the best overall effect. It meets the following conditions:
[0322] ,
[0323] in, For the set of candidate response strategies, For the first The overall evaluation value of the strategy, The optimal response strategy;
[0324] If multiple strategies exist to satisfy Further compare the priority of key indicators in the strategy and judge them in order:
[0325] Prioritize the time for fire extinguishing The shortest;
[0326] If they are still the same, compare the evacuation times. ;
[0327] If you still cannot distinguish, select resource consumption. The lowest;
[0328] Sub-step 6.2: Based on the optimal response strategy selected in sub-step 6.1 and its operation instruction set Call the adaptive linkage instruction generation function Generate a scheduling instruction set for multiple execution terminals. ;
[0329] The logic for generating scheduling instructions is defined as follows:
[0330] ,
[0331] in, The first in the optimal response strategy Item operation, For the first The current physical location of the execution terminal. To determine the current load status of the execution terminal, This is the current system timestamp. Fire risk level, For scheduling instructions;
[0332] Simultaneously, an execution feasibility judgment function is introduced. If an instruction does not meet one of the following conditions, the scheduling instruction will be rejected:
[0333] The current load must not exceed the maximum load capacity.
[0334] The distance from the fire source must not exceed the executable range;
[0335] Command response time must be within the allowed timeframe;
[0336] in, This is the maximum load threshold for the terminal. It is a spatial distance function. The maximum executable space distance threshold. To estimate command transmission and response latency, This refers to the latest execution time of the operation instruction;
[0337] Sub-step 6.3 involves processing the scheduling instruction set generated in sub-step 6.2. Group by terminal device category and construct control distribution table ,each It contains all the instructions that similar devices need to execute;
[0338] Define the scheduling control logic for each group as follows:
[0339] ,
[0340] in, For scheduling instructions The type code of the device being pointed to. To distribute to the first The instruction set for this type of device;
[0341] Ultimately through the control interface module Push the scheduling control table to the terminal device system and upper-level management platform To ensure the automatic execution of response actions:
[0342] ,
[0343] If the device reports a status of instruction rejection, the system will transfer the instruction to the alternative execution device;
[0344] Step 7: Send the scheduling instructions to the target execution terminal device and management system to achieve automated response actions;
[0345] Sub-step 7.1, based on the scheduling control table constructed in step 6.3 Call the control interface module Send corresponding instruction sets to various terminal devices;
[0346] Define the success rate of sending each command as follows:
[0347] ,
[0348] in, For scheduling instructions, For the initial transmission success rate of the interface, For transmission attenuation coefficient, Send the path distance for the command;
[0349] like ,in, To minimize the acceptable probability of successful communication, the transmission is aborted, the command is marked as needing retransmission, and added to the retransmission queue. ;
[0350] Sub-step 7.2: All successfully issued instructions Upon reaching the target device, the terminal device activates the automatic response logic module according to the instructions. Execute the specified action, and define the probability of the terminal completing the execution as follows:
[0351] ,
[0352] in, This is the terminal baseline execution capability coefficient. This indicates the current load status of the device. This is the maximum load that the equipment can bear. The device status availability function is defined as follows:
[0353] ,
[0354] like The instruction is transferred to the backup device and the device abnormality information is recorded and transmitted to the feedback analysis module in step 8.
[0355] The actions performed include physical operations and logical control, and a response status code is actively returned after the operation. The possible values are as follows:
[0356] Response successful;
[0357] Response failed;
[0358] The terminal refused to execute.
[0359] Terminal disconnected;
[0360] Sub-step 7.3: After the device performs the action, it calls the status reporting module. , will include equipment Operation type, execution timestamp Status codes Execution feedback data Execution status data packet Report to the upper management platform and control center ;
[0361] Define state synchronization delay for:
[0362] ,
[0363] in, The timestamp of the W2 device completing the operation. The time when the W2 management platform receives the status packet;
[0364] like If the delay exceeds the maximum tolerance value for state synchronization, the state is marked as out of sync, triggering the synchronization retransmission logic. ;
[0365] After all received execution statuses are summarized, the management platform will write them into the response log, which will serve as input data for feedback data analysis in the subsequent step 8.
[0366] Step 8: Collect feedback data after the response action is executed and input it into the system for analysis. Based on the feedback results, dynamically optimize the subsequent response strategy to achieve closed-loop control.
[0367] Sub-step 8.1: Based on the device execution status data packet received by the management platform in step 7.3. Data collection includes terminal devices Response status codes Feedback data Execution timestamp Synchronization delay The original feedback data set, including ;
[0368] The feedback data is standardized to construct analysis vectors in a unified format. :
[0369] ,
[0370] in, , To provide the mean and standard deviation of the feedback values, , The mean and standard deviation of the synchronization delay. In response to the status code, This is the normalized value of the equipment load. This is the maximum load that the equipment can withstand.
[0371] like If any standardized indicator in any dimension exceeds the preset anomaly range, the data is marked as an anomaly and added to the anomaly feedback set. ;
[0372] Sub-step 8.2 involves standardizing the feedback data vector set from step 8.1. Perform cluster analysis and multidimensional evaluation, and define a scoring function for the effectiveness of each strategy in the current fire instance. :
[0373] ,
[0374] in, For the success rate,
[0375] This represents the average load percentage. For average synchronization delay, For strategy execution evaluation weighting coefficients, For indicator functions;
[0376] like ,in, The lower limit of acceptable performance is used to trigger the policy adjustment logic.
[0377] Sub-step 8.3 addresses the response strategies identified as inefficient in step 8.2. According to the abnormal feedback set Based on the error distribution characteristics, optimize the response strategy parameters and generate a new set of candidate strategies. ;
[0378] The specific adjustment methods are as follows:
[0379] If device response failures are concentrated in a specific type of terminal, reduce the call weight of that type of device in the policy;
[0380] For situations where the failure rate is high in high-load areas, increase the resource reservation factor for devices in that area;
[0381] If the synchronization delay is severe, terminal equipment that is geographically closer to the control center should be given priority.
[0382] Define the strategy parameter adjustment function It acts on the original policy parameter vector. Generate new parameters :
[0383] ,
[0384] in, For learning rate, The loss function is constructed based on abnormal feedback samples. The gradient of the policy parameters with respect to the loss function;
[0385] After optimization, a new set of candidate strategies will be created. The results are then resubmitted to the virtual exercise module in step 5.1 for effect evaluation, thereby achieving cyclical self-optimization of the strategy.
[0386] The benefits of Step 1 include: systematically collecting heterogeneous data on images, temperature, gas concentration, and audio from the field environment by constructing a multimodal sensing channel; ensuring the integrity and reliability of the collected information through data alignment and validity detection mechanisms; and laying a solid foundation for subsequent fire assessment. This design effectively reduces the risk of false alarms and missed alarms caused by anomalies in a single sensor. Data cleaning and window filtering processes improve data quality, making the data entering subsequent model processing more representative and valuable.
[0387] The benefit of step 2 is that by fusing and analyzing information such as images, temperature, gas, and sound through a multimodal semantic recognition module, it breaks through the limitations of isolated data processing in traditional sensing systems, achieving deep integration of multi-source information at the semantic level. This mechanism enables the system to construct fire descriptions with interpretability and context awareness, effectively improving the recognition accuracy in complex, ambiguous, or data-constrained scenarios, and providing strong support for the accurate classification and assessment of fires.
[0388] The benefits of step 3 lie in its introduction of a dynamic risk assessment mechanism. This mechanism integrates the semantic description of the fire situation with factors such as time, space, and population density to intelligently generate fire risk levels. This method accurately depicts the dynamic evolution of fire risk, enhances the system's adaptability to different scenarios, and makes risk assessments in various fire situations more targeted and operable. This provides a scientific basis for differentiated deployment of response strategies and significantly improves the overall predictive capability of the system.
[0389] Step 4 maps the actual fire situation to a digital twin space, achieving a high degree of realism and multi-dimensional visualization of the situation through virtual reconstruction of thermal fields, gas fields, image hotspots, and sound source directions. This design enhances the intuitiveness of fire scene perception and provides an accurate environmental simulation platform for the formulation and evaluation of subsequent response strategies. The system has anomaly detection and fault tolerance mechanisms in the virtual space, ensuring that it can maintain overall control of the fire situation even when limited by real-world constraints.
[0390] The benefit of step 5 is that it allows for virtual drills and effectiveness evaluations of multiple response strategies within a digital twin environment. Simulations rehearse the execution results of different response paths in advance, avoiding blind decision-making. The system's quantitative analysis of response effectiveness helps select the optimal response plan from multiple dimensions, balancing response speed, personnel evacuation efficiency, and resource utilization. This provides accurate and reliable decision support for subsequent actions, effectively improving the system's responsiveness to sudden fires.
[0391] The benefit of step 6 is that, through an adaptive linkage instruction generation mechanism, a multi-terminal scheduling instruction set is automatically constructed based on the virtual simulation evaluation results, providing flexible response and real-time adjustment capabilities. This mechanism considers actual factors such as device load, distance, and timeliness, ensuring that the instructions are feasible to execute and efficient to implement, greatly improving the intelligence level of response deployment.
[0392] The benefit of step 7 is that it accurately distributes the generated response commands to various terminal devices through a systematic scheduling and control mechanism, tracks their execution status in real time, and ensures that all emergency response actions are implemented as planned. The automatic feedback of device execution results constructs a closed-loop control chain for the response process, providing key data support for dynamic system monitoring and subsequent strategy adjustments.
[0393] The benefit of step 8 is that by collecting and analyzing feedback data during the device's execution process, the effectiveness of the response strategy can be evaluated and parameters optimized, achieving continuous iteration and closed-loop self-optimization at the strategy level. This mechanism enables the system to continuously adjust strategy weights and execution configurations based on actual operational results, improving the adaptability and robustness of subsequent responses.
[0394] Please see the appendix Figure 2 A multimodal-based A large-scale intelligent fire management system, which includes:
[0395] The multimodal semantic recognition module is used to fuse and process video, temperature, gas concentration and voice information from the scene and output a unified fire description;
[0396] The dynamic risk assessment module is used to generate fire risk scores based on semantic descriptions and environmental context factors.
[0397] The digital twin response module is used to build a virtual space model and perform response strategy rehearsals.
[0398] The adaptive linkage instruction generation module is used to generate scheduling instructions based on the exercise results and control the relevant equipment or systems to execute them;
[0399] The feedback analysis module is used to collect response result data and dynamically optimize subsequent response strategies.
[0400] This invention proposes a smart fire management system based on a multimodal large model, constructing an intelligent fire protection solution that spans the entire process of perception, judgment, response, and optimization. This system significantly improves the accuracy of fire identification and the intelligence of response decisions, achieving a fundamental transformation in fire response from passive pre-planning to proactive learning, and from static response to dynamic collaboration. It possesses high adaptability, scalability, and visualization capabilities, and can be widely applied to fire safety management in various complex scenarios, demonstrating significant social benefits and industrial application value.
[0401] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal-based The large-scale intelligent fire management method is characterized by, include: Step 1: Collect multimodal perception data from the field environment and transmit the multimodal perception data as input data to the multimodal semantic recognition module; Step 2: Use the multimodal semantic recognition module to perform modal parsing and semantic fusion on the input data to generate a unified semantic description containing fire status elements; Step 3: Input the generated unified semantic description into the dynamic risk judgment module, and combine it with the current time information, spatial location parameters of the collection point and population density data to output the fire risk level that reflects the on-site risk status. Step 4: Based on the obtained fire risk level, call the pre-built digital twin space model to realize the synchronous mapping of the fire status in the virtual space; Step 5: Perform virtual drills of various response strategies in the constructed digital twin space model, and generate evaluation results of the response effects of each strategy. Step 6: Select the optimal response strategy based on the evaluation results, and generate scheduling instructions through the adaptive linkage instruction generation module; Step 7: Send the scheduling instructions to the target execution terminal device and management system to achieve automated response actions; Step 8: Collect feedback data after the response action is executed and input it into the system for analysis. Based on the feedback results, dynamically optimize the subsequent response strategy to achieve closed-loop control.
2. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 1 further includes: Sub-step 1.1 involves constructing image acquisition channels, temperature acquisition channels, gas concentration acquisition channels, and audio acquisition channels by deploying multiple types of sensor nodes within the target building or location to acquire video frame sequences. Temperature sequence Gas concentration sequence and sound amplitude sequence ,in: For time sampling points, for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time; The sampling frequency of each channel must meet the condition of uniform timestamp alignment: , in, , Maximum tolerance; If the data acquisition delay of any channel exceeds the tolerance range, the synchronous re-acquisition mechanism will be triggered; Sub-step 1.2 involves performing validity and continuity checks on the multimodal data sequence obtained in sub-step 1.1, and determining whether it meets the basic conditions for being sent to the multimodal semantic recognition module based on the following rules: Image frame sharpness Greater than the set threshold Image sharpness can be defined using the Laplacian operator variance method: ,in, For the Laplace operator, It is the variance function; The change in the maximum value of three consecutive frames in the temperature sequence satisfies: ,in, The threshold for temperature abrupt change. for Temperature value at any given time; The rate of change in gas concentration is greater than the upper limit of background fluctuation: ,in, The threshold for the rate of change in gas concentration. It is the unit of time differentiation; The proportion of spectral energy in the sound signal within the alarm frequency band is greater than a preset ratio: , in, , This is the frequency band for fire alarm sounds. , For the recording system frequency band, For sound spectral power density, The threshold for the ratio of audio signal energy. It is the unit of frequency differentiation; When all four conditions above are met, a valid input window is formed. : ; Sub-step 1.3 involves processing the valid multimodal data filtered through sub-step 1.
2. According to spatial location information Reorganize to construct a unified modal input tensor. The structure is as follows: , in, The feature tensor after image modality encoding, , , Numerical encoding for temperature, gas, and audio modes. This is the spatial location information vector corresponding to the sensor; Then tensor The input is fed into the encoding network of the multimodal semantic recognition module, preparing it for the semantic fusion process in step 2.
3. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 2 further includes: Sub-step 2.1: The input tensor constructed in step 1... The input is fed into a modality-specific encoding network to generate modality feature representation vectors. , , , ,Right now: , in, For visual modal encoders, It is a temperature mode encoder. It is a gas mode encoder. For audio modal encoders, for Image frames at any given time, for Temperature value at time, for The target flammable gas concentration at any given time. for The intensity of the ambient sound signal at any given time; The output feature vectors are all mapped to a uniform dimensional space; Sub-step 2.2 involves performing relevance matching on the modal features obtained in sub-step 2.1, and constructing a weighted fusion vector based on the following fusion weight calculation strategy. : , in, , , , The fusion weight coefficients for visual, temperature, gas, and audio modalities are listed in order. The weighting coefficients are determined through a soft attention mechanism: , in, For learnable weight vectors, It is the transpose of the weight vector. , The feature vector of the mode; If any of the fusion weights , To minimize the effective fusion weight threshold, the system determines that the modality information has a high noise ratio and performs feature suppression operation to retain the main modality information for the next step; Sub-step 2.3 involves fusing the feature vectors. The input is fed into the semantic generation module to generate a structured semantic description vector. ,in: The results of the fire source location, The fire development level, This is an abnormal phenomenon type; Among them, the fire development level The judgment is made based on the following logic: , in, This is the upper limit of normal temperature. The threshold for the rate of temperature rise. For image clarity, The amplitude value of the sound. This is the critical value of the temperature gradient. This represents the critical lower limit for image sharpness. This is the minimum loudness threshold for alarm sounds; Each level forms a fire severity decision chain based on a joint threshold logic; The final output semantic description vector This will be used as input for step 3, participating in the further calculation of the on-site fire risk level.
4. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 3 further includes: Sub-step 3.1: The unified semantic description vector generated in step 2... As input, extract the location of the fire source. Fire development level and types of phenomena At the same time, combined with the spatial location vector obtained in step 1 and current time information Integrate and construct scene analysis input vectors The structure is as follows: , in, The coordinates of the fire source location. Fire intensity level This is a label for the phenomenon type. The spatial coordinates of the sampling point This is the current timestamp. Human flow density; vector This provides input data for the next sub-step, risk factor calculation. Sub-step 3.2 involves processing the input vector obtained in sub-step 3.
1. Input the risk factor calculation module to calculate the following risk factors in sequence: Fire risk factor Population exposure risk factors Spatiotemporal coupling risk factors ,in: Fire risk factors Defined as: , in, Assigning weights to fire intensity levels Weights for phenomenon types A unique code for the phenomenon label; Population exposure risk factors Defined as: , in, For population density risk coefficient, For human traffic density, The spatial decay factor, The distance between the fire source and the sampling point is expressed as the Euclidean distance. Spatiotemporal coupling risk factor Defined as: , in, For spatiotemporal fluctuation coefficient, For time frequency, This is the current timestamp; like and The situation was determined to be high-risk for fire, and the next step of the risk level assessment process was initiated. Sub-step 3.3: Calculate the result obtained in sub-step 3.
2. , , The judgment result is input into the risk level mapping function, and the fire risk level is output. The judgment logic is as follows: , in, The fire risk threshold For the population exposure risk threshold, This is the threshold for spatiotemporal coupling risk. Ultimately, the fire risk level will be determined. As input to step 4, it is used to invoke the digital twin space model and trigger the corresponding response strategy.
5. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 4 further includes: Sub-step 4.1: Based on the fire risk level generated in step 3. and the spatial location of the fire source Select the corresponding building instance from the pre-built digital twin model database. And load the current scene; Initialize the fire situation field tensor It is used to simulate the thermal diffusion state of a fire source and is defined as follows: , in, For the spatial coordinates of the fire source, This is the ignition source intensity coefficient. The heat decay coefficient, This refers to the current time point; If no instance matching the fire source location can be found in the model database. The system stops synchronization mapping and generates a fault indication event for feedback analysis in step 8; Sub-step 4.2, based on the multimodal raw data tensors obtained in steps 1 and 2... and fused feature vectors The temperature field, concentration field, image thermal area, and sound source direction are sequentially mapped onto the virtual space scene to construct a multimodal visualization field. ; The definition is as follows: Temperature distribution field : , Combustible gas concentration field : , Image thermal projection map Image mode encoder The output features are obtained by reverse mapping. Sound source direction cone Calculated by inversion of spatial directivity in audio features: , in, The angle of the main direction of the sound source. For audio signals at frequencies Power spectral density on; All mapping fields together form a three-dimensional multimodal visual space. ; If any mode is missing during the mapping process, automatic interpolation is performed at that location. The interpolation algorithm uses three-dimensional Gaussian kernel interpolation. , in, For three-dimensional space points Interpolated modal values at the location, The spatial smoothing coefficient is... Original observation point Modal data values at the location, For the first Spatial coordinates of the original modal sampling points, It is an exponential function; Sub-step 4.3 involves constructing the multimodal field of view in sub-step 4.
2. Fire situation tensor The digital twin dynamic driving engine is jointly input to update the virtual fire situation status at the current moment; Define virtual fire scene state frames for: , in, For twin state generation functions, For twin space structure model, Given the current state of fire heat spread, For modal data fields; status frame After outputting the results, proceed to step 5, the virtual simulation phase of the response strategy. If the state frame update fails, the system records the current tensor and configuration state, and proceeds to the feedback analysis and processing module in step 8 for model correction.
6. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 5 further includes: Sub-step 5.1, based on the virtual fire status frame generated in step 4 and fire risk level The system selects a set of multi-response strategies. Conduct virtual drills; Response strategies All contain the set of operation instructions to be executed. According to the location of the fire source and risk level Choose the most suitable set of strategies Input into the digital twin engine for simulation; During this phase, all response strategies will proceed according to the predetermined simulation timeframe. The simulation is conducted internally, and the simulation process is based on the time schedule. Update fire status and operational feedback; Sub-step 5.2, each response strategy When executed in a digital twin environment, the following performance metrics are recorded in real time: Fire extinguishing time : Representation strategy The time required for the fire to be completely extinguished after execution is calculated using the following formula: , in, The initial time when the response strategy begins to be executed. Time for extinguishing the fire This represents the thermal diffusion state of the ignition source. Personnel evacuation time : Representation strategy The formula for calculating the time required for the evacuation of personnel involved in the incident is as follows: , in, In order to evacuate the number of people, For personnel evacuation distance For strategy The average evacuation speed of people being evacuated in the middle of the evacuation; Device response time : Representation strategy The response time of the device is calculated using the following formula: , in, For the number of devices involved, For equipment Distance for mission execution For device response speed; Resource consumption : Representation strategy The resource consumption required for execution is calculated using the following formula: , in, The number of resource types required. For resources Consumption coefficient, For strategy China's resources Demand; Sub-step 5.3: Based on the various effect indicators recorded in sub-step 5.2, construct a comprehensive effect evaluation index for the response strategy. This is used to compare the merits of different strategies, and the evaluation formula is: , in, , , , The corresponding weighting coefficients for each indicator, , In order of priority, these are fire extinguishing time, evacuation time, equipment response time, and resource consumption. Based on the evaluation results The system selects the response strategy with the best overall effect. Proceed to the next step: generating scheduling instructions.
7. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 6 further includes: Sub-step 6.1: Based on the response strategy evaluation index calculated in step 5.
3. Select the strategy with the best overall effect. It meets the following conditions: , in, For a set of candidate response strategies, For the first The overall evaluation value of the strategy, The optimal response strategy; If multiple strategies exist to satisfy Further compare the priority of key indicators in the strategy and judge them in order: Prioritize the time for fire extinguishing The shortest; If they are still the same, compare the evacuation times. ; If you still cannot distinguish, select resource consumption. The lowest; Sub-step 6.2: Based on the optimal response strategy selected in sub-step 6.1 and its operation instruction set Call the adaptive linkage instruction generation function Generate a scheduling instruction set for multiple execution terminals. ; The logic for generating scheduling instructions is defined as follows: , in, The first in the optimal response strategy Item operation, For the first The current physical location of the execution terminal. To determine the current load status of the execution terminal, This is the current system timestamp. Fire risk level, For scheduling instructions; Simultaneously, an execution feasibility judgment function is introduced. If an instruction does not meet one of the following conditions, the scheduling instruction will be rejected: The current load must not exceed the maximum load capacity. The distance from the fire source must not exceed the executable range; Command response time must be within the allowed timeframe; in, This is the maximum load threshold for the terminal. It is a spatial distance function. The maximum executable space distance threshold. To estimate command transmission and response latency, This refers to the latest execution time of the operation instruction; Sub-step 6.3 involves processing the scheduling instruction set generated in sub-step 6.
2. Group by terminal device category and construct control distribution table ,each It contains all the instructions that similar devices need to execute; Define the scheduling control logic for each group as follows: , in, For scheduling instructions The type code of the device being pointed to. To distribute to the first The instruction set for this type of device; Ultimately through the control interface module Push the scheduling control table to the terminal device system and upper-level management platform To ensure the automatic execution of response actions: , If the device reports a status of instruction rejection, the system will transfer the instruction to the alternative execution device.
8. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, In step 7, the scheduling instruction is sent to the target execution terminal device and management system to realize automated response actions, which further includes: Sub-step 7.1, based on the scheduling control table constructed in step 6.3 Call the control interface module Send corresponding instruction sets to various terminal devices; Define the success rate of sending each command as follows: , in, For scheduling instructions, For the initial transmission success rate of the interface, The transmission attenuation coefficient, Send the path distance for the command; like ,in, To minimize the acceptable probability of successful communication, the transmission is aborted, the command is marked as needing retransmission, and added to the retransmission queue. ; Sub-step 7.2: All successfully issued instructions Upon reaching the target device, the terminal device activates the automatic response logic module according to the instructions. Execute the specified action, and define the probability of the terminal completing the execution as follows: , in, This is the terminal baseline execution capability coefficient. This indicates the current load status of the device. This is the maximum load that the equipment can bear. The device status availability function is defined as follows: , like The instruction is transferred to the backup device and the device abnormality information is recorded and transmitted to the feedback analysis module in step 8. The actions performed include physical operations and logical control, and a response status code is actively returned after the operation. The possible values are as follows: Response successful; Response failed; The terminal refused to execute. Terminal lost connection; Sub-step 7.3: After the device performs the action, it calls the status reporting module. , will include equipment Operation type, execution timestamp Status codes Execution feedback data Execution status data packet Report to the upper management platform and control center ; Define state synchronization delay for: , in, The timestamp of the W2 device completing the operation. The time when the W2 management platform receives the status packet; like If the state delay exceeds the maximum tolerance value for state synchronization, the state is marked as out of sync, triggering the synchronization retransmission logic. ; After all received execution statuses are summarized, the management platform will write them into the response log, which will serve as input data for feedback data analysis in subsequent step 8.
9. A multimodal-based method according to claim 1 The large-scale intelligent fire management method is characterized by, Step 8 further includes: Sub-step 8.1: Based on the device execution status data packet received by the management platform in step 7.
3. Data collection includes terminal devices Response status codes Feedback data Execution timestamp Synchronization delay The original feedback data set, including ; The feedback data is standardized to construct analysis vectors in a unified format. : , in, , To provide the mean and standard deviation of the feedback values, , The mean and standard deviation of the synchronization delay. In response to the status code, This is the normalized value of the equipment load. This is the maximum load that the equipment can withstand. like If any standardized indicator in any dimension exceeds the preset anomaly range, the data is marked as an anomaly and added to the anomaly feedback set. ; Sub-step 8.2 involves standardizing the feedback data vector set from step 8.
1. Perform cluster analysis and multidimensional evaluation, and define a scoring function for the effectiveness of each strategy in the current fire instance. : , in, For the success rate, This represents the average load percentage. For average synchronization delay, For strategy execution evaluation weighting coefficients, For indicator functions; like ,in, The lower limit of acceptable performance is used to trigger the policy adjustment logic. Sub-step 8.3 addresses the response strategies identified as inefficient in step 8.
2. According to the abnormal feedback set Based on the error distribution characteristics, optimize the response strategy parameters and generate a new set of candidate strategies. ; The specific adjustment methods are as follows: If device response failures are concentrated in a specific type of terminal, reduce the call weight of that type of device in the policy; For situations where the failure rate is high in high-load areas, increase the resource reservation factor for devices in that area; If the synchronization delay is severe, terminal equipment that is geographically closer to the control center should be given priority. Define the strategy parameter adjustment function It acts on the original policy parameter vector. Generate new parameters : , in, For learning rate, The loss function is constructed based on abnormal feedback samples. The gradient of the policy parameters with respect to the loss function; After optimization, a new set of candidate strategies will be created. The results are then resubmitted to the virtual exercise module in step 5.1 for effect evaluation, thereby achieving cyclical self-optimization of the strategy.
10. A multimodal-based approach A large-scale intelligent fire management system, according to any one of claims 1-9, is based on a multimodal... The large-scale intelligent fire management method is characterized by, The intelligent fire management system includes: The multimodal semantic recognition module is used to fuse and process video, temperature, gas concentration and voice information from the scene and output a unified fire description; The dynamic risk assessment module is used to generate fire risk scores based on semantic descriptions and environmental context factors. The digital twin response module is used to build a virtual space model and perform response strategy rehearsals. The adaptive linkage instruction generation module is used to generate scheduling instructions based on the exercise results and control the relevant equipment or systems to execute them; The feedback analysis module is used to collect response result data and dynamically optimize subsequent response strategies.
Citation Information
Patent Citations
Intelligent fire management method and system based on multimodal AI big model
CN117910811B
Intelligent fire-fighting remote management system
CN119992805A
Cited By
Intelligent fire-fighting AI dynamic early warning method and system fusing edge computing and digital twinning
CN122336932A