Intelligent fire-fighting hidden danger identification method based on multi-modal fusion

By constructing a spatiotemporally coupled causal alignment graph and a self-evolving semantic decision graph, and combining BIM data and air duct structure, the self-learning and spatial constraint fusion of multimodal data are achieved, solving the problems of identification accuracy and adaptability of intelligent fire protection systems in complex environments, and improving the accuracy and stability of fire hazard identification.

CN121765637APending Publication Date: 2026-03-31HANGZHOU LIANKE TIANCHEN SECURITY TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing smart fire protection systems suffer from problems such as insufficient accuracy of single-modal recognition, high false alarm rate, lack of self-learning ability and poor adaptability in fire hazard identification, especially in complex environments where it is difficult to accurately locate the location and spread path of hazards.

Method used

By employing a spatiotemporally coupled causal alignment graph, a spatial topological awareness attention mechanism, and a self-evolving semantic decision graph, event-level alignment, spatial constraint fusion, and self-learning updates of multimodal data are achieved. Multimodal events are aligned by constructing a spatiotemporally coupled causal alignment graph, feature fusion is performed by combining BIM data and wind tunnel structure, and dynamic adjustments are made using the self-evolving semantic decision graph.

Benefits of technology

It significantly improves the accuracy and stability of hazard identification, reduces the false alarm rate, and has high accuracy, low false alarm rate and strong robustness. It also achieves cross-scenario adaptability, reduces maintenance costs and extends the model's lifespan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765637A_ABST
    Figure CN121765637A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent fire-fighting hidden danger identification method based on multi-modal fusion, and the method comprises the steps: collecting multi-source heterogeneous data in a building environment, and carrying out the unified processing of the data, and obtaining standardized multi-modal original input data; event anchor point detection is carried out on the multi-modal input, key events are extracted, and a corresponding multi-modal event sequence is generated; constructing an event graph containing time and causal edges, and outputting an alignment event stream by using a space-time coupling causal alignment module; importing BIM and air duct structure information to establish a spatial topological graph, and fusing event streams to generate spatial constraint features; constructing a self-evolution semantic decision map based on the fusion features, and dynamically adjusting node weights and edge connection relationships; and carrying out hidden danger identification and grade division on the real-time data by using the decision map, and outputting an early warning signal and positioning information. According to the method, high-precision identification and intelligent early warning of fire-fighting hidden dangers are realized through space-time causal alignment, spatial topology fusion and self-evolution semantic decision of multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and smart safety management technology, and in particular to a smart fire hazard identification method based on multimodal fusion. Background Technology

[0002] Existing smart fire protection systems typically rely on video surveillance, sensor data acquisition, and IoT communication technologies. They identify and analyze characteristic data such as flames, smoke, temperature, and gas concentrations to achieve preliminary monitoring and early warning of fire hazards. However, these systems generally employ single-modal or simple weighted multimodal fusion methods, lacking unified modeling of the temporal relationships and causal logic between different data sources. Due to differences in sampling frequency, response latency, and triggering sequence among video, audio, and environmental sensor data, misalignment or misjudgment of hazard signals often occurs, making it difficult for the system to accurately distinguish between real fires and environmental interference. This results in insufficient identification accuracy and a high false alarm rate.

[0003] Existing technologies primarily rely on static feature recognition models, failing to incorporate the building's internal spatial structure and its dynamic changes into the intelligent analysis framework. Traditional methods depend solely on two-dimensional video images or sensor locations, neglecting to utilize BIM data, duct structures, and spatial topology information such as fire compartments to establish a computable spatial constraint model. This results in the system struggling to accurately determine the location and spread path of hazards in complex environments. When doors and windows are open, airflow direction changes, or building structures are adjusted, the recognition model cannot dynamically adjust weights based on spatial connectivity, easily leading to missed detections or incorrect localizations.

[0004] Existing smart fire protection algorithms are mostly fixed-structure models trained offline, lacking self-learning and adaptive capabilities in actual operation. When application scenarios change or new types of hazards emerge, the system cannot evolve its structure and update parameters based on historical operation records and real-time feedback, requiring retraining before it can be put into use. This not only leads to high maintenance costs and poor adaptability but also causes the model to gradually become ineffective over long-term operation, making it difficult to meet the actual needs of smart city fire protection systems for continuous intelligence and dynamic self-evolution.

[0005] Therefore, how to provide a smart fire hazard identification method based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a smart fire hazard identification method based on multimodal fusion. This invention comprehensively employs spatiotemporally coupled causal alignment graphs, spatial topological awareness attention mechanisms, and self-evolving semantic decision graphs to perform event-level alignment of heterogeneous data such as video, audio, and environmental sensors, feature fusion under spatial topological constraints such as BIM and ventilation ducts, and achieves structured self-learning and evolutionary updates by combining historical operation and real-time feedback. This invention provides a detailed implementation process for multi-source data processing, spatiotemporal causal alignment, spatial constraint fusion, graph-based decision-making, and early warning linkage. It can significantly reduce false alarms and missed alarms caused by multimodal misalignment, improve the accuracy and stability of hazard location and level determination, and has the advantages of high accuracy, low false alarm rate, strong robustness, good interpretability, and rapid adaptation across scenarios.

[0007] A smart fire hazard identification method based on multimodal fusion according to an embodiment of the present invention includes: Collect multi-source heterogeneous data from the built environment, process the multi-source heterogeneous data, and obtain multimodal raw input data; Event anchor point detection is performed on the multimodal raw input data to extract key event points for each mode, including brightness abrupt changes, sound pressure peaks, smoke concentration changes, and temperature inflection points, and to generate event sequences for each mode. Based on the event sequences of each modality, an initial event graph containing temporal and causal edges is constructed. The event graph is then jointly aligned using the spatiotemporal coupling causal alignment graph module to ensure that the multimodal events maintain consistency in temporal order and causal relationship, and outputs the spatiotemporally aligned multimodal event stream. Import BIM data, duct structure data and detector layout information to establish a spatial topology map. Based on the spatial topology awareness attention mechanism, the spatiotemporally aligned multimodal event flow is fused with the spatial topology map. Multimodal fusion features with spatial constraints are generated with topology embedding as the bias. Based on multimodal fusion features, a self-evolving semantic decision graph is constructed, semantic nodes are generated and decision edges are established to form an initial graph structure. Based on historical operation records and real-time feedback information, the node weights and edge connection relationships are dynamically adjusted to complete the self-learning and evolution update of the structure. By utilizing the self-evolving semantic decision graph, the system identifies and classifies potential hazards from real-time input multimodal data. When the identification results reach a preset risk threshold, the system outputs corresponding warning signals and spatial location information of potential hazards.

[0008] Optionally, the multi-source heterogeneous data includes visible light video data, infrared video data, audio data, temperature data, humidity data, smoke concentration data, and gas concentration data.

[0009] Optionally, the processing of multi-source heterogeneous data includes unified time synchronization, noise filtering, and standardized coding of the multi-source heterogeneous data.

[0010] Optionally, generating the sequence of events for each modality includes: The data of visible light video, infrared video, audio, temperature, humidity, smoke concentration and gas concentration are denoised, de-drifted and uniformly sampled to establish a single time series index and generate a change trend sequence and change amplitude sequence for each data channel. An event prototype library is constructed based on historical calibration data. The event prototype library contains a set of constraint parameters for the duration, amplitude, rising or falling direction, duration and recovery time of sudden brightness changes, infrared radiation changes, sound pressure peaks, sudden increases in smoke concentration and temperature inflection points. For each data channel, a candidate event search is performed within the sliding time window. When the change amplitude and duration simultaneously satisfy the corresponding constraint parameter set, it is marked as a candidate event point of the channel, and the occurrence time, amplitude level, direction attribute and duration are recorded. Within the cross-modal consistency verification window, causal order consistency screening and delay distribution constraint verification are performed on candidate event points for each channel. The screening is based on a preset trigger sequence table and the allowed time difference interval. Candidate event points that are inconsistent with the trigger sequence table or exceed the time difference interval are eliminated. Conflicting candidate event points within the same window are uniquely retained according to priority rules. The selected candidate event points are summarized into event sequences by channel. The event sequences include the occurrence time, channel identifier, amplitude level, direction attribute, duration, and cross-modal consistency marker.

[0011] Optionally, the output spatiotemporally aligned multimodal event stream includes: The spatiotemporal coupled causal alignment graph module is invoked, taking the event sequences of each modality as input, to initialize the spatiotemporal event graph, the dual-anchor time grid, and the causal delay envelope library. The spatiotemporal coupled causal alignment graph module consists of a time-preserving unit, a causal-first inference unit, and a conflict arbitration and consistency mapping unit. The time-preserving unit performs order preservation, deduplication, and merging on event nodes within the dual-anchor time grid, deletes connections that do not meet the requirement that the source anchor point precedes the response anchor point, limits the event interval to within a preset lower and upper limit, and outputs a calibrated time connection set. The causal inference unit retrieves precursor events that satisfy minimum delay, maximum delay, and preferred delay constraints for each candidate successor event based on the causal delay envelope library, generates a candidate causal connection set, and establishes a local causal evidence table for each candidate connection. The local causal evidence table includes cross-modal consistency markers, amplitude level, direction attribute, duration, and degree of matching with delay constraints. In the spatiotemporal coupled causal alignment graph module, the counterfactual occlusion stability check is enabled. For each candidate causal connection, the predecessor occlusion and successor recalculation are performed. Within a fixed verification window, the persistence and amplitude stability of the successor event are checked to see if they are lower than the preset threshold. If they are lower than the threshold, the causal connection is confirmed. If they are higher than the threshold, the causal connection is removed, and the set of causal connections that pass the check is obtained. The conflict arbitration and consistency mapping unit performs priority sorting and one-to-one mapping on the verified causal connections. The priority is determined based on the comprehensive score of cross-modal consistency matching degree, proximity to the preferred delay, and amplitude and duration. A unique predecessor event is selected for each subsequent event, circular dependencies are cleared, a spatiotemporal coupled causal alignment graph is generated, and it is expanded into an aligned event chain in ascending time order to form a spatiotemporally aligned multimodal event flow.

[0012] Optionally, the generation of spatially constrained multimodal fusion features with topological embedding as a bias includes: A spatial topology map is created based on BIM data, duct structure data and detector layout information. The spatial topology map consists of spatial nodes and connecting edges. Spatial nodes record the location, fire compartment, visible range and obstruction mark. Connecting edges record the physical connection relationship, air flow direction and door and window valve status. Topological embedding information is generated on the spatial topology graph. For each spatial node, the connectivity, airflow direction and obstacle status of adjacent nodes and adjacent connected edges are aggregated to obtain topological description information for attention calculation. The real-time status of doors, windows and valves is written into the gate control tag of the corresponding connected edge. Multimodal event streams are mapped to spatial nodes of the spatial topology graph according to time windows. The mapping is determined based on the camera's field of view, sensor coverage, and spatial proximity. Events falling into the same spatial node are weighted and aggregated according to time sequence and source confidence to form a node event set. The spatial topological awareness attention mechanism is used to fuse events of spatial nodes, specifically as follows: Based on the visible range and occlusion markers of spatial nodes, suppression weights are set for events that are not within the visible range or are occluded, and pass weights are set for events that are within the visible range and are not occluded. The influence of events propagates from upstream nodes to downstream nodes along the airflow direction. Propagation is only allowed when the doors, windows, and valves on the connected edges are open, and propagation is blocked when they are closed. Events in the downstream direction are given a higher propagation weight, while events in the upstream direction are given a suppression weight. Based on the topological description information of fire compartment boundaries, floor shafts and air duct branches, state biases are assigned to spatial nodes and connecting edges, and the biases are updated as the state of doors, windows and valves changes. After completing the visual field occlusion gating, airflow path propagation and state adaptive bias, a multimodal fusion feature with spatial constraints is output for each spatial node.

[0013] Optionally, the step of dynamically adjusting node weights and edge connectivity based on historical operation records and real-time feedback information includes: Semantic elements are extracted from multimodal fusion features, and semantic nodes are generated one-to-one with spatial nodes. Each semantic node contains a unique identifier, its spatial node number, a list of modal components, historical hit counts, and the most recent update time. Semantic nodes are also classified into three categories according to their source: source nodes, transit nodes, and convergence nodes. Establish initial decision edges between semantic nodes, and execute them one by one according to the following three types of rules: Based on the time-cause-effect edge of the alignment event chain established by the output, the edge is connected only when the timestamp of the predecessor node is earlier than that of the successor node and both belong to the same alignment event chain. Spatial-adjacent edges are established based on spatial adjacency relationships, and edges are connected only when the spatial nodes corresponding to two semantic nodes are directly connected and the doors, windows, and valves are allowed to connect. Cross-modal support edges are established based on modal complementarity. Edges are connected only when the modal compositions of two semantic nodes are complementary and they co-occur within the same time window, forming an initial graph structure that includes semantic nodes and decision edges. The initial graph structure is stored and indexed in a graph format. A node index table, an edge index table, and a spatial-temporal inverted index are established. The graph version number and construction time are recorded. Graph integrity verification is enabled. Parts with circular dependencies or duplicate edges are removed or merged. The self-evolving semantic decision graph is output. During the operation phase, historical records and real-time feedback are received, and node weights and edge strengths are updated at fixed time steps. The update takes into account current activation, temporal order compliance, spatial accessibility and cross-modal consistency, and penalties are imposed on edges that do not conform to spatial topology constraints. Based on evolutionary rules, the structure of the self-evolutionary semantic decision graph is modified by adding, deleting, and altering elements: When the strength of a decision edge is not lower than the new threshold in multiple consecutive updates, the decision edge is included in the stable set. When the strength of a decision edge does not exceed the deletion threshold in multiple consecutive updates, the decision edge is removed from the self-evolving semantic decision graph. When the unexplained activation of a spatial node continuously meets the new generation threshold, a new semantic node is added and candidate edges with adjacent nodes are initialized.

[0014] Optionally, the output of the corresponding early warning signal and the spatial location information of the potential hazard includes: It receives real-time input multimodal data streams, performs spatiotemporal synchronization and spatial mapping, generates real-time event sequences, and projects them onto the corresponding semantic nodes of the self-evolving semantic decision graph. Real-time activation values ​​are calculated based on the weights of semantic nodes and node state vectors, and edge collaboration values ​​are calculated based on the edge strength and connectivity between nodes. The risk score is obtained by combining the node activation values ​​and edge collaboration values. The risk score of the hazard is compared with the updated set of risk thresholds. If the score is higher than the first-level threshold, it is marked as a high-risk hazard; if the score is between the first-level and second-level thresholds, it is marked as a medium-risk hazard; and if the score is lower than the second-level threshold but higher than the minimum threshold, it is marked as a low-risk hazard. Based on the node location information and spatial topology index in the self-evolving semantic decision graph, the spatial region, channel and impact range corresponding to the hidden danger are determined, and an identification result set containing the hidden danger level, location, trigger node and trigger time is generated; When any hazard risk level in the identified results set reaches the preset alarm threshold, an early warning signal and hazard spatial location information are automatically output, and the identification results and real-time feedback records are written into the historical operation record.

[0015] The beneficial effects of this invention are: This invention achieves dynamic alignment of multi-source heterogeneous data in both temporal and causal dimensions by constructing a spatiotemporally coupled causal alignment graph. Compared with traditional methods that rely solely on time synchronization or fixed window matching, the spatiotemporal coupling mechanism of this invention can accurately identify the sequential logical relationships of events such as flames, smoke, temperature rises, and acoustic anomalies, establishing a real physical causal chain. This effectively solves the problems of misalignment and mismatch between video, audio, and environmental sensor data, significantly improving the temporal and causal consistency of hazard identification results.

[0016] This invention further introduces a spatial topology-aware attention mechanism, mapping BIM data, duct structure, sensor distribution, and door / window valve status into computable topological biases, enabling the identification process to recognize spatial structures. By dynamically updating spatial topology embedding information, the system can perceive the direction of airflow, regional connectivity, and occlusion relationships within the building in real time, achieving precise propagation and constraint of potential hazards in space. This effectively reduces false alarms and missed alarms caused by environmental changes or view obstruction, improving the system's positioning accuracy and stability in complex building environments.

[0017] This invention constructs a self-evolving semantic decision graph, enabling the system to automatically adjust its structure based on historical records and real-time feedback during long-term operation. Through dynamic optimization of node weights and edge connections, the system can learn hazard feature patterns in different scenarios and automatically generate new semantic nodes when new hazard types are discovered. This achieves continuous model evolution and cross-scenario adaptation, improving the system's continuous intelligence level and transforming fire hazard identification from passive detection to proactive perception and adaptive learning. This effectively reduces maintenance costs and extends the model's lifespan. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0019] Figure 1 This is a flowchart of a smart fire hazard identification method based on multimodal fusion proposed in this invention; Figure 2 This is a schematic diagram illustrating the construction and updating process of the self-evolving semantic decision graph of a smart fire hazard identification method based on multimodal fusion proposed in this invention. Detailed Implementation

[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0021] refer to Figure 1 and Figure 2 A smart fire hazard identification method based on multimodal fusion includes: Collect multi-source heterogeneous data from the built environment, process the multi-source heterogeneous data, and obtain multimodal raw input data; Event anchor point detection is performed on the multimodal raw input data to extract key event points for each mode, including brightness abrupt changes, sound pressure peaks, smoke concentration changes, and temperature inflection points, and to generate event sequences for each mode. Based on the event sequences of each modality, an initial event graph containing temporal and causal edges is constructed. The event graph is then jointly aligned using the spatiotemporal coupling causal alignment graph module to ensure that the multimodal events maintain consistency in temporal order and causal relationship, and outputs the spatiotemporally aligned multimodal event stream. Import BIM data, duct structure data and detector layout information to establish a spatial topology map. Based on the spatial topology awareness attention mechanism, the spatiotemporally aligned multimodal event flow is fused with the spatial topology map. Multimodal fusion features with spatial constraints are generated with topology embedding as the bias. Based on multimodal fusion features, a self-evolving semantic decision graph is constructed, semantic nodes are generated and decision edges are established to form an initial graph structure. Based on historical operation records and real-time feedback information, the node weights and edge connection relationships are dynamically adjusted to complete the self-learning and evolution update of the structure. By utilizing the self-evolving semantic decision graph, the system identifies and classifies potential hazards from real-time input multimodal data. When the identification results reach a preset risk threshold, the system outputs corresponding warning signals and spatial location information of potential hazards.

[0022] In this embodiment, the multi-source heterogeneous data includes visible light video data, infrared video data, audio data, temperature data, humidity data, smoke concentration data, and gas concentration data.

[0023] In this embodiment, the processing of multi-source heterogeneous data includes unified time synchronization, noise filtering, and standardized encoding of the multi-source heterogeneous data.

[0024] In this embodiment, generating the sequence of events for each modality includes: The data of visible light video, infrared video, audio, temperature, humidity, smoke concentration and gas concentration are denoised, de-drifted and uniformly sampled to establish a single time series index and generate a change trend sequence and change amplitude sequence for each data channel. An event prototype library is constructed based on historical calibration data. The event prototype library contains a set of constraint parameters for the duration, amplitude, rising or falling direction, duration and recovery time of sudden brightness changes, infrared radiation changes, sound pressure peaks, sudden increases in smoke concentration and temperature inflection points. For each data channel, a candidate event search is performed within the sliding time window. When the change amplitude and duration simultaneously satisfy the corresponding constraint parameter set, it is marked as a candidate event point of the channel, and the occurrence time, amplitude level, direction attribute and duration are recorded. Within the cross-modal consistency verification window, causal order consistency screening and delay distribution constraint verification are performed on candidate event points for each channel. The screening is based on a preset trigger sequence table and the allowed time difference interval. Candidate event points that are inconsistent with the trigger sequence table or exceed the time difference interval are eliminated. Conflicting candidate event points within the same window are uniquely retained according to priority rules. The selected candidate event points are summarized into event sequences by channel. The event sequences include the occurrence time, channel identifier, amplitude level, direction attribute, duration, and cross-modal consistency marker.

[0025] In this embodiment, the output spatiotemporally aligned multimodal event stream includes: The spatiotemporal coupled causal alignment graph module is invoked, taking the event sequences of each modality as input, to initialize the spatiotemporal event graph, the dual-anchor time grid, and the causal delay envelope library. The spatiotemporal coupled causal alignment graph module consists of a time-preserving unit, a causal-first inference unit, and a conflict arbitration and consistency mapping unit. The time-preserving unit performs order preservation, deduplication, and merging on event nodes within the dual-anchor time grid, deletes connections that do not meet the requirement that the source anchor point precedes the response anchor point, limits the event interval to within a preset lower and upper limit, and outputs a calibrated time connection set. The causal inference unit retrieves precursor events that satisfy minimum delay, maximum delay, and preferred delay constraints for each candidate successor event based on the causal delay envelope library, generates a candidate causal connection set, and establishes a local causal evidence table for each candidate connection. The local causal evidence table includes cross-modal consistency markers, amplitude level, direction attribute, duration, and degree of matching with delay constraints. In the spatiotemporal coupled causal alignment graph module, the counterfactual occlusion stability check is enabled. For each candidate causal connection, the predecessor occlusion and successor recalculation are performed. Within a fixed verification window, the persistence and amplitude stability of the successor event are checked to see if they are lower than the preset threshold. If they are lower than the threshold, the causal connection is confirmed. If they are higher than the threshold, the causal connection is removed, and the set of causal connections that pass the check is obtained. The conflict arbitration and consistency mapping unit performs priority sorting and one-to-one mapping on the verified causal connections. The priority is determined based on the comprehensive score of cross-modal consistency matching degree, proximity to the preferred delay, and amplitude and duration. A unique predecessor event is selected for each subsequent event, circular dependencies are cleared, a spatiotemporal coupled causal alignment graph is generated, and it is expanded into an aligned event chain in ascending time order to form a spatiotemporally aligned multimodal event flow.

[0026] In this embodiment, the step of generating spatially constrained multimodal fusion features using topological embedding as a bias includes: A spatial topology map is created based on BIM data, duct structure data and detector layout information. The spatial topology map consists of spatial nodes and connecting edges. Spatial nodes record the location, fire compartment, visible range and obstruction mark. Connecting edges record the physical connection relationship, air flow direction and door and window valve status. Topological embedding information is generated on the spatial topology graph. For each spatial node, the connectivity, airflow direction and obstacle status of adjacent nodes and adjacent connected edges are aggregated to obtain topological description information for attention calculation. The real-time status of doors, windows and valves is written into the gate control tag of the corresponding connected edge. Multimodal event streams are mapped to spatial nodes of the spatial topology graph according to time windows. The mapping is determined based on the camera's field of view, sensor coverage, and spatial proximity. Events falling into the same spatial node are weighted and aggregated according to time sequence and source confidence to form a node event set. The spatial topological awareness attention mechanism is used to fuse events of spatial nodes, specifically as follows: Based on the visible range and occlusion markers of spatial nodes, suppression weights are set for events that are not within the visible range or are occluded, and pass weights are set for events that are within the visible range and are not occluded. The influence of events propagates from upstream nodes to downstream nodes along the airflow direction. Propagation is only allowed when the doors, windows, and valves on the connected edges are open, and propagation is blocked when they are closed. Events in the downstream direction are given a higher propagation weight, while events in the upstream direction are given a suppression weight. Based on the topological description information of fire compartment boundaries, floor shafts and air duct branches, state biases are assigned to spatial nodes and connecting edges, and the biases are updated as the state of doors, windows and valves changes. After completing the visual field occlusion gating, airflow path propagation and state adaptive bias, a multimodal fusion feature with spatial constraints is output for each spatial node.

[0027] In this embodiment, the step of dynamically adjusting node weights and edge connectivity based on historical operation records and real-time feedback information includes: Semantic elements are extracted from multimodal fusion features, and semantic nodes are generated one-to-one with spatial nodes. Each semantic node contains a unique identifier, its spatial node number, a list of modal components, historical hit counts, and the most recent update time. Semantic nodes are also classified into three categories according to their source: source nodes, transit nodes, and convergence nodes. Establish initial decision edges between semantic nodes, and execute them one by one according to the following three types of rules: Based on the time-cause-effect edge of the alignment event chain established by the output, the edge is connected only when the timestamp of the predecessor node is earlier than that of the successor node and both belong to the same alignment event chain. Spatial-adjacent edges are established based on spatial adjacency relationships, and edges are connected only when the spatial nodes corresponding to two semantic nodes are directly connected and the doors, windows, and valves are allowed to connect. Cross-modal support edges are established based on modal complementarity. Edges are connected only when the modal compositions of two semantic nodes are complementary and they co-occur within the same time window, forming an initial graph structure that includes semantic nodes and decision edges. The initial graph structure is stored and indexed in a graph format. A node index table, an edge index table, and a spatial-temporal inverted index are established. The graph version number and construction time are recorded. Graph integrity verification is enabled. Parts with circular dependencies or duplicate edges are removed or merged. The self-evolving semantic decision graph is output. During the operation phase, historical records and real-time feedback are received, and node weights and edge strengths are updated at fixed time steps. The update takes into account current activation, temporal order compliance, spatial accessibility and cross-modal consistency, and penalties are imposed on edges that do not conform to spatial topology constraints. Based on evolutionary rules, the structure of the self-evolutionary semantic decision graph is modified by adding, deleting, and altering elements: When the strength of a decision edge is not lower than the new threshold in multiple consecutive updates, the decision edge is included in the stable set. When the strength of a decision edge does not exceed the deletion threshold in multiple consecutive updates, the decision edge is removed from the self-evolving semantic decision graph. When the unexplained activation of a spatial node continuously meets the new generation threshold, a new semantic node is added and candidate edges with adjacent nodes are initialized.

[0028] In this embodiment, the output of the corresponding early warning signal and spatial location information of the potential hazard includes: It receives real-time input multimodal data streams, performs spatiotemporal synchronization and spatial mapping, generates real-time event sequences, and projects them onto the corresponding semantic nodes of the self-evolving semantic decision graph. Real-time activation values ​​are calculated based on the weights of semantic nodes and node state vectors, and edge collaboration values ​​are calculated based on the edge strength and connectivity between nodes. The risk score is obtained by combining the node activation values ​​and edge collaboration values. The risk score of the hazard is compared with the updated set of risk thresholds. If the score is higher than the first-level threshold, it is marked as a high-risk hazard; if the score is between the first-level and second-level thresholds, it is marked as a medium-risk hazard; and if the score is lower than the second-level threshold but higher than the minimum threshold, it is marked as a low-risk hazard. Based on the node location information and spatial topology index in the self-evolving semantic decision graph, the spatial region, channel and impact range corresponding to the hidden danger are determined, and an identification result set containing the hidden danger level, location, trigger node and trigger time is generated; When any hazard risk level in the identified results set reaches the preset alarm threshold, an early warning signal and hazard spatial location information are automatically output, and the identification results and real-time feedback records are written into the historical operation record.

[0029] Example 1: To verify the feasibility of this invention in practice, it was applied to a smart fire protection renovation project in a science and technology park. The testing period was from May to September 2025, covering three buildings with a total construction area of ​​approximately 68,000 square meters, including office areas, laboratories, a canteen, and an underground parking lot. The test system deployed 92 high-definition visible light cameras, 18 infrared cameras, 64 environmental sensing nodes (including smoke, temperature, gas concentration, and humidity sensors), and 12 audio acquisition units. The system operated collaboratively on an edge server and a central control platform, with a real-time sampling frequency of 1Hz and a video frame rate of 15fps. The core objective of the test was to verify the recognition accuracy and adaptive capability of the three innovative modules proposed in this invention—the spatiotemporally coupled causal alignment graph, the spatial topological awareness attention mechanism, and the self-evolving semantic decision graph—in complex environments.

[0030] In practical applications, the spatiotemporal coupled causal alignment graph module first performs temporal and causal alignment on video, audio, and sensor data. By establishing event anchor points (such as sudden brightness changes, smoke concentration jumps, sound pressure peaks, and rapid temperature increases) and generating a multimodal event graph based on the causal delay relationships between events, the system can distinguish between real fire source events and non-fire source disturbances. For example, in a test conducted in the B2 level parking garage in mid-July 2025, the system performed temporal causal separation on the events of "increased vehicle exhaust concentration" and "visible light flickering" during the same time period, eliminating false fire alarm signals; the traditional system misjudged 4 times under the same conditions, while the spatiotemporal coupled causal alignment graph module only misjudged 1 time. The average alignment delay was reduced from 3.6 seconds in the traditional method to 0.8 seconds, and the event temporal consistency was improved by approximately 78%.

[0031] Subsequently, the spatial topology awareness attention mechanism module maps the aligned event flow onto the spatial topology map generated from the building BIM model. The topology map nodes contain detectors, cameras, and duct structures, while the edges record airflow direction and door / window status. The system dynamically adjusts the feature propagation direction through topology attention bias, accurately reflecting the diffusion path of smoke or hot air. In the experimental area on the third floor of Building A, when the duct valves were closed and the air conditioning external circulation was on, the spatial topology awareness attention mechanism module automatically reduced the event propagation weight between adjacent experiments, avoiding the "cross-zone false alarms" common in traditional systems. Experimental results show that the average error in hazard location using the method of this invention is 2.7 meters, a reduction of approximately 76% compared to the traditional system (11.3 meters); the false alarm rate decreased from 7.8% to 1.6%.

[0032] During a 45-day continuous test, the self-evolving semantic decision graph module continuously adjusted node weights and decision edge relationships based on historical data and real-time feedback, achieving model self-evolution and stable optimization. The system recorded 112 valid potential hazard events, including 83 triggered by simulated experiments and 29 by actual equipment anomalies. Initially, the model automatically added 6 semantic nodes after the third week, including new pattern nodes such as "laboratory odor diffusion" and "abnormal slow equipment temperature rise." Throughout the cycle, the model performed 35 graph structure updates, deleting 8 redundant edges and adding 12. As the learning process progressed, the system's recognition accuracy improved from 93.5% at the initial deployment stage to 97.3% during the stable period, achieving cross-scenario adaptation and long-term stable operation.

[0033] Table 1 Comparison of Measured Performance Data of the Smart Fire Protection System in the Science and Technology Park

[0034] As shown in Table 1, the intelligent fire hazard identification method based on multimodal fusion proposed in this invention exhibits significantly better performance than traditional systems in various scenarios. In the office area test in Building A, the accuracy rate of this invention reached 96.8%, approximately 10 percentage points higher than the traditional system, with an average response time of 1.3 seconds. This indicates that in areas with dense office personnel and frequent changes in ambient light, the spatiotemporal coupling causal alignment mechanism of this invention can effectively suppress misjudgments caused by personnel movement and changes in light. In the underground parking garage scenario on level B2, due to significant interference from exhaust gas concentration, noise, and vehicle reflections, the false alarm rate of the traditional system was high. However, this invention, through self-evolving semantic decision graphs, performs causal separation of events such as "exhaust gas smoke" and "exhaust temperature rise," maintaining an accuracy rate of 97.5%, reducing the false alarm rate to 1.5%, and controlling the average positioning error to within 3.1 meters.

[0035] In the testing of the C-building experimental area, due to the involvement of multi-source gases and high-temperature equipment, the characteristics of potential hazards were complex and rapidly changing. The results showed that the system of this invention achieved a maximum identification accuracy of 98.1%, with a false negative rate of only 0.9%, indicating its strong robustness in highly dynamic and high-risk scenarios. Testing in the catering area further verified the effectiveness of the spatial topological awareness attention mechanism. When the state of the kitchen exhaust duct valve changes, the system can automatically adjust the event propagation path according to the topological bias, avoiding false alarms across areas and maintaining an overall accuracy of 97.0%.

[0036] Overall, across 112 events in four test areas, the system of this invention achieved an average recognition accuracy of 97.3%, a false alarm rate of 1.7%, and a false negative rate of 0.9%, while the traditional system achieved an accuracy of only 87.6%, with false alarm and false negative rates of 6.4% and 4.8%, respectively. The average response time of the system of this invention was 1.42 seconds, approximately twice that of the traditional system's 3.12 seconds. These results demonstrate that the present invention achieves substantial improvements in multimodal data fusion accuracy, spatial positioning accuracy, and response speed, enabling stable operation in complex building environments and effectively supporting the real-time monitoring and early warning needs of smart fire protection.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A smart fire hazard identification method based on multimodal fusion, characterized in that, include: Collect multi-source heterogeneous data from the built environment, process the multi-source heterogeneous data, and obtain multimodal raw input data; Event anchor point detection is performed on the multimodal raw input data to extract key event points for each mode, including brightness abrupt changes, sound pressure peaks, smoke concentration changes, and temperature inflection points, and to generate event sequences for each mode. Based on the event sequences of each modality, an initial event graph containing temporal and causal edges is constructed. The event graph is then jointly aligned using the spatiotemporal coupling causal alignment graph module to ensure that the multimodal events maintain consistency in temporal order and causal relationship, and outputs the spatiotemporally aligned multimodal event stream. Import BIM data, duct structure data and detector layout information to establish a spatial topology map. Based on the spatial topology awareness attention mechanism, the spatiotemporally aligned multimodal event flow is fused with the spatial topology map. Multimodal fusion features with spatial constraints are generated with topology embedding as the bias. Based on multimodal fusion features, a self-evolving semantic decision graph is constructed, semantic nodes are generated and decision edges are established to form an initial graph structure. Based on historical operation records and real-time feedback information, the node weights and edge connection relationships are dynamically adjusted to complete the self-learning and evolution update of the structure. By utilizing the self-evolving semantic decision graph, the system identifies and classifies potential hazards from real-time input multimodal data. When the identification results reach a preset risk threshold, the system outputs corresponding warning signals and spatial location information of potential hazards.

2. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The multi-source heterogeneous data includes visible light video data, infrared video data, audio data, temperature data, humidity data, smoke concentration data, and gas concentration data.

3. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The processing of multi-source heterogeneous data includes unified time synchronization, noise filtering, and standardized coding of the multi-source heterogeneous data.

4. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The generation of each modal event sequence includes: The data of visible light video, infrared video, audio, temperature, humidity, smoke concentration and gas concentration are denoised, de-drifted and uniformly sampled to establish a single time series index and generate a change trend sequence and change amplitude sequence for each data channel. An event prototype library is constructed based on historical calibration data. The event prototype library contains a set of constraint parameters for the duration, amplitude, rising or falling direction, duration and recovery time of sudden brightness changes, infrared radiation changes, sound pressure peaks, sudden increases in smoke concentration and temperature inflection points. For each data channel, a candidate event search is performed within the sliding time window. When the change amplitude and duration simultaneously satisfy the corresponding constraint parameter set, it is marked as a candidate event point of the channel, and the occurrence time, amplitude level, direction attribute and duration are recorded. Within the cross-modal consistency verification window, causal order consistency screening and delay distribution constraint verification are performed on candidate event points for each channel. The screening is based on a preset trigger sequence table and the allowed time difference interval. Candidate event points that are inconsistent with the trigger sequence table or exceed the time difference interval are eliminated. Conflicting candidate event points within the same window are uniquely retained according to priority rules. The selected candidate event points are summarized into event sequences by channel. The event sequences include the occurrence time, channel identifier, amplitude level, direction attribute, duration, and cross-modal consistency marker.

5. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The output spatiotemporally aligned multimodal event stream includes: The spatiotemporal coupled causal alignment graph module is invoked, taking the event sequences of each modality as input, to initialize the spatiotemporal event graph, the dual-anchor time grid, and the causal delay envelope library. The spatiotemporal coupled causal alignment graph module consists of a time-preserving unit, a causal-first inference unit, and a conflict arbitration and consistency mapping unit. The time-preserving unit performs order preservation, deduplication, and merging on event nodes within the dual-anchor time grid, deletes connections that do not meet the requirement that the source anchor point precedes the response anchor point, limits the event interval to within a preset lower and upper limit, and outputs a calibrated time connection set. The causal inference unit retrieves precursor events that satisfy minimum delay, maximum delay, and preferred delay constraints for each candidate successor event based on the causal delay envelope library, generates a candidate causal connection set, and establishes a local causal evidence table for each candidate connection. The local causal evidence table includes cross-modal consistency markers, amplitude level, direction attribute, duration, and degree of matching with delay constraints. In the spatiotemporal coupled causal alignment graph module, the counterfactual occlusion stability check is enabled. For each candidate causal connection, the predecessor occlusion and successor recalculation are performed. Within a fixed verification window, the persistence and amplitude stability of the successor event are checked to see if they are lower than the preset threshold. If they are lower than the threshold, the causal connection is confirmed. If they are higher than the threshold, the causal connection is removed, and the set of causal connections that pass the check is obtained. The conflict arbitration and consistency mapping unit performs priority sorting and one-to-one mapping on the verified causal connections. The priority is determined based on the comprehensive score of cross-modal consistency matching degree, proximity to the preferred delay, and amplitude and duration. A unique predecessor event is selected for each subsequent event, circular dependencies are cleared, a spatiotemporal coupled causal alignment graph is generated, and it is expanded into an aligned event chain in ascending time order to form a spatiotemporally aligned multimodal event flow.

6. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The generation of spatially constrained multimodal fusion features using topological embedding as a bias includes: A spatial topology map is created based on BIM data, duct structure data and detector layout information. The spatial topology map consists of spatial nodes and connecting edges. Spatial nodes record the location, fire compartment, visible range and obstruction mark. Connecting edges record the physical connection relationship, air flow direction and door and window valve status. Topological embedding information is generated on the spatial topology graph. For each spatial node, the connectivity, airflow direction and obstacle status of adjacent nodes and adjacent connected edges are aggregated to obtain topological description information for attention calculation. The real-time status of doors, windows and valves is written into the gate control tag of the corresponding connected edge. Multimodal event streams are mapped to spatial nodes of the spatial topology graph according to time windows. The mapping is determined based on the camera's field of view, sensor coverage, and spatial proximity. Events falling into the same spatial node are weighted and aggregated according to time sequence and source confidence to form a node event set. The spatial topological awareness attention mechanism is used to fuse events of spatial nodes, specifically as follows: Based on the visible range and occlusion markers of spatial nodes, suppression weights are set for events that are not within the visible range or are occluded, and pass weights are set for events that are within the visible range and are not occluded. The influence of events propagates from upstream nodes to downstream nodes along the airflow direction. Propagation is only allowed when the doors, windows, and valves on the connected edges are open, and propagation is blocked when they are closed. Events in the downstream direction are given a higher propagation weight, while events in the upstream direction are given a suppression weight. Based on the topological description information of fire compartment boundaries, floor shafts and air duct branches, state biases are assigned to spatial nodes and connecting edges, and the biases are updated as the state of doors, windows and valves changes. After completing the visual field occlusion gating, airflow path propagation and state adaptive bias, a multimodal fusion feature with spatial constraints is output for each spatial node.

7. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The process of dynamically adjusting node weights and edge connectivity based on historical operation records and real-time feedback information includes: Semantic elements are extracted from multimodal fusion features, and semantic nodes are generated one-to-one with spatial nodes. Each semantic node contains a unique identifier, its spatial node number, a list of modal components, historical hit counts, and the most recent update time. Semantic nodes are also classified into three categories according to their source: source nodes, transit nodes, and convergence nodes. Establish initial decision edges between semantic nodes, and execute them one by one according to the following three types of rules: Based on the alignment event chain establishment time-causal edge of the output, the edge is connected only when the timestamp of the predecessor node is earlier than that of the successor node and both belong to the same alignment event chain. Spatial-adjacent edges are established based on spatial adjacency relationships, and edges are connected only when the spatial nodes corresponding to two semantic nodes are directly connected and the doors, windows, and valves are allowed to connect. Cross-modal support edges are established based on modal complementarity. Edges are connected only when the modal compositions of two semantic nodes are complementary and they co-occur within the same time window, forming an initial graph structure that includes semantic nodes and decision edges. The initial graph structure is stored and indexed in a graph format. A node index table, an edge index table, and a spatial-temporal inverted index are established. The graph version number and construction time are recorded. Graph integrity verification is enabled. Parts with circular dependencies or duplicate edges are removed or merged. The self-evolving semantic decision graph is output. During the operation phase, historical records and real-time feedback are received, and node weights and edge strengths are updated at fixed time steps. The update takes into account current activation, temporal order compliance, spatial accessibility and cross-modal consistency, and penalties are imposed on edges that do not conform to spatial topology constraints. Based on evolutionary rules, the structure of the self-evolutionary semantic decision graph is modified by adding, deleting, and altering elements: When the strength of a decision edge is not lower than the new threshold in multiple consecutive updates, the decision edge is included in the stable set. When the strength of a decision edge does not exceed the deletion threshold in multiple consecutive updates, the decision edge is removed from the self-evolving semantic decision graph. When the unexplained activation of a spatial node continuously meets the new generation threshold, a new semantic node is added and candidate edges with adjacent nodes are initialized.

8. The intelligent fire hazard identification method based on multimodal fusion according to claim 1, characterized in that, The output of corresponding early warning signals and spatial location information of potential hazards includes: It receives real-time input multimodal data streams, performs spatiotemporal synchronization and spatial mapping, generates real-time event sequences, and projects them onto the corresponding semantic nodes of the self-evolving semantic decision graph. Real-time activation values ​​are calculated based on the weights of semantic nodes and node state vectors, and edge collaboration values ​​are calculated based on the edge strength and connectivity between nodes. The risk score is obtained by combining the node activation values ​​and edge collaboration values. The risk score of the hazard is compared with the updated set of risk thresholds. If the score is higher than the first-level threshold, it is marked as a high-risk hazard; if the score is between the first-level and second-level thresholds, it is marked as a medium-risk hazard; and if the score is lower than the second-level threshold but higher than the minimum threshold, it is marked as a low-risk hazard. Based on the node location information and spatial topology index in the self-evolving semantic decision graph, the spatial region, channel and impact range corresponding to the hidden danger are determined, and an identification result set containing the hidden danger level, location, trigger node and trigger time is generated; When any hazard risk level in the identified results set reaches the preset alarm threshold, an early warning signal and hazard spatial location information are automatically output, and the identification results and real-time feedback records are written into the historical operation record.

Citation Information

Cited By

  • Quadruped robot fire-fighting inspection study and judgment disposal closed-loop method and system and medium

    CN121988000A