Fire-fighting early warning system based on image data relevance
By analyzing the correlation of multi-source image data and optimizing resource scheduling, a fire warning system is constructed, which solves the shortcomings of traditional fire warning systems in fire monitoring and resource scheduling, achieves more accurate fire feature identification and scientific emergency response decisions, and improves the stability of the system and the reliability of warnings.
Patent Information
- Application Number
- CN202510682998.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional fire warning systems have shortcomings in fire monitoring and resource scheduling. They are unable to fully obtain fire scene information, lack the ability to deeply analyze image data, and statically allocate computing resources and lack a collaborative mechanism. This leads to insufficient timeliness and accuracy in warnings, and makes it impossible to accurately predict fire development trends and provide scientific emergency response decisions.
A fire warning system based on image data correlation is adopted. Through the multi-source image acquisition module, correlation feature analysis module, spatiotemporal dynamic modeling module and resource optimization scheduling library, cross-data source correlation analysis of multi-source heterogeneous image data is realized, spatiotemporal dynamic feature fusion space is constructed, and real-time fire warning instructions and emergency response decisions are generated through a hierarchical reinforcement learning framework.
It improves the accuracy and comprehensiveness of fire feature identification, enhances the accuracy of fire development trend prediction and the scientific nature of emergency response, ensures stable and efficient operation of the system, and provides more targeted early warning and emergency response decisions.
Smart Images

Figure CN120636124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fire warning, and in particular to a fire warning system based on image data correlation. Background Art
[0002] Fire accidents occur frequently, posing a huge threat to people's lives and property. Traditional fire warning systems have many limitations in terms of function and performance, making it difficult to meet the growing demand for fire safety.
[0003] Early fire warnings primarily rely on single-type detectors, such as smoke and temperature sensors. These detectors monitor only the smoke concentration or temperature fluctuations generated by a fire, sounding an alarm when a preset threshold is reached. However, this approach has significant shortcomings. For one thing, single sensors are susceptible to interference from environmental factors. For example, in environments like kitchens and boiler rooms, normal fluctuations in smoke or temperature can cause false alarms. Furthermore, they fail to fully capture the fire scene, lacking effective monitoring of key information such as the fire's development trends, the distribution and movement of people, and changes in building structure. This makes it difficult to provide accurate and comprehensive early warnings.
[0004] With the development of video surveillance technology, some fire warning systems have incorporated video surveillance equipment. However, these systems often simply use video footage for post-event review, failing to fully exploit the key information contained in the video images. Faced with complex fire scenarios, due to a lack of in-depth analysis and processing capabilities for image data, they are unable to promptly extract effective fire-related features from the massive amount of video data, such as the dynamic texture changes of flames, the diffusion pattern and speed of smoke, etc., significantly compromising the timeliness and accuracy of warnings.
[0005] Furthermore, existing fire warning systems also suffer from flaws in resource scheduling. Computing resource allocation is typically static, failing to account for the dynamic changes in tasks under different fire scenarios and the load balancing of edge nodes. This can result in some computing resources being idle during a fire, while critical tasks cannot be processed promptly due to insufficient resources, hindering the efficiency of early warning and emergency response. Furthermore, there is a lack of effective coordination mechanisms between different types of sensors and cameras, preventing them from fully leveraging their respective strengths and achieving data complementarity and integration.
[0006] Existing fire risk assessment methods are mostly based on empirical evidence or simple statistical models, failing to fully integrate real-time environmental parameters and dynamic information from the fire scene. This approach makes it difficult to accurately predict fire development trends and potential damage, resulting in a lack of scientific basis for emergency response decisions and a lack of targeted action, potentially delaying optimal firefighting and rescue efforts.
[0007] Furthermore, traditional fire warning systems are inadequate for detecting structural anomalies in complex structures. They are unable to detect structural deformation in real time during a fire, making it difficult to identify potential structural safety hazards in advance. This inability to provide crucial building safety information for evacuation and rescue operations increases the difficulty and risk of rescue operations. Summary of the Invention
[0008] The purpose of the present invention is to provide a fire warning system based on image data correlation to solve the problems raised in the above background technology.
[0009] To achieve the above-mentioned object, the present invention provides the following technical solution: a fire warning system based on image data correlation, the system comprising:
[0010] Multi-source image acquisition module: used to collect multi-source heterogeneous image data of fire scenes in real time;
[0011] Correlation feature analysis module: performs cross-data source correlation analysis on the multi-source heterogeneous image data to generate feature vectors of each image modality, including smoke diffusion morphological features, flame dynamic texture features, thermal distribution gradient features, personnel movement trajectory patterns, and abnormal building structure contours;
[0012] Spatiotemporal dynamic modeling module: constructs a multi-dimensional feature fusion space based on the spatiotemporal graph convolutional network, maps the feature vectors to a unified spatiotemporal dynamic coordinate system, and generates associated spatiotemporal feature tensors;
[0013] Resource optimization and scheduling library: Build a two-layer collaborative library of computing resource pools and fire risk assessment models. The two-layer collaborative library stores dynamic task scheduling rules, edge node load balancing strategies, and camera-sensor device topology relationship maps.
[0014] Intelligent early warning control module: Based on the associated spatiotemporal feature tensor and the resource optimization scheduling library, it generates real-time fire warning instructions and emergency response decisions through a hierarchical reinforcement learning framework.
[0015] Preferably, the multi-source heterogeneous image data includes infrared thermal imaging sequences, visible light video streams, smoke concentration distribution maps, lidar point clouds and personnel positioning trajectory data;
[0016] The performing cross-data source correlation analysis on the multi-source heterogeneous image data includes:
[0017] The infrared thermal imaging sequence is decomposed into frequency domain components of thermal distribution by using discrete cosine transform, and the smoke diffusion trend is extracted by using time series association algorithm;
[0018] Using a dense attention network to segment abnormal areas of the building structure on the lidar point cloud, and combining it with a motion compensation algorithm to predict the risk of structural collapse;
[0019] The escape path features are extracted from the personnel movement trajectory data using a spatiotemporal alignment network, and the personnel gathering density pattern is matched based on a dynamic programming algorithm.
[0020] Preferably, the cross-data source correlation analysis further includes:
[0021] A three-dimensional convolutional network is used to model the spatiotemporal evolution of the dynamic texture of the flame on the visible light video stream, and a generative adversarial network is used to enhance the detailed features of the flame edge;
[0022] A graph embedding algorithm is used to model the topological dependency of the smoke diffusion path on the smoke concentration distribution map, and the probability distribution of the smoke source position is inferred based on a hidden Markov model.
[0023] Preferably, the construction of a two-layer collaborative library of a computing resource pool and a fire risk assessment model includes:
[0024] Based on the dynamic bandwidth allocation protocol and the heterogeneous architecture of edge nodes, a federated learning framework is used to build a distributed computing resource pool;
[0025] Based on historical fire data and real-time environmental parameter distribution, a spatiotemporal variational autoencoder is used to generate the latent variable representation of the fire risk assessment model.
[0026] Preferably, the hierarchical reinforcement learning framework adopts a multi-agent collaborative architecture, including:
[0027] The state space is defined as the joint encoding of the associated spatiotemporal feature tensor and the computational load of the resource pool, and the action space is the decision sequence of image acquisition priority allocation and processing frequency;
[0028] The decision conflicts of each agent are coordinated through the asynchronous advantage actor-critic algorithm, and a multi-level reward function is designed based on delay constraints.
[0029] Preferably, the hierarchical reinforcement learning framework further includes:
[0030] A distributed double-delayed deep deterministic policy gradient algorithm is used to implement multi-node collaborative policy updates, and curriculum learning technology is combined to optimize global training efficiency.
[0031] The communication bandwidth limitation is transformed into a dynamic constraint of the policy network through an adaptive penalty mechanism.
[0032] Preferably, the system further comprises:
[0033] Performing real-time deformation detection on the abnormal contour of the building structure and using a spectral segmentation algorithm to identify structural vulnerable areas that exceed a preset deformation threshold;
[0034] When abnormal deformation is detected, the emergency avoidance mechanism of the intelligent early warning control module is triggered, generating evacuation path planning instructions and collaborative monitoring signals of adjacent sensor equipment.
[0035] Preferably, the emergency avoidance mechanism adopts a bidirectional spatiotemporal graph network modeling of the human-environment interaction relationship, including:
[0036] Extract the escape direction features from the forward propagation path, and extract the obstacle distribution impact features from the reverse propagation path;
[0037] The optimal evacuation plan is generated through multi-head attention mechanism fusion, and the risk area weights of the fire risk assessment model are updated.
[0038] Preferably, the system further comprises:
[0039] A smoke diffusion prediction model was constructed based on historical fire event data and real-time meteorological parameters. A gated spatiotemporal convolutional network was used to fuse the environmental diffusion factor and dynamically adjust the resolution threshold of the image acquisition module.
[0040] Preferably, the smoke diffusion prediction model further includes:
[0041] Based on the conditional generative adversarial network, multi-sensor noise parameters are fused to generate a smoke diffusion confidence index and embed it into the real-time fire warning instruction.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] In terms of data acquisition and analysis, the system is equipped with a multi-source image acquisition module that can collect heterogeneous image data from multiple sources in real time, including infrared thermal imaging sequences, visible light video streams, smoke concentration distribution maps, lidar point clouds, and personnel location trajectory data. This multi-source data acquisition approach greatly enriches the information sources and can more comprehensively reflect the fire scene compared to traditional single or small-scale data collection. The correlation feature analysis module uses specialized algorithms to deeply explore the characteristics of different data types. For infrared thermal imaging sequences, the discrete cosine transform is used to decompose the frequency domain components of the thermal distribution, and a time series correlation algorithm is used to accurately extract smoke diffusion trends. For lidar point clouds, a dense attention network is used to segment abnormal areas of the building structure, and a motion compensation algorithm is used to predict the risk of structural collapse. This cross-data source correlation analysis generates multiple feature vectors, such as smoke diffusion morphological characteristics and flame dynamic texture characteristics, which provide a precise basis for subsequent fire diagnosis and significantly improve the accuracy and comprehensiveness of fire feature identification.
[0044] The spatiotemporal dynamic modeling module utilizes a spatiotemporal graph convolutional network to construct a multidimensional feature fusion space, mapping each feature vector to a unified spatiotemporal dynamic coordinate system to generate a correlated spatiotemporal feature tensor. This process breaks down the barriers between different data modalities and achieves deep fusion of features across the spatiotemporal dimensions. This enables the system to comprehensively grasp the dynamic changes in fire scenarios and accurately capture the spatiotemporal patterns of fire development, providing more valuable information for subsequent early warning and decision-making, effectively improving the accuracy of fire development trend predictions.
[0045] The resource optimization and scheduling library constructs a two-layer collaborative library of computing resource pools and fire risk assessment models. Based on a dynamic bandwidth allocation protocol and a heterogeneous architecture of edge nodes, a federated learning framework is employed to construct a distributed computing resource pool. This dynamically allocates computing resources based on each node's task requirements and network conditions, achieving load balancing among edge nodes. Furthermore, based on historical fire data and real-time environmental parameter distributions, a spatiotemporal variational autoencoder is used to generate latent variable representations of the fire risk assessment model, enhancing the scientific nature and accuracy of fire risk assessments. This resource optimization and scheduling mechanism not only improves the utilization efficiency of the system's computing resources and ensures stable and efficient system operation, but also provides a more scientific basis for fire risk assessments, making early warning and emergency response decisions more targeted.
[0046] The intelligent early warning and control module generates real-time fire warning instructions and emergency response decisions through a hierarchical reinforcement learning framework, based on the associated spatiotemporal feature tensor and resource optimization scheduling library. The hierarchical reinforcement learning framework adopts a multi-agent collaborative architecture, rationally defines the state space and action space, and uses advanced algorithms such as the asynchronous dominant actor-critic algorithm to coordinate agent decision conflicts, effectively improving the timeliness and accuracy of decision-making. In addition, the system also has a variety of auxiliary functions, such as real-time deformation detection of abnormal contours of building structures. Once an anomaly is detected, the emergency avoidance mechanism is triggered, and evacuation path planning instructions and collaborative monitoring signals of nearby sensor equipment are generated to ensure personnel safety; a smoke diffusion prediction model is constructed, and the resolution threshold of the image acquisition module is dynamically adjusted according to the prediction results. The smoke diffusion confidence index is generated and embedded in the real-time fire warning instructions, further improving the reliability and accuracy of the warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a working principle diagram of the fire warning system based on image data correlation according to the present invention;
[0048] Figure 2 Flowchart for feature analysis of visible light video stream and smoke concentration distribution map;
[0049] Figure 3 Flowchart for building structure anomaly detection and emergency triggering;
[0050] Figure 4Schematic diagram of the human-environment interaction analysis in emergency avoidance mechanism. DETAILED DESCRIPTION
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0052] See also Figures 1-4 The present invention provides a fire warning system based on image data correlation, and specific embodiments include:
[0053] Example 1:
[0054] The system's multi-source heterogeneous image data includes infrared thermal imaging sequences, visible light video streams, smoke concentration distribution maps, lidar point clouds, and personnel location trajectory data. For image data of different modalities, cross-data source correlation analysis uses specific algorithms and models to extract key features. The specific operations are as follows:
[0055] Infrared thermal imaging sequences acquire temperature distribution information within a scene through non-contact measurement, providing a visual representation of areas of thermal anomaly. Processing of these sequences involves two steps: frequency domain component decomposition and smoke diffusion trend extraction.
[0056] The discrete cosine transform (DCT) is used to perform frequency domain decomposition on each frame of the infrared thermal imaging sequence. As an orthogonal transform, the discrete cosine transform can convert the pixel values of the image from the spatial domain to the frequency domain, where the low-frequency components correspond to the overall brightness of the image and the slowly changing background, and the high-frequency components correspond to the edges, details, and mutation information of the image. In firefighting scenarios, the heat generated by the fire can cause the temperature in the local area to rise significantly, forming abnormal high-frequency components of the thermal distribution. Through the discrete cosine transform, the frequency domain components of the thermal distribution can be decomposed into components of different frequencies, which facilitates the subsequent analysis of the spatial distribution and temporal variation of thermal anomalies.
[0057] After completing the frequency domain component decomposition, the smoke diffusion trend is extracted using a time series association algorithm. The time series association algorithm is used to analyze the correlation between thermal anomaly areas in infrared thermal imaging sequences at different times. Specifically, the high-frequency components of thermal anomalies extracted from each frame are used as sample points in the time series. By calculating parameters such as the spatial position overlap and temperature change gradient of sample points at adjacent moments, the corresponding relationship between thermal anomaly areas at different times is established. For example, if a certain area shows high-frequency thermal components in multiple consecutive frames of images, and the position of this area gradually moves in a certain direction over time, it can be determined that this area is the source of smoke diffusion, and its movement direction is the trend of smoke diffusion. In this way, the dynamic trend characteristics of smoke diffusion can be extracted from infrared thermal imaging sequences, providing a key basis for fire early warning.
[0058] LiDAR (Light Detection and Ranging) generates three-dimensional point cloud data of a target scene by emitting laser pulses and measuring their reflected signals. This data accurately describes the geometric shape and spatial position of building structures. Processing LiDAR point clouds involves two steps: segmenting abnormal areas of the building structure and predicting the risk of structural collapse.
[0059] A Dense Attention Network (DAN) is used to segment abnormal building structures from LiDAR point clouds. A DAN is a deep learning-based image segmentation model that effectively captures local details and global structural features in point cloud data through densely connected network layers and an attention mechanism. During training, the network is fed LiDAR point cloud data labeled with normal building structure areas and abnormal areas. The network automatically identifies and segments abnormal areas by learning features such as point cloud density, geometry, and spatial distribution. For example, a building structure may experience abnormalities such as cracked walls and deformed beams and columns during a fire. These abnormalities can cause a decrease in point cloud density and a sudden change in geometry in the corresponding area. The DAN can learn these features to accurately segment abnormal structural areas.
[0060] After the abnormal area segmentation is completed, the motion compensation algorithm is combined to predict the risk of structural collapse. The motion compensation algorithm is used to analyze the dynamic changes of the lidar point cloud in the time series. By calculating the position offset, shape change and other parameters of the abnormal area in the point cloud data at adjacent moments, it compensates for the noise interference caused by equipment movement or scene changes, thereby more accurately capturing the development trend of structural abnormalities. Specifically, for each segmented abnormal area, its motion trajectory in continuous multi-frame point cloud data is established, and the trajectory's dynamic parameters such as speed and acceleration are calculated. If the motion trajectory of an abnormal area shows that its displacement is accelerating over time, it indicates that the structural stability of the area is declining sharply and there is a high risk of collapse. In this way, the collapse risk of building structures can be predicted from lidar point cloud data, providing an important reference for emergency response decisions.
[0061] Personnel movement trajectory data is captured using indoor positioning technologies (such as RFID, Bluetooth, and Wi-Fi), recording the real-time location and movement paths of personnel in firefighting scenarios. Processing this data involves extracting escape path features and matching personnel density patterns.
[0062] The spatio-temporal alignment network (Spatio-TemporalAlignmentNetwork) is used to extract escape path features. The spatio-temporal alignment network is a deep learning model for spatio-temporal sequence data. It can effectively capture the temporal sequence features and spatial topological features of personnel movement trajectories through convolution operations in the time dimension and graph neural networks in the spatial dimension. Specifically, the movement trajectory of personnel is represented as a spatio-temporal graph, where nodes represent the positions of personnel at different times, and edges represent the time transfer relationship between positions. The spatio-temporal alignment network extracts key nodes (such as turning points, stop points), path length, movement speed and other features of the personnel escape path by performing convolution operations on the spatio-temporal graph, and encodes these features into low-dimensional feature vectors for subsequent analysis and matching.
[0063] The personnel gathering density pattern is matched based on the dynamic programming algorithm. The dynamic programming algorithm decomposes the complex optimization problem into multiple sub-problems and uses the optimal solutions of the sub-problems to construct the global optimal solution. It is suitable for matching and identifying personnel gathering density patterns. Specifically, the characteristic vector of the personnel movement trajectory is used as input, and the gathering density pattern is defined as the spatial distribution density threshold of the personnel position in different time periods (such as high-density area, medium-density area, low-density area). Through the dynamic programming algorithm, the subsequence that best matches the preset gathering density pattern is found in the personnel movement trajectory. For example, when a subsequence of densely distributed personnel positions is detected in a certain area at multiple consecutive moments, the area is judged to be a personnel gathering area, and the characteristics such as the duration of the gathering and the density change trend are analyzed. In this way, the escape path characteristics and gathering density pattern can be extracted from the personnel movement trajectory data, providing a basis for the formulation of personnel evacuation strategies in fires.
[0064] After processing the infrared thermal imaging sequence, lidar point cloud and personnel motion trajectory data separately, it is necessary to further establish a cross-data source correlation analysis mechanism to integrate the feature information of different modal data. Specifically, through timestamp alignment technology, the processing results of different data sources are mapped to a unified time coordinate system to ensure that the feature vectors of each modality at the same moment have temporal consistency. Then, a feature fusion algorithm (such as concatenation, weighted summation, etc.) is used to fuse the smoke diffusion trend characteristics, structural collapse risk characteristics, escape path characteristics and aggregation density pattern characteristics to generate a composite feature vector that comprehensively reflects the safety status of the fire scene. This cross-data source correlation analysis can make full use of the complementarity of multimodal data and improve the accuracy and reliability of fire warning.
[0065] Example 2:
[0066] This embodiment further processes the visible light video stream and the smoke concentration distribution map, extracting the dynamic texture features of the flame and the topological dependency of the smoke diffusion path through a specific algorithm and model. The specific operations are as follows:
[0067] Visible light video streams capture real-time color image sequences of fire scenes through cameras, visually displaying the shape, color, and dynamic characteristics of flames. Processing of this video stream involves two steps: modeling the spatiotemporal evolution of the fire scene and enhancing edge detail features.
[0068] A 3D Convolutional Neural Network (3DCNN) was used to model the spatiotemporal evolution of the dynamic texture of flames. A 3D Convolutional Neural Network (3DCNN) is a deep learning model specifically designed for processing spatiotemporal sequence data. By performing convolution operations simultaneously in the temporal and spatial dimensions, it can capture the temporal dependencies between adjacent frames in a video sequence and the spatial features within a single frame. In firefighting scenarios, the dynamic texture of flames manifests as spatiotemporal features such as color changes (e.g., red, orange, and yellow), edge jitter, and shape expansion and contraction. Specifically, the visible light video stream is divided into consecutive video segments, each containing several consecutive frames, which serve as the input to the 3D Convolutional Network. The first layer of the network performs a two-dimensional convolution on a single frame to extract spatial texture features (e.g., the gradient of the flame edge and the color distribution). Subsequent layers convolve the feature maps of adjacent frames in the temporal dimension using a 3D convolution kernel to extract dynamic features in the temporal dimension (e.g., the rate of change of the flame texture over time and the flickering frequency). Through layer-by-layer abstraction of a multi-layer network, the final output is a feature vector that can characterize the spatiotemporal evolution law of the dynamic texture of the flame, such as the average movement speed of the flame and the periodicity of the texture change.
[0069] A generative adversarial network (GAN) is used to enhance flame edge detail. A GAN consists of a generator and a discriminator. The generator is responsible for generating fake data that closely resembles real samples, while the discriminator distinguishes between real samples and fake data. In the flame edge detail enhancement task, the input is an original image from a visible light video stream. The generator's goal is to generate a flame image with clearer edge details, while the discriminator determines whether the input image is the original image or the enhanced image generated by the generator. The specific training process is as follows: the generator uses an encoder-decoder architecture, first encoding the original image into a low-dimensional feature vector through convolutional layers, then decoding it through deconvolutional layers to generate the enhanced image. The discriminator, using a multi-layer convolutional network, assesses the input image based on metrics such as edge clarity and texture detail. During training, the generator and discriminator are continuously optimized through adversarial training. The generator gradually learns to enhance the contrast and sharpen the edge contours of the flame, making the flame edge details more prominent in the generated image. For example, the boundary between the flame and the background is clearly distinguished, and the texture layering within the flame is enhanced. This processing can effectively improve the recognizability of flame characteristics and provide a more accurate basis for subsequent fire warnings.
[0070] The smoke concentration distribution map is generated using data collected by the smoke sensor network, reflecting the smoke concentration distribution in various areas of the fire scene. Processing this distribution map involves two steps: topological dependency modeling and inferring the probability distribution of smoke source locations.
[0071] A graph embedding algorithm is used to model the topological dependencies of smoke diffusion paths. Graph embedding algorithms are used to map graph-structured data into a low-dimensional vector space while preserving the topological structure and associations of the nodes in the graph. In the smoke concentration distribution graph, each smoke sensor node is considered a node in the graph. Edges between nodes represent the spatial proximity between sensors (for example, edges are established between sensor nodes with a Euclidean distance less than a preset threshold). The edge weights represent the smoke concentration gradient between adjacent sensor nodes. Using graph embedding algorithms (such as DeepWalk and Node2Vec), random walk sampling is performed on the graph structure to generate a node sequence. The Skip-gram model is then used to map the node sequence into a low-dimensional embedding vector, ensuring that nodes that are adjacent in the graph are also close to each other in the vector space. These embedding vectors capture the topological dependencies between smoke sensor nodes. For example, if two sensor nodes are adjacent in the graph and have a large smoke concentration gradient, their embedding vectors will be close in the low-dimensional space, indicating a strong correlation between them in the smoke diffusion path. In this way, the topological structure of the smoke concentration distribution map can be converted into a computable vector representation, which facilitates the subsequent analysis of the path characteristics of smoke diffusion.
[0072] The probability distribution of smoke source locations is inferred based on a hidden Markov model (HMM). A hidden Markov model is a statistical model that describes the relationship between a hidden state sequence and an observable output sequence. In a smoke diffusion scenario, the hidden state represents the location of the smoke source, and the observable output represents the smoke concentration data at each sensor node. Specifically, assuming that the smoke source location follows a certain probability distribution (such as a Gaussian distribution), each hidden state corresponds to a smoke diffusion model that describes the distribution of smoke concentration from the smoke source at that location to surrounding sensor nodes. The parameters of the hidden Markov model (such as the state transition probability matrix and the observation probability matrix) are trained using historical smoke concentration data. When real-time smoke concentration data is acquired, a forward-backward algorithm is used to calculate the posterior probability of each hidden state (i.e., the smoke source location) given the observed data, thereby inferring the probability distribution of the smoke source location. For example, if the calculated posterior probability of a certain area is significantly higher than that of other areas, then that area is more likely to be the smoke source. This inference method can combine the spatiotemporal characteristics of smoke concentration distribution to achieve probabilistic estimation of the smoke source location, providing key information for fire fighting and personnel evacuation.
[0073] After completing the processing of the visible light video stream and the smoke concentration distribution map, it is necessary to integrate the correlation between its feature vector and the smoke diffusion trend feature, structural collapse risk feature, and personnel movement trajectory feature extracted in Example 1 across data sources. Specifically, through feature standardization processing, the numerical ranges of different modal feature vectors are unified to the same interval to avoid feature fusion deviations caused by dimensional differences. Then, an attention mechanism is used to assign weights to feature vectors of different modalities, and the size of the weight reflects the importance of the modal feature in the current fire scene. For example, when obvious flame dynamic texture features are detected in the visible light video stream, a higher weight is given to the flame feature vector; when the smoke concentration distribution map shows that the probability distribution of the smoke source position is relatively concentrated, a higher weight is given to the smoke feature vector. Through this weighted fusion method, a comprehensive feature vector containing multi-dimensional information such as flame, smoke, heat, structure, and personnel is generated, thereby more comprehensively describing the fire risk status of the fire scene.
[0074] 3D Convolutional Network Architecture Design: A lightweight 3DCNN architecture, such as the C3D network, is used, consisting of multiple 3D convolutional layers, pooling layers, and fully connected layers. The 3D convolution kernel size is set to 3×3×3 (3×3 spatial dimension, 3 temporal dimension) to capture the spatiotemporal features of three adjacent image frames. The pooling layer uses a 2×2×2 max pooling to reduce the spatial and temporal dimensions of the feature map, thus reducing computational effort.
[0075] Loss function of the generative adversarial network: The loss function of the least squares generative adversarial network (LSGAN) is adopted to avoid the gradient vanishing problem that occurs in the traditional GAN training process, enabling the generator to generate high-quality flame edge enhanced images more stably.
[0076] Parameter settings of the graph embedding algorithm: In the Node2Vec algorithm, the step size of the random walk is set to 80, the number of walks for each node is set to 10, and the dimension of the embedding vector is set to 128 to balance the computational efficiency and feature representation capability of the model.
[0077] State space of the hidden Markov model: The fire scene is divided into grid-like discrete areas. Each area corresponds to a hidden state of the hidden Markov model. The total number of states is determined by the size of the scene, usually dozens to hundreds, to ensure that the model can accurately describe the distribution of smoke source locations.
[0078] 5. Interaction with other system modules
[0079] The processed visible light video stream features and smoke concentration distribution map features are transmitted via an interface to the spatiotemporal dynamic modeling module, where they serve as input data for generating the associated spatiotemporal feature tensor. Specifically, the flame dynamic texture features and smoke diffusion path features, as modality-specific feature vectors, are mapped together with other modal feature vectors into a unified spatiotemporal dynamic coordinate system. Through multi-layer operations of the spatiotemporal graph convolutional network, multidimensional feature fusion and spatiotemporal correlation modeling are achieved. These features are also input into the hierarchical reinforcement learning framework of the intelligent early warning and control module as components of the state space, generating real-time fire warning instructions and emergency response decisions. For example, when the flame dynamic texture features indicate an accelerated flame spread and a concentrated smoke source location probability distribution, the intelligent early warning and control module adjusts the image acquisition priority and processing frequency based on these features, strengthening monitoring of high-risk areas and generating corresponding warning level upgrade instructions.
[0080] Example 3:
[0081] When constructing a two-layer collaborative library of computing resource pool and fire risk assessment model in the resource optimization and scheduling library, the specific implementation method is as follows:
[0082] Based on the dynamic bandwidth allocation protocol and edge node heterogeneous architecture, a federated learning framework is used to build a distributed computing resource pool to achieve efficient management and dynamic scheduling of computing resources.
[0083] The dynamic bandwidth allocation protocol is used to dynamically adjust the communication bandwidth allocation between edge nodes according to the real-time computing task requirements and network status. In the fire warning system, the image data processing of different modalities (such as frequency domain decomposition of infrared thermal imaging sequences and abnormal area segmentation of lidar point clouds) has different requirements for computing resources and communication bandwidth. For example, the three-dimensional convolutional network processing of visible light video streams requires higher computing power, while the graph embedding algorithm processing of smoke concentration distribution maps has higher requirements for memory capacity. The dynamic bandwidth allocation protocol establishes an optimization model for bandwidth allocation by monitoring the computing load of each edge node (such as CPU utilization, memory occupancy) and the bandwidth utilization of the network link in real time. Suppose the edge node set is The current computational load of each node is L(n i )∈[0,1], the link bandwidth utilization is B(n i ,n j )∈[0,1], the goal is to maximize the overall bandwidth utilization of the system while satisfying the latency constraints of each task. Through heuristic algorithms (such as greedy algorithms) or optimization algorithms (such as linear programming), the bandwidth allocation ratio between nodes is dynamically adjusted to ensure that high-priority tasks (such as flame dynamic texture feature extraction) have sufficient communication resources and avoid computing task delays caused by insufficient bandwidth.
[0084] A heterogeneous edge node architecture refers to a system in which edge nodes have varying hardware configurations (such as CPU models, GPU computing power, and memory capacity) and software environments (such as different versions of deep learning frameworks). This heterogeneity results in varying processing efficiency for different types of computing tasks. For example, GPU-equipped nodes are more suitable for processing compute-intensive tasks such as 3D convolutional networks, while CPU-equipped nodes are more suited for logic control tasks. To fully leverage the heterogeneous nature of edge nodes, a federated learning framework is employed to build a distributed computing resource pool. Federated learning is a distributed machine learning paradigm that allows each node to train models locally, sharing only model parameter updates rather than raw data. This enables collaborative model optimization while protecting data privacy. During the computational resource pool construction process, each edge node's hardware resource information (such as peak computing power and memory size) and software capabilities (such as supported algorithm and model types) are first registered and modeled to form a node capability profile. Then, based on the real-time computing task requirements (such as algorithm type, computing power requirements, and data size), a task scheduling algorithm allocates the task to the most suitable edge node. For example, the three-dimensional convolutional network task of visible light video stream is assigned to nodes equipped with high-performance GPUs, and the spatiotemporal alignment network task of personnel motion trajectory data is assigned to nodes with stronger CPU performance to minimize the processing delay and energy consumption of the task.
[0085] Under the federated learning framework, each edge node performs local model training based on locally stored fire scene data (such as infrared thermal imaging sequences and lidar point clouds), and regularly sends model parameter updates to the central server. The central server aggregates the parameter updates of each node (such as using the FedAvg algorithm), generates a global model update, and distributes the updated model to each node. This collaborative training mechanism enables the computing resource pool to continuously optimize the model performance of each node and improve the accuracy of feature parsing without exposing the original data. For example, in the associated feature parsing module, each node collaboratively optimizes the parameters of the dense attention network through federated learning to improve the accuracy of segmenting abnormal areas of the lidar point cloud.
[0086] Based on historical fire data and real-time environmental parameter distribution, a spatiotemporal variational autoencoder is used to generate the latent variable representation of the fire risk assessment model, so as to achieve low-dimensional feature modeling and dynamic assessment of fire risk.
[0087] The input data includes multi-source image data of historical fire events (such as infrared thermal imaging sequences, visible light video streams), corresponding fire parameters (such as fire location, burning area, smoke concentration peak) and real-time environmental parameters (such as temperature, humidity, wind speed, and wind direction). First, the historical fire data is cleaned and labeled to remove noise data and incorrectly labeled samples; then, feature extraction is performed on the multi-source image data (such as using the algorithms in Example 1 and Example 2 to extract features such as smoke diffusion trends and flame dynamic textures) to form a high-dimensional feature vector. Real-time environmental parameters are collected in real time by sensors, normalized, and spliced with historical feature vectors as input to the spatiotemporal variational autoencoder.
[0088] The spatiotemporal variational autoencoder is a generative model that combines time series modeling with variational autoencoders. Its architecture consists of three parts: encoder, latent variable space, and decoder.
[0089] Encoder: Bidirectional long short-term memory network (Bi-LSTM) is used as the main structure of the encoder to capture the spatiotemporal dependencies of the input data. Bi-LSTM extracts contextual information of the past and future moments through forward and backward neurons respectively, and outputs a hidden state sequence Then, the hidden state is mapped to the mean μ and logarithmic variance logσ of the latent variable through the fully connected layer. 2 , that is, μ,logσ 2 =FC(h T ) where FC() represents the fully connected layer, h T The hidden state at the last moment.
[0090] Latent variable space: From the normal distribution N(μ,σ) through reparameterization techniques 2 ) to obtain the latent variable z, namely z = μ + σ⊙∈,∈~N(0,I), where ⊙ represents element-wise multiplication and ∈ is a random noise vector from a standard normal distribution. The latent variable z is a low-dimensional representation of the fire risk assessment model, with a dimension K << D. It can capture the core characteristics of fire risk (such as fire development trend, smoke diffusion speed, and evacuation difficulty).
[0091] Decoder: A multi-layer perceptron (MLP) is used as a decoder to decode the latent variable z into a reconstructed feature vector Right now By minimizing the reconstruction loss (such as mean square error) and KL divergence loss, the parameters of the spatiotemporal variational autoencoder are optimized so that the latent variable space can accurately capture the distribution characteristics of the input data.
[0092] The generated latent variable z serves as the core input of the fire risk assessment model and is used to calculate the fire risk level of each area. Specifically, by training a classifier (such as a logistic regression model) or a regression model, the latent variable is mapped to a fire risk score (such as 0-100 points) or a risk level (such as low, medium, and high). In addition, latent variables can also be used for spatiotemporal prediction of fire risks. By analyzing the changing trend of latent variables in time series, the risk status at future moments can be predicted. For example, if the value of a certain dimension in the latent variable continues to increase, it indicates that the corresponding risk factor (such as the speed of smoke diffusion) is intensifying. The system can issue an early warning to prompt relevant personnel to take preventive measures.
[0093] The computing resource pool and fire risk assessment model work together through dynamic task scheduling rules, edge node load balancing strategies and device topology relationship maps.
[0094] Dynamic task scheduling rules: Based on the latent variable representation output by the fire risk assessment model, the risk level of the current fire scene is determined and the priority and allocation strategy of computing tasks are dynamically adjusted. For example, when the risk level is high, computing resources are prioritized for fire feature extraction tasks (such as flame dynamic texture analysis and smoke source location inference) to ensure that monitoring data in high-risk areas can be processed in a timely manner.
[0095] Edge node load balancing strategy: By real-time monitoring of the computing load of each edge node and the resource demand characteristics in the hidden variables (such as computing power demand and memory demand), load balancing algorithms (such as polling algorithm and minimum load first algorithm) are used to redistribute tasks to avoid the situation where some nodes are overloaded while other nodes are idle, thereby improving the utilization of computing resources.
[0096] Camera-Sensor Device Topology Map: This constructs a topology map of cameras, sensor devices, their spatial locations, and communication links to optimize data collection and transmission paths. By combining risk area information from latent variables, the system dynamically adjusts camera acquisition angles and sensor sampling frequencies. For example, it increases the acquisition frame rate for cameras in high-risk areas and increases the number of smoke concentration sampling times for adjacent sensors to obtain more intensive monitoring data.
[0097] Optimizing communication efficiency of federated learning: To reduce the communication overhead between edge nodes and the central server, model parameter compression techniques (such as gradient quantization and sparse update) are adopted to transmit only key parameter update information.
[0098] Training data enhancement of spatiotemporal variational autoencoders: Due to the relative scarcity of historical fire data, data enhancement techniques (such as time series shifting and noise injection) are used to expand the training dataset and improve the generalization ability of the model.
[0099] Compatibility design for heterogeneous nodes: In the management system of the computing resource pool, containerization technology (such as Docker) is used to encapsulate different algorithm models to ensure that they can run on edge nodes with different hardware configurations, reducing compatibility issues caused by heterogeneous architectures.
[0100] Example 4:
[0101] The hierarchical reinforcement learning framework uses a multi-agent collaborative architecture to define state and action spaces, coordinate decision conflicts, and optimize policy updates to achieve real-time fire warning instructions and emergency response decisions. The following describes its implementation in detail based on specific application scenarios:
[0102] 1. Definition of State Space and Action Space
[0103] In a fire warning system, the state space is defined as a joint encoding of a correlated spatiotemporal feature tensor and the resource pool's computational load. Taking a commercial complex fire monitoring scenario as an example, the correlated spatiotemporal feature tensor integrates multimodal data features: dynamic flame texture features extracted from visible light video streams (such as flame area growth rate and edge jitter frequency), thermal distribution gradient features extracted from infrared thermal imaging sequences (such as the diffusion velocity of high-temperature areas), abnormal building structural contours extracted from lidar point clouds (such as the deformation and displacement of beams and columns on a specific floor), and escape path features extracted from occupant trajectory data (such as the density of occupants in a specific evacuation corridor). These features are mapped to a unified spatiotemporal coordinate system through a spatiotemporal dynamic modeling module, forming a multidimensional tensor containing spatial locations (such as floor numbers and area coordinates) and timestamps (such as millisecond sampling intervals). The resource pool's computational load collects real-time computing power usage at each edge node, such as the remaining video memory capacity of a GPU node processing a 3D convolutional network and the CPU utilization of a CPU node executing a dynamic programming algorithm. This is represented in vector form as load vector = [node 1 load, node 2 load, …, node N load].
[0104] The action space is defined as a sequence of decisions regarding image acquisition priorities and processing frequencies. For example, if a suspected fire area is detected on a floor, the system can generate the following actions: raise the acquisition priority of the camera corresponding to that area to the highest level and increase the processing frequency of its video stream from the default 10 frames per second to 30 frames per second; simultaneously, reduce the acquisition frame rate of cameras in non-high-risk areas to conserve computing resources. Actions are executed through dynamic task scheduling rules in the resource optimization scheduling library, such as sending instructions to edge nodes to adjust data acquisition parameters and reallocate computing task queues.
[0105] 2. Coordination of Multi-Agent Decision-Making Conflicts
[0106] In a multi-agent collaborative architecture, each agent corresponds to an independent computing node or functional module, such as an agent responsible for flame signature analysis, an agent responsible for personnel trajectory analysis, or an agent responsible for resource scheduling. Each agent makes decisions based on its own local observations, which may lead to resource competition or task priority conflicts. For example, in a factory workshop fire scenario, agent A (flame detection) requests an increase in the camera acquisition frequency in a certain area to capture flame dynamics, while agent B (structural safety) simultaneously requests the same edge node to process a structural collapse prediction task based on a lidar point cloud. The computing resource demands of both agents may exceed the processing capacity of the node.
[0107] At this point, the system uses an asynchronous dominant actor-critic algorithm to mediate conflicting decisions. The A3C algorithm allows agents to update their policies asynchronously on different threads, enabling collaboration by sharing global model parameters. The specific process is as follows: Each agent generates an action (such as requesting resources or adjusting task priorities) based on its current state (e.g., associating spatiotemporal feature tensors with its own observed local load) and calculates an advantage function for the action (a measure of its superiority relative to the average policy). The global model collects the gradient updates of each agent and updates the policy network parameters using a weighted average, enabling the agents to gradually learn to prioritize high-risk tasks given limited resources. For example, in the aforementioned conflict, the system dynamically adjusts the agent action weights using the A3C algorithm based on latent variables in the fire risk assessment model (e.g., rapid flame spread and increased structural deformation). This prioritizes the high-priority requests of the flame detection agent while delaying or downgrading the tasks of the structural safety agent to prevent critical task failures due to resource overload.
[0108] 3. Design of Multi-level Reward Function
[0109] The reward function is designed based on delay constraints to guide the agent to make decisions that meet real-time requirements. Taking the high-rise residential fire scenario as an example, the reward function is divided into three levels:
[0110] Emergency Response Reward: High rewards are awarded for emergency actions that impact human life safety, such as triggering emergency avoidance mechanisms and generating evacuation route instructions. For example, if the agent detects that the density of people gathering on a certain floor exceeds a threshold and smoke concentration is rising rapidly, it will immediately initiate an emergency broadcast for that floor, earning a +100 reward.
[0111] Resource Efficiency Reward: A medium reward is awarded for actions that utilize computing resources effectively (e.g., balancing node loads and reducing redundant data processing). For example, if an agent reduces the average load of edge nodes from 85% to 70% by adjusting the acquisition frequency of different cameras without affecting key feature extraction, it will receive a +50 reward.
[0112] Delay Penalty: Negative rewards are assigned to actions that exceed a delay threshold (e.g., fire feature extraction takes longer than the maximum allowed time for an alert). For example, if the flame dynamic texture analysis task takes longer than 2 seconds, causing a delay in sending the alert command, the agent will receive a -80 penalty.
[0113] Guided by multi-level reward functions, the intelligent agent needs to comprehensively consider the urgency of the task and the availability of resources when making decisions. For example, it prioritizes optimizing resource efficiency in low-risk scenarios and increases resource consumption to ensure response speed in high-risk scenarios.
[0114] 4. Multi-node collaborative strategy update
[0115] The system uses a distributed double-delayed deep deterministic policy gradient algorithm to achieve multi-node collaborative policy updates. Taking a multi-floor warehouse park as an example, edge nodes on different floors each deploy an agent responsible for fire monitoring and resource scheduling in their local area. The D4PG algorithm achieves collaboration through the following mechanisms:
[0116] Distributed training: Each node's agent trains its strategy based on local observation data (such as camera video streams and sensor data on that floor) and sends gradient updates to a central parameter server. The central server aggregates the updates from each node, generates global policy parameters, and synchronizes them to all nodes, enabling agents to learn cross-regional collaborative strategies. For example, if a fire breaks out on one floor, agents on adjacent floors can use global policy parameters to sense the risk of fire spread and adjust the camera acquisition priority on that floor in advance, preparing for subsequent fire spread monitoring.
[0117] Dual-Delay Network: This introduces a delayed update mechanism for the target and evaluation networks to reduce variance in policy training and improve convergence stability. The evaluation network calculates the value function of the current policy in real time, while the target network periodically copies parameters from the evaluation network to calculate the target value, avoiding training oscillations caused by frequent parameter updates.
[0118] Curriculum Learning Technology: Training begins with a gradual transition from simple to complex scenarios. Initially, only the agent on a single floor is activated to train its decision-making capabilities in simple fire scenarios. As training progresses, multi-floor linkage scenarios are gradually added, requiring the agent to learn to coordinate cross-regional resource scheduling and early warning responses. For example, the agent is first trained to optimize camera acquisition frequency during a single-floor fire, and then trained to coordinate multi-node processing tasks when the fire spreads to adjacent floors, achieving global efficiency optimization.
[0119] 5. Dealing with Communication Bandwidth Constraints
[0120] An adaptive penalty mechanism transforms communication bandwidth limitations into dynamic constraints for the policy network. For example, in an urban fire monitoring system, multiple edge nodes connect to a central server via a bandwidth-limited communication network. Simultaneously transmitting large amounts of feature data or model parameters can lead to network congestion. The adaptive penalty mechanism dynamically adjusts the policy network's loss function based on real-time bandwidth utilization. When bandwidth utilization exceeds a threshold (e.g., 80%), it penalizes high-data-transfer actions (e.g., full-resolution video streaming) generated by the agent, reducing the probability of selecting such actions. When bandwidth utilization is low, the penalty is relaxed to allow for higher data transfer requirements.
[0121] Taking fire monitoring in a hospital ward as an example, the execution process of the hierarchical reinforcement learning framework is as follows:
[0122] State Perception: The multi-source image acquisition module captures real-time visible light video streams from the ward area (showing smoke coming from electrical appliances in a particular ward), infrared thermal imaging sequences (detecting an abnormally high temperature in that ward), and occupant location trajectory data (indicating a gathering of people in a nearby corridor). The associated feature analysis module extracts the initial texture features of the flames, thermal diffusion trends, and occupant escape path characteristics, generating an associated spatiotemporal feature tensor.
[0123] Resource scheduling decisions: Agent A (Flame Detection) detects a suspected fire source and requests an increase in the ward camera's acquisition frequency to 25 frames per second, allocating GPU nodes for 3D convolutional network processing. Agent B (Evacuation) detects a gathering of people and requests CPU nodes to calculate evacuation routes. At this point, the edge node's GPU load is 60% and its CPU load is 70%, both within the threshold. The system approves both requests, and after execution, the loads rise to 80% and 85%, respectively.
[0124] Conflict Resolution and Reward Feedback: As the CPU load approaches the threshold, the system uses the A3C algorithm to adjust the priority of subsequent actions. For example, it reduces the frequency of non-urgent environmental parameter collection tasks to free up CPU resources. The agent receives an Emergency Response Reward of +100 and a Resource Efficiency Reward of +30 for promptly responding to fire source detection and evacuation.
[0125] Strategy update: Each node agent stores the experience data (state, action, reward, next state) of this decision into the experience replay buffer, regularly updates the policy network parameters through the D4PG algorithm, and learns to allocate resources more efficiently in similar scenarios.
[0126] Example 5:
[0127] The system has expanded its functional modules in terms of building structure safety monitoring, emergency avoidance mechanism, and smoke diffusion prediction. The following describes its implementation in detail based on specific scenarios:
[0128] Taking a high-rise building fire scene as an example, LiDAR point cloud data collects real-time 3D coordinate information of the building's exterior and interior structures. The system uses a spectral segmentation algorithm to detect abnormal structural contours in real time. The specific process is as follows:
[0129] LiDAR point cloud data is divided into multiple structural units (such as walls, beams, columns, and floor slabs). Each unit corresponds to a node in the graph structure. Edges between nodes represent spatial proximity (for example, nodes are connected if the distance between them is less than 1 meter). Edge weights are calculated based on point cloud density and geometric smoothness, reflecting the continuity and stability of the structural unit. Under normal conditions, the point cloud of each structural unit is evenly distributed, and the edge weights in the graph structure are high and stable. However, when the structure deforms (such as cracking walls or bending beams), the point cloud density in the corresponding area decreases, and the edge weights drop significantly.
[0130] The spectral graph segmentation algorithm calculates the eigenvalues and eigenvectors of the graph's Laplacian matrix, partitioning the graph into multiple subgraphs and identifying structurally vulnerable areas where edge weights fall below a preset deformation threshold (e.g., 30% of the initial weight). For example, if a load-bearing column on a floor slightly bends under the intense heat of a fire, the corresponding point cloud region will show a discretization trend. The spectral graph segmentation algorithm detects a sudden drop in the edge weight in this region to 25% of the initial value, identifying it as a structurally vulnerable area and triggering an alert.
[0131] When abnormal structural deformation is detected, the system triggers the emergency avoidance mechanism of the intelligent early warning control module. Taking the risk of structural floor collapse in a shopping mall fire as an example, the emergency avoidance mechanism models the human-environment interaction relationship through a bidirectional spatiotemporal graph network:
[0132] Forward propagation path: Extracts escape direction features. Individual location trajectory data shows that people in a certain area are moving toward the stairwell, with 80% directional consistency, indicating a concentrated escape direction. The forward layer of the bidirectional spatiotemporal graph network uses convolution operations to capture features such as the speed and directional angle of movement in the area, generating an escape direction vector.
[0133] Backward propagation path: Extracts obstacle distribution and impact features. LiDAR point cloud data indicates an obstacle on the escape path, caused by collapsed shelves, occupying 60% of the aisle width. The backward layer uses a graph neural network to analyze the degree of conflict between the obstacle's spatial position and the occupant's movement trajectory, generating an obstacle impact vector.
[0134] The system uses a multi-head attention mechanism to fuse the escape direction vector and the obstacle impact vector to generate an optimal evacuation plan. For example, if the system calculates that the original escape route is blocked by an obstacle, resulting in a 40% decrease in efficiency, it will generate new evacuation instructions: directing people to the side exits and updating the route guidance in real time through emergency broadcasts. Simultaneously, the system sends collaborative monitoring signals to nearby sensor devices, such as increasing the sampling frequency of smoke concentration sensors in the area to 5 times per second, to ensure real-time monitoring of environmental changes.
[0135] When building a smoke diffusion prediction model, taking a fire scene in a chemical park as an example, the model is based on historical fire event data (such as the smoke diffusion range under different wind speeds) and real-time meteorological parameters (such as the current wind speed of 5m / s and the wind direction of southeast by south), and uses a gated spatiotemporal convolutional network to integrate environmental diffusion factors (such as wind speed and temperature gradient). The model input includes:
[0136] Historical smoke diffusion speed series (diffusion distance per minute in the past 30 minutes);
[0137] Real-time weather data (wind speed, wind direction, air humidity);
[0138] Building layout characteristics (such as the blocking effect of building height and spacing on airflow).
[0139] The gated spatiotemporal convolutional network uses temporal gating units and spatial convolution kernels to capture the temporal dependence of smoke diffusion (e.g., increased diffusion velocity over time) and spatial correlation (e.g., changes in diffusion direction due to obstruction by buildings). The model outputs predicted smoke concentrations in each area for the next 10 minutes, which are used to dynamically adjust the resolution threshold of the image acquisition module. For example, if the prediction indicates that smoke will spread northwest to the storage area, the system automatically increases the resolution of the cameras in that area from 1080P to 4K to clearly capture the movement details of smoke particles, while simultaneously reducing the resolution of cameras in non-diffusion paths to 720P to conserve storage resources.
[0140] Based on the conditional generative adversarial network, the noise parameters of multiple sensors (such as the temperature measurement error of the infrared sensor and the point cloud density fluctuation of the lidar) are integrated to generate the smoke diffusion confidence index. Taking the forest fire scene as an example, the input of the conditional generative adversarial network includes:
[0141] Condition variables: real-time meteorological parameters and sensor types (such as smoke sensor numbered S1);
[0142] Noise variation: The statistical distribution of the historical measurement errors of each sensor (such as the mean and variance of Gaussian noise).
[0143] The generator generates simulated smoke diffusion paths based on the conditional variables and noise variables. The discriminator compares the differences between real sensor data and simulated data, forcing the generator to output a diffusion pattern that is closer to reality. The resulting confidence index is expressed as a value between 0 and 1, with higher values indicating greater reliability of the smoke diffusion prediction. For example, when the confidence index for a certain area is 0.9, the system embeds the smoke diffusion prediction results for that area into real-time fire warning instructions, prompting firefighters, "There is a high probability that smoke will spread to the northeast, and special precautions are required." If the confidence index is only 0.5, it is marked as "The prediction is uncertain, and additional on-site monitoring is recommended."
[0144] Taking the underground parking lot fire scenario as an example, the collaborative process of the system's functional modules is as follows:
[0145] Structural detection triggers an early warning: The lidar point cloud detects that the deformation displacement of a load-bearing column reaches 5 cm (exceeding the preset threshold of 3 cm). The spectral segmentation algorithm identifies the area where the column is located as a structurally fragile area and sends an emergency signal to the intelligent early warning control module.
[0146] Emergency evacuation mechanism response: A bidirectional spatiotemporal network analysis of personnel location data revealed that people in the parking lot were moving along the main corridor toward the exit, but a burning vehicle had created an obstruction in the middle of the corridor. The system generated an evacuation plan: using a broadcast to direct people to evacuate in two separate groups through the side emergency corridors, it also sent instructions to the nearby sprinkler system to increase water flow in the area to suppress the spread of the fire.
[0147] Smoke Diffusion Prediction and Data Collection Optimization: A gated spatiotemporal convolutional network, combined with the enclosed environment of the underground parking lot and real-time ventilation data (wind speed 2 m / s, blowing toward the exit), predicted that smoke would fill the main corridor within 5 minutes. The system upgraded the resolution of cameras near the exit to the highest level to monitor smoke diffusion in real time. Based on the conditions, the system generated a confidence index (0.85) for the adversarial network output to confirm the reliability of the prediction. The system then sent a comprehensive warning instruction to the fire command center, including the smoke diffusion path and evacuation status.
[0148] Risk model weight update: After the emergency avoidance mechanism is executed, the system updates the risk area weights of the fire risk assessment model through a multi-head attention mechanism based on the actual data of personnel evacuation speed and smoke diffusion. For example, the risk weights of structurally fragile areas and high smoke concentration areas are increased by 20% and 15% respectively, so that these areas can be given priority in subsequent monitoring.
[0149] Dynamically adjust the deformation detection threshold: The deformation threshold for spectrum segmentation is automatically adjusted based on the building type (e.g., steel structure, concrete structure) and the duration of the fire. For example, steel structures deform faster at high temperatures, so the threshold is set at 20% of the initial value, while the threshold for concrete structures is set at 30%.
[0150] Real-time correction of evacuation plans: The bidirectional spatiotemporal graph network updates the human-environment interaction characteristics every 10 seconds. If new obstacles or gathering points are detected, the optimal path is immediately recalculated and the instructions are updated.
[0151] Priority strategy for sensor collaboration: Sensors adjacent to hazardous areas have the highest priority, followed by sensors on escape routes, and finally peripheral monitoring sensors, to ensure the frequency and accuracy of key data collection.
Claims
1. A fire warning system based on image data correlation, characterized in that: include: Multi-source image acquisition module: used to collect multi-source heterogeneous image data of fire scenes in real time; Correlation feature analysis module: performs cross-data source correlation analysis on the multi-source heterogeneous image data to generate feature vectors of each image modality, including smoke diffusion morphological features, flame dynamic texture features, thermal distribution gradient features, personnel movement trajectory patterns, and abnormal building structure contours; Spatiotemporal dynamic modeling module: constructs a multi-dimensional feature fusion space based on the spatiotemporal graph convolutional network, maps the feature vectors to a unified spatiotemporal dynamic coordinate system, and generates associated spatiotemporal feature tensors; Resource optimization and scheduling library: Build a two-layer collaborative library of computing resource pools and fire risk assessment models. The two-layer collaborative library stores dynamic task scheduling rules, edge node load balancing strategies, and camera-sensor device topology relationship maps. Intelligent early warning control module: Based on the associated spatiotemporal feature tensor and the resource optimization scheduling library, it generates real-time fire warning instructions and emergency response decisions through a hierarchical reinforcement learning framework.
2. The fire warning system according to claim 1, characterized in that: The multi-source heterogeneous image data includes infrared thermal imaging sequences, visible light video streams, smoke concentration distribution maps, lidar point clouds and personnel positioning trajectory data; The performing cross-data source correlation analysis on the multi-source heterogeneous image data includes: The infrared thermal imaging sequence is decomposed into frequency domain components of thermal distribution by using discrete cosine transform, and the smoke diffusion trend is extracted by using time series association algorithm; Using a dense attention network to segment abnormal areas of the building structure on the lidar point cloud, and combining it with a motion compensation algorithm to predict the risk of structural collapse; The escape path features are extracted from the personnel movement trajectory data using a spatiotemporal alignment network, and the personnel gathering density pattern is matched based on a dynamic programming algorithm.
3. The fire warning system according to claim 2, characterized in that: The cross-data source correlation analysis further includes: A three-dimensional convolutional network is used to model the spatiotemporal evolution of the dynamic texture of the flame on the visible light video stream, and a generative adversarial network is used to enhance the detailed features of the flame edge; A graph embedding algorithm is used to model the topological dependency of the smoke diffusion path on the smoke concentration distribution map, and the probability distribution of the smoke source position is inferred based on a hidden Markov model.
4. The fire warning system according to claim 1, characterized in that: The two-layer collaborative library of the computing resource pool and the fire risk assessment model is constructed, including: Based on the dynamic bandwidth allocation protocol and the heterogeneous architecture of edge nodes, a federated learning framework is used to build a distributed computing resource pool; Based on historical fire data and real-time environmental parameter distribution, a spatiotemporal variational autoencoder is used to generate the latent variable representation of the fire risk assessment model.
5. The fire warning system according to claim 1, characterized in that: The hierarchical reinforcement learning framework adopts a multi-agent collaborative architecture, including: The state space is defined as the joint encoding of the associated spatiotemporal feature tensor and the computational load of the resource pool, and the action space is the decision sequence of image acquisition priority allocation and processing frequency; The decision conflicts of each agent are coordinated through the asynchronous advantage actor-critic algorithm, and a multi-level reward function is designed based on delay constraints.
6. The fire warning system according to claim 5, characterized in that: The hierarchical reinforcement learning framework also includes: A distributed double-delayed deep deterministic policy gradient algorithm is used to implement multi-node collaborative policy updates, and curriculum learning technology is combined to optimize global training efficiency. The communication bandwidth limitation is transformed into a dynamic constraint of the policy network through an adaptive penalty mechanism.
7. The fire warning system according to claim 1, characterized in that: The system further comprises: Performing real-time deformation detection on the abnormal contour of the building structure and using a spectral segmentation algorithm to identify structural vulnerable areas that exceed a preset deformation threshold; When abnormal deformation is detected, the emergency avoidance mechanism of the intelligent early warning control module is triggered, generating evacuation path planning instructions and collaborative monitoring signals of adjacent sensor equipment.
8. The fire warning system according to claim 7, characterized in that: The emergency avoidance mechanism uses a bidirectional spatiotemporal graph network to model the human-environment interaction relationship, including: Extract the escape direction features from the forward propagation path, and extract the obstacle distribution impact features from the reverse propagation path; The optimal evacuation plan is generated through multi-head attention mechanism fusion, and the risk area weights of the fire risk assessment model are updated.
9. The fire warning system according to claim 1, characterized in that: The system further comprises: A smoke diffusion prediction model was constructed based on historical fire event data and real-time meteorological parameters. A gated spatiotemporal convolutional network was used to fuse the environmental diffusion factor and dynamically adjust the resolution threshold of the image acquisition module.
10. The fire warning system according to claim 9, characterized in that: The smoke diffusion prediction model also includes: Based on the conditional generative adversarial network, multi-sensor noise parameters are fused to generate a smoke diffusion confidence index and embed it into the real-time fire warning instruction.
Citation Information
Patent Citations
Multi-sensor collaborative intelligent fire early warning system and method thereof
CN116386247A
Intelligent detection method and system for water bleeding and coal injection of blast-furnace tuyere
CN118839291A
Edge calculation method based on AI
CN119046010A
Fire remote monitoring and multistage early warning mechanism integrated system
CN119445796A
Fire accident investigation method, device, equipment, storage medium and product
CN119963386A
Cited By
Fire prevention and control method and system using computer vision
CN121147856A
Fire prevention and control method and system using computer vision
CN121147856B
Internet of Things fire-fighting facility early warning method and system based on data processing
CN121214676A
An internet of things fire-fighting facility early warning method and system based on data processing
CN121214676B
VCSEL epitaxial wafer warping degree detection method
CN121297720A