A fire smoke image early identification method based on multi-modal fusion

By utilizing multimodal fusion technology and the synchronous acquisition and dynamic optimization of visible light, infrared, and photoacoustic spectral data, the real-time and adaptability issues in early fire identification have been resolved. This has enabled accurate identification of smoke characteristics and fire source location, thereby improving the accuracy and real-time performance of fire monitoring.

CN119964080BActive Publication Date: 2025-12-05HANGZHOU ZIPENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510040940.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-12-05
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing fire monitoring technologies suffer from problems in early fire identification, such as poor real-time performance, limited coverage, weak adaptability to complex environments, poor data fusion, and insufficient prediction of fire source location and spread path. In particular, they are difficult to effectively capture smoke signals and adapt to dynamic background interference in the case of multimodal data.

Method used

Employing multimodal fusion technology, through the synchronous acquisition and time alignment of visible light, infrared and photoacoustic spectral data, combined with a spatiotemporal tracking module, low-rank matrix factorization, lightweight self-attention mechanism and sparse feature enhancement algorithm, the shared feature pool is dynamically optimized to achieve real-time and accurate identification of smoke features and fire source location.

Benefits of technology

It significantly improves the accuracy and real-time performance of early fire identification, reduces the false alarm and missed alarm rates, enhances the accuracy of fire source location and the ability to predict smoke diffusion paths, and is suitable for complex fire monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964080B_ABST
    Figure CN119964080B_ABST
Patent Text Reader

Abstract

The application discloses a kind of early identification methods of fire smoke image based on multi-modal fusion, comprising the following steps: S1, acquisition multi-modal data and carry out time alignment;S2, construct visible light branch, infrared branch and photoacoustic branch, generate multi-modal feature and store into shared feature pool;S3, embedding space-time tracking module in shared feature pool, generate global feature representation;S4, generate dynamic background template using low-rank matrix decomposition algorithm, separate and remove background noise;S5, by light self-attention mechanism in shared feature pool fusion multi-modal feature and compensate missing features;S6, by sparse feature enhancement algorithm enhances early smoke signal, and applies multi-task classification network to fire state classification and smoke diffusion path prediction;S7, by dynamic feedback mechanism updates shared feature pool and space-time tracking module.The application utilizes multi-modal fusion, space-time tracking and dynamic feedback method, realizes the early accurate identification of fire smoke.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fire monitoring technology, and in particular to an early identification method for fire smoke images based on multimodal fusion. Background Technology

[0002] Fire, as a common disaster, not only poses a significant threat to life and property but also has long-term destructive effects on the environment. Traditional fire monitoring methods mainly rely on devices such as smoke sensors, temperature sensors, or single optical cameras. These methods typically suffer from poor real-time performance, limited coverage, and weak adaptability to complex environments. Especially in the early stages of a fire, smoke signals are often sparse and inconspicuous, making it difficult for traditional single-modal monitoring devices to effectively identify them, resulting in delayed fire alarms and high false alarm rates. As fire monitoring scenarios become increasingly complex, there is an urgent need for a technical solution that can integrate information from multiple data sources to achieve high-precision fire monitoring.

[0003] In recent years, multimodal data fusion technology has been increasingly applied in the field of fire monitoring. Multimodal technology integrates and analyzes data from different sources such as visible light, infrared spectroscopy, sound, and gas sensors to overcome the problem of insufficient information from single-modal data. However, existing multimodal fusion technologies still have some significant limitations. First, the temporal differences and spatial inconsistencies between data acquisition devices can easily lead to feature alignment deviations during data fusion, affecting the data integration effect. Second, due to the diversity and complexity of multimodal data itself, existing technologies lack fine-grained processing mechanisms for the characteristics of each modality during the fusion process, which may result in the dilution or even loss of key information, making it difficult to comprehensively reflect fire characteristics. Furthermore, multimodal data contains a certain degree of noise and missing data, especially in complex environments (such as industrial sites or high-humidity environments). Data missing data can significantly impact monitoring results, and existing compensation methods are mostly static rules, making it difficult to adapt to dynamic environmental changes in real time.

[0004] Another technical bottleneck lies in the fact that smoke characteristics in the early stages of a fire are often sparse and change rapidly. Existing technologies typically use static image processing algorithms or traditional machine learning methods, which are weak at capturing early, sparse smoke signals and are prone to missed or false alarms due to noise interference. Furthermore, traditional methods struggle to effectively model smoke diffusion paths and changes in the location of the fire source. Fire monitoring not only requires accurate identification of the fire's occurrence but also real-time location of the fire source and prediction of its future diffusion trends, which is crucial for optimizing firefighting strategies and improving emergency response efficiency. However, existing technologies lack effective models and technical support for fire source location and diffusion path prediction.

[0005] In existing technologies, some studies attempt to model fire characteristics using time series analysis models, such as Long Short-Term Memory (LSTM) networks. However, these methods typically have limited capabilities in modeling time series data, especially when dealing with multimodal data, failing to effectively capture the dynamic interactions between different modalities. Other studies have attempted to introduce attention mechanisms to improve the efficiency of feature selection and fusion, but their high complexity and computational cost make them difficult to implement in resource-constrained scenarios. Furthermore, most current methods rely on static background modeling techniques for background noise removal, which are ill-suited to dynamic background interference in fire scenarios (such as changes in lighting and wind flow).

[0006] Therefore, how to provide an early identification method for fire smoke images based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] One objective of this invention is to propose an early identification method for fire smoke images based on multimodal fusion. This invention utilizes advanced methods such as multimodal fusion, spatiotemporal tracking, low-rank matrix factorization, and dynamic feedback to comprehensively utilize visible light, infrared, and photoacoustic spectral data, achieving accurate early identification of fire smoke. Through sparse feature enhancement algorithms and a lightweight self-attention mechanism, this invention effectively improves the smoke signal capture capability and solves the problems of complex background interference and data gaps. It possesses advantages such as strong real-time performance, high identification accuracy, and strong environmental adaptability, making it suitable for complex fire monitoring scenarios such as industrial plants and large public buildings, and significantly reducing the false alarm and missed alarm rates.

[0008] An early identification method for fire smoke images based on multimodal fusion according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect multimodal data using a visible light camera, an infrared camera, and a photoacoustic spectroscopy sensor, and synchronize the multimodal data based on timestamps.

[0010] S2. Construct visible light branch, infrared branch and photoacoustic branch, respectively extract smoke diffusion morphology, temperature gradient and dynamic change features of gas composition from the multimodal data, generate multimodal features and store them in the shared feature pool;

[0011] S3. Embed a spatiotemporal tracking module in the shared feature pool to capture the variation patterns of the multimodal features on the time axis and spatial distribution, and generate a global feature representation with spatiotemporal correlation characteristics;

[0012] S4. Based on the global feature representation, a low-rank matrix factorization algorithm is used to generate a dynamic background template, separate and remove background noise, and dynamically optimize the multimodal features in the shared feature pool.

[0013] S5. By using a lightweight self-attention mechanism, multimodal features are fused in the shared feature pool and missing features are compensated to generate a complete multimodal feature vector.

[0014] S6. Based on the complete multimodal feature vector, the early smoke signal is enhanced by a sparse feature enhancement algorithm, and a multi-task classification network is applied to classify the fire status and predict the smoke diffusion path.

[0015] S7. Based on the deviation of real-time environmental data, update the shared feature pool and spatiotemporal tracking module through a dynamic feedback mechanism to optimize the accuracy of fire source positioning.

[0016] Optionally, S2 specifically includes:

[0017] S21. Acquire visible light image data using a visible light camera and construct a visible light branch. The visible light image data is used to obtain the diffusion morphology characteristics of the smoke, including the edge features, diffusion speed, and diffusion direction of the smoke.

[0018] S22. Infrared image data is collected by an infrared camera to construct an infrared branch. The infrared image data is used to obtain the temperature gradient characteristics of the smoke area, including the temperature change of the smoke diffusion area and the intensity of the hot spot area.

[0019] S23. Collect gas composition data through a photoacoustic spectral sensor and construct a photoacoustic branch. The gas composition data is used to obtain the dynamic change characteristics of the gas in the smoke.

[0020] S24. Standardize the features extracted from each branch to generate multimodal features and store them in the shared feature pool.

[0021] Optionally, S3 specifically includes:

[0022] S31. Construct a spatiotemporal tracking module, which is embedded in a shared feature pool. The spatiotemporal tracking module includes a temporal convolutional network and a spatiotemporal attention mechanism. The temporal convolutional network is used to capture the change patterns of multimodal features in the time dimension, and the spatiotemporal attention mechanism is used to capture the distribution characteristics of multimodal features in the spatial dimension.

[0023] S32. Apply a temporal convolutional network to perform temporal modeling on the multimodal features, generating temporal feature representations at each time step:

[0024] T(t)=σ(W t *X(t)+b t );

[0025] Where T(t) represents the temporal feature representation at time step t, σ represents the activation function, and Wt The weights represent the temporal convolution kernel weights, * represents the convolution operation, X(t) represents the multimodal feature input at time step t, and b represents the multimodal feature input at time step t. t Represents the bias vector;

[0026] S33. Apply a spatiotemporal attention mechanism to weight the multimodal features in the spatial dimension to generate spatial feature representations for each spatial location:

[0027]

[0028] Where A(s) represents the spatial feature representation of spatial location s, Q(s) represents the query matrix of spatial location s, K(s) represents the key matrix of spatial location s, V(s) represents the value matrix of spatial location s, Softmax represents the normalization operation, and d k Indicates the dimension of the key vector;

[0029] S34. Fuse the temporal feature representation T(t) and the spatial feature representation A(s) to generate a global feature representation with spatiotemporal correlation:

[0030]

[0031] Where G(t,s) represents a global feature representation with spatiotemporal correlation, and α, β, and γ represent weight parameters. This represents the tensor product operation, used to capture the interaction between temporal and spatial features.

[0032] Optionally, S4 specifically includes:

[0033] S41. Based on the global feature representation, the global feature representation is decomposed into a low-rank matrix L and a sparse matrix S using a low-rank matrix factorization algorithm.

[0034] S42. Optimize the low-rank matrix L and the sparse matrix S using the nuclear norm minimization method:

[0035] min L,S ∥L∥ * +λ∥S∥1;

[0036] Among them, ∥L∥ * Let represent the nuclear norm of the low-rank matrix L, ∥S∥1 represent the L1 norm of the sparse matrix S, and λ represent the regularization parameter.

[0037] S43. The iterative solution is performed using the alternating direction multiplier method, with the specific update rules as follows:

[0038]

[0039] Y k+1 =Y k+G(t,s)-L k+1 -S k+1 ;

[0040] Among them, L k+1 Let S represent the low-rank matrix in the (k+1)th iteration. k+1 Y represents the sparse matrix in the (k+1)th iteration. k +1 Let ∥ represent the Lagrange multipliers in the (k+1)th iteration. F S represents the Frobenius norm. k Y represents the sparse matrix in the k-th iteration. k Let represent the Lagrange multipliers in the k-th iteration, ρ represent the augmented Lagrange parameters, and G(t,s) represent the global feature representation with spatiotemporal correlation.

[0041] S44. After completing the iteration, obtain the optimized low-rank matrix L. * and the optimized sparse matrix S * Then, background noise is removed from the global feature representation to obtain the denoised feature representation:

[0042] G′(t,s)=G(t,s)-L * ;

[0043] Where G'(t,s) represents the denoising feature representation;

[0044] S45. Dynamically optimize the denoised feature representation G'(t,s) by adjusting the weights of multimodal features in the shared feature pool using a weighting mechanism based on gradient boosting decision trees:

[0045]

[0046] in, This represents the updated weight of the i-th modal feature in the shared feature pool. Let represent the original weight of the i-th modality feature in the shared feature pool, exp represent the exponential function, and η represent the learning rate. This represents the gradient of the loss function J;

[0047] S46. Update the optimized denoised feature representation G'(t,s) back into the shared feature pool.

[0048] Optionally, S5 specifically includes:

[0049] S51. Construct a lightweight self-attention mechanism module, which includes a multi-head attention unit and a feature reconstruction unit for real-time fire monitoring;

[0050] S52. Input the extracted multimodal feature vectors into the multi-head attention unit in the lightweight self-attention mechanism module to perform the correlation analysis between multimodal features, and dynamically calculate the importance weight of each modality feature through the self-attention mechanism to generate a weighted feature representation.

[0051] S53. The weighted feature representation is reconstructed using a feature reconstruction unit to generate an enhanced modal feature vector. The reconstruction process includes feature fusion and redundancy suppression.

[0052] S54. Employing a context-dependent self-attention compensation strategy, this approach analyzes the correlations between existing modal features, infers and compensates for missing modal feature information, and generates a complete multimodal feature vector.

[0053]

[0054] Among them, F complete This represents the complete multimodal eigenvector, where ω represents the nonlinear transformation function. Indicates the concatenation operation, Q i K represents the query matrix for the i-th modal feature. i V represents the key matrix of the i-th modal feature. i Let θ represent the value matrix of the i-th modality feature, n represent the total number of modalities, M represent the index set of missing modalities, and θ represent the value matrix of the i-th modality feature. j C represents the compensation weight factor for the j-th missing mode. j (X j ) represents the compensation feature generation function for the j-th missing mode.

[0055] Optionally, S6 specifically includes:

[0056] S61. Use the complete multimodal feature vector as input to the sparse feature enhancement algorithm, and initialize the regularization parameter and sparsity threshold.

[0057] S62. The complete multimodal feature vector is sparsified by a sparse feature enhancement algorithm. A sparse coding mechanism is adopted to represent each modal feature as a linear combination of the original features. The optimization objective is to minimize the reconstruction error while maintaining the sparsity of the modal feature representation.

[0058] S63. Solve the sparse coding matrix through iterative optimization and combine it with the feature dictionary to generate an enhanced sparse feature matrix;

[0059] S64. Input the enhanced sparse feature matrix into the multi-task classification network, where task 1 is fire status classification and task 2 is smoke diffusion path prediction. The multi-task classification network extracts common features through a shared layer and processes the feature representations of the two tasks in a task-specific layer, forming a parallel structure of classification and prediction.

[0060] S65. Define the objective function for multi-task classification and perform joint optimization by combining the sparse feature matrix. The objective function is expressed as a weighted sum of the loss functions of the two tasks, where the weight coefficients are dynamically adjusted according to the importance of the tasks.

[0061] S66. Fire status classification results and smoke diffusion path prediction results are generated through a multi-task classification network.

[0062] Optionally, S7 specifically includes:

[0063] S71. Collect environmental data in real time and calculate the deviation matrix E between the current multimodal characteristic data of fire monitoring and the prediction results;

[0064] S72. Based on the deviation matrix E, the feature weights in the shared feature pool are optimized and adjusted through a dynamic feedback mechanism:

[0065]

[0066] in, This represents the updated weight of the i-th modal feature in the shared feature pool. η represents the original weight of the i-th modal feature in the shared feature pool, and η1 represents the learning rate. This represents the gradient of the deviation matrix with respect to the modal feature weights;

[0067] S73. Input the optimized shared feature pool into the embedded spatiotemporal tracking module to update the fire source location;

[0068] S74. Based on the optimized fire source location and combined with environmental characteristics, predict the future location of the fire source and generate trajectory prediction results, including the future fire source location and spread range.

[0069] S75. Evaluate the trajectory prediction results, monitor the error change trend in real time, and judge whether the prediction accuracy has reached the expected level by the change amplitude of the deviation matrix E. If the amplitude of the deviation matrix E exceeds the set threshold, further optimize the spatiotemporal tracking module parameters to improve the positioning accuracy.

[0070] S76. The optimized trajectory prediction results are fed back to the spatiotemporal tracking module of the shared feature pool, thereby improving the response speed and accuracy to changes in the location of the fire source in the next round of monitoring, forming a dynamically optimized closed-loop feedback mechanism.

[0071] The beneficial effects of this invention are:

[0072] First, this invention achieves comprehensive extraction of smoke diffusion patterns, temperature gradients, and dynamic changes in gas composition by fusing visible light, infrared, and photoacoustic spectral data, overcoming the problems of incomplete information and low recognition accuracy in traditional single-modal methods. By constructing a shared feature pool to uniformly store multimodal features and embedding a spatiotemporal tracking module to capture the dynamic changes of multimodal features in time and space, this invention can efficiently model the spatiotemporal correlation characteristics of smoke signals. Compared to existing technologies that only extract static features, this invention provides a solution that dynamically adapts to the complex changes in the early stages of a fire, significantly improving the accuracy of fire source location and the predictive ability of smoke diffusion paths.

[0073] Secondly, this invention employs a low-rank matrix factorization algorithm to generate a dynamic background template, effectively eliminating background noise from the global feature representation and addressing complex background issues such as light variations and wind interference in fire monitoring scenarios. Unlike traditional static background modeling methods, this invention dynamically adjusts the background template based on real-time changes in multimodal data, ensuring the stability and accuracy of feature extraction. Furthermore, the sparse feature enhancement algorithm, tailored to the characteristics of sparse smoke signals in early fires, amplifies key features using a sparse coding mechanism, preventing weak smoke signals from being masked by background interference. Through this feature enhancement process, this invention significantly reduces the false alarm and false alarm rates, providing strong support for the accurate identification of early fires.

[0074] Furthermore, this invention introduces a lightweight self-attention mechanism, effectively solving the dynamic optimization problem of feature weight allocation during multimodal data fusion. By weighting multimodal features through the self-attention mechanism, this invention can adaptively highlight the importance of different modalities in different scenarios, thereby improving the efficiency of feature fusion. Simultaneously, a context compensation strategy is employed to address the problem of missing multimodal data, ensuring the integrity of feature fusion and the robustness of the system. Compared to traditional static fusion methods, this invention significantly improves both fusion efficiency and missing data compensation capabilities, exhibiting stronger adaptability, especially in complex environments and scenarios with incomplete data.

[0075] Finally, this invention constructs a closed-loop optimization system through a dynamic feedback mechanism. This system compares real-time collected environmental data with the system's prediction results, dynamically adjusting the parameters of the shared feature pool and the spatiotemporal tracking module to continuously optimize the accuracy of fire source location. This mechanism can correct prediction deviations in real time, ensuring the system's adaptability and accuracy in dynamic environments. Furthermore, by combining fire state classification and smoke diffusion path prediction tasks, this invention achieves comprehensive monitoring of fire development, providing forward-looking support for fire safety decision-making. The overall system demonstrates significant advantages in accuracy, response speed, and robustness, making it particularly suitable for complex scenarios such as industrial plants and large public buildings. Attached Figure Description

[0076] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0077] Figure 1 This is an overall flowchart of a fire smoke image early recognition method based on multimodal fusion proposed in this invention;

[0078] Figure 2 This is a schematic diagram of the shared feature pool structure of the embedded spatiotemporal tracking module in the fire smoke image early recognition method based on multimodal fusion proposed in this invention. Detailed Implementation

[0079] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0080] refer to Figure 1 and Figure 2 An early identification method for fire smoke images based on multimodal fusion includes the following steps:

[0081] S1. Collect multimodal data using a visible light camera, an infrared camera, and a photoacoustic spectroscopy sensor, and synchronize the multimodal data based on timestamps.

[0082] S2. Construct visible light branch, infrared branch and photoacoustic branch, respectively extract smoke diffusion morphology, temperature gradient and dynamic change features of gas composition from the multimodal data, generate multimodal features and store them in the shared feature pool;

[0083] S3. Embed a spatiotemporal tracking module in the shared feature pool to capture the variation patterns of the multimodal features on the time axis and spatial distribution, and generate a global feature representation with spatiotemporal correlation characteristics;

[0084] S4. Based on the global feature representation, a low-rank matrix factorization algorithm is used to generate a dynamic background template, separate and remove background noise, and dynamically optimize the multimodal features in the shared feature pool.

[0085] S5. By using a lightweight self-attention mechanism, multimodal features are fused in the shared feature pool and missing features are compensated to generate a complete multimodal feature vector.

[0086] S6. Based on the complete multimodal feature vector, the early smoke signal is enhanced by a sparse feature enhancement algorithm, and a multi-task classification network is applied to classify the fire status and predict the smoke diffusion path.

[0087] S7. Based on the deviation of real-time environmental data, update the shared feature pool and spatiotemporal tracking module through a dynamic feedback mechanism to optimize the accuracy of fire source positioning.

[0088] In this embodiment, S2 specifically includes:

[0089] S21. Acquire visible light image data using a visible light camera and construct a visible light branch. The visible light image data is used to obtain the diffusion morphology characteristics of the smoke, including the edge features, diffusion speed, and diffusion direction of the smoke.

[0090] S22. Infrared image data is collected by an infrared camera to construct an infrared branch. The infrared image data is used to obtain the temperature gradient characteristics of the smoke area, including the temperature change of the smoke diffusion area and the intensity of the hot spot area.

[0091] S23. Collect gas composition data through a photoacoustic spectral sensor and construct a photoacoustic branch. The gas composition data is used to obtain the dynamic change characteristics of the gas in the smoke.

[0092] S24. Standardize the features extracted from each branch to generate multimodal features and store them in the shared feature pool.

[0093] In this embodiment, S3 specifically includes:

[0094] S31. Construct a spatiotemporal tracking module, which is embedded in a shared feature pool. The spatiotemporal tracking module includes a temporal convolutional network and a spatiotemporal attention mechanism. The temporal convolutional network is used to capture the change patterns of multimodal features in the time dimension, and the spatiotemporal attention mechanism is used to capture the distribution characteristics of multimodal features in the spatial dimension.

[0095] S32. Apply a temporal convolutional network to perform temporal modeling on the multimodal features, generating temporal feature representations at each time step:

[0096] T(t)=σ(W t *X(t)+bt );

[0097] Where T(t) represents the temporal feature representation at time step t, σ represents the activation function, and W t The weights represent the temporal convolution kernel weights, * represents the convolution operation, X(t) represents the multimodal feature input at time step t, and b represents the multimodal feature input at time step t. t Represents the bias vector;

[0098] S33. Apply a spatiotemporal attention mechanism to weight the multimodal features in the spatial dimension to generate spatial feature representations for each spatial location:

[0099]

[0100] Where A(s) represents the spatial feature representation of spatial location s, Q(s) represents the query matrix of spatial location s, K(s) represents the key matrix of spatial location s, V(s) represents the value matrix of spatial location s, Softmax represents the normalization operation, and d k Indicates the dimension of the key vector;

[0101] S34. Fuse the temporal feature representation T(t) and the spatial feature representation A(s) to generate a global feature representation with spatiotemporal correlation:

[0102]

[0103] Where G(t,s) represents a global feature representation with spatiotemporal correlation, and α, β, and γ represent weight parameters. This represents the tensor product operation, used to capture the interaction between temporal and spatial features.

[0104] In this embodiment, S4 specifically includes:

[0105] S41. Based on the global feature representation, the global feature representation is decomposed into a low-rank matrix L and a sparse matrix S using a low-rank matrix factorization algorithm.

[0106] S42. Optimize the low-rank matrix L and the sparse matrix S using the nuclear norm minimization method:

[0107] min L,S ∥L∥ * +λ∥S∥1;

[0108] Among them, ∥L∥ * Let represent the nuclear norm of the low-rank matrix L, ∥S∥1 represent the L1 norm of the sparse matrix S, and λ represent the regularization parameter.

[0109] S43. The iterative solution is performed using the alternating direction multiplier method, with the specific update rules as follows:

[0110]

[0111] Y k+1 =Y k +G(t,s)-L k+1 -S k+1 ;

[0112] Among them, L k+1 Let S represent the low-rank matrix in the (k+1)th iteration. k+1 Y represents the sparse matrix in the (k+1)th iteration. k +1 Let ∥ represent the Lagrange multipliers in the (k+1)th iteration. F S represents the Frobenius norm. k Y represents the sparse matrix in the k-th iteration. k Let represent the Lagrange multipliers in the k-th iteration, ρ represent the augmented Lagrange parameters, and G(t,s) represent the global feature representation with spatiotemporal correlation.

[0113] S44. After completing the iteration, obtain the optimized low-rank matrix L. * and the optimized sparse matrix S * Then, background noise is removed from the global feature representation to obtain the denoised feature representation:

[0114] G′(t,s)=G(t,s)-L * ;

[0115] Where G'(t,s) represents the denoising feature representation;

[0116] S45. Dynamically optimize the denoised feature representation G'(t,s) by adjusting the weights of multimodal features in the shared feature pool using a weighting mechanism based on gradient boosting decision trees:

[0117]

[0118] in, This represents the updated weight of the i-th modal feature in the shared feature pool. Let represent the original weight of the i-th modality feature in the shared feature pool, exp represent the exponential function, and η represent the learning rate. This represents the gradient of the loss function J;

[0119] S46. Update the optimized denoised feature representation G'(t,s) back into the shared feature pool.

[0120] In this embodiment, S5 specifically includes:

[0121] S51. Construct a lightweight self-attention mechanism module, which includes a multi-head attention unit and a feature reconstruction unit for real-time fire monitoring;

[0122] S52. Input the extracted multimodal feature vectors into the multi-head attention unit in the lightweight self-attention mechanism module to perform the correlation analysis between multimodal features, and dynamically calculate the importance weight of each modality feature through the self-attention mechanism to generate a weighted feature representation.

[0123] S53. The weighted feature representation is reconstructed using a feature reconstruction unit to generate an enhanced modal feature vector. The reconstruction process includes feature fusion and redundancy suppression.

[0124] S54. Employing a context-dependent self-attention compensation strategy, this approach analyzes the correlations between existing modal features, infers and compensates for missing modal feature information, and generates a complete multimodal feature vector.

[0125]

[0126] Among them, F complete This represents the complete multimodal eigenvector, where ω represents the nonlinear transformation function. Indicates the concatenation operation, Q i K represents the query matrix for the i-th modal feature. i V represents the key matrix of the i-th modal feature. i Let θ represent the value matrix of the i-th modality feature, n represent the total number of modalities, M represent the index set of missing modalities, and θ represent the value matrix of the i-th modality feature. j C represents the compensation weight factor for the j-th missing mode. j (X j ) represents the compensation feature generation function for the j-th missing mode.

[0127] In this embodiment, S6 specifically includes:

[0128] S61. Use the complete multimodal feature vector as input to the sparse feature enhancement algorithm, and initialize the regularization parameter and sparsity threshold.

[0129] S62. The complete multimodal feature vector is sparsified by a sparse feature enhancement algorithm. A sparse coding mechanism is adopted to represent each modal feature as a linear combination of the original features. The optimization objective is to minimize the reconstruction error while maintaining the sparsity of the modal feature representation.

[0130] S63. Solve the sparse coding matrix through iterative optimization and combine it with the feature dictionary to generate an enhanced sparse feature matrix;

[0131] S64. Input the enhanced sparse feature matrix into the multi-task classification network, where task 1 is fire status classification and task 2 is smoke diffusion path prediction. The multi-task classification network extracts common features through a shared layer and processes the feature representations of the two tasks in a task-specific layer, forming a parallel structure of classification and prediction.

[0132] S65. Define the objective function for multi-task classification and perform joint optimization by combining the sparse feature matrix. The objective function is expressed as a weighted sum of the loss functions of the two tasks, where the weight coefficients are dynamically adjusted according to the importance of the tasks.

[0133] S66. Fire status classification results and smoke diffusion path prediction results are generated through a multi-task classification network.

[0134] In this embodiment, S7 specifically includes:

[0135] S71. Collect environmental data in real time and calculate the deviation matrix E between the current multimodal characteristic data of fire monitoring and the prediction results;

[0136] S72. Based on the deviation matrix E, the feature weights in the shared feature pool are optimized and adjusted through a dynamic feedback mechanism:

[0137]

[0138] in, This represents the updated weight of the i-th modal feature in the shared feature pool. η represents the original weight of the i-th modal feature in the shared feature pool, and η1 represents the learning rate. This represents the gradient of the deviation matrix with respect to the modal feature weights;

[0139] S73. Input the optimized shared feature pool into the embedded spatiotemporal tracking module to update the fire source location;

[0140] S74. Based on the optimized fire source location and combined with environmental characteristics, predict the future location of the fire source and generate trajectory prediction results, including the future fire source location and spread range.

[0141] S75. Evaluate the trajectory prediction results, monitor the error change trend in real time, and judge whether the prediction accuracy has reached the expected level by the change amplitude of the deviation matrix E. If the amplitude of the deviation matrix E exceeds the set threshold, further optimize the spatiotemporal tracking module parameters to improve the positioning accuracy.

[0142] S76. The optimized trajectory prediction results are fed back to the spatiotemporal tracking module of the shared feature pool, thereby improving the response speed and accuracy to changes in the location of the fire source in the next round of monitoring, forming a dynamically optimized closed-loop feedback mechanism.

[0143] Example 1:

[0144] To verify the feasibility of this invention in practice, it was applied to an early fire warning system in a large industrial park. Located in the suburbs of a city, the industrial park covers approximately 50,000 square meters and contains multiple high-temperature production workshops and flammable material storage warehouses, posing a high fire risk and presenting a complex environment with factors such as high temperatures, smoke interference, and dynamic light changes. This test aims to evaluate the practical performance of the proposed multimodal fusion-based early fire smoke image recognition method, particularly its performance under conditions of sparse smoke signals, complex background noise, and missing data.

[0145] In this embodiment, the system deploys a visible light camera, an infrared camera, and a photoacoustic spectroscopy sensor to collect the diffusion pattern, temperature gradient, and dynamic changes in gas composition of the smoke, respectively. All devices synchronize and align data via timestamps, and the multimodal features are stored in a shared feature pool, where an embedded spatiotemporal tracking module performs dynamic modeling and feature optimization. This method combines a low-rank matrix factorization algorithm to generate a dynamic background template, removing background noise from the global feature representation and optimizing the multimodal features. Furthermore, a lightweight self-attention mechanism is used for multimodal feature fusion, and a sparse feature enhancement algorithm amplifies the early sparse smoke signal, thereby completing fire state classification and smoke diffusion path prediction.

[0146] During the experiment, various fire scenarios were simulated, including three stages: normal state, early smoke diffusion, and gradual fire development. Under normal state, the system's anti-interference capability and false alarm rate were tested. During the early smoke diffusion stage, the ability of the sparse feature enhancement algorithm to identify weak smoke signals was verified. During the fire development stage, the accuracy of fire state classification and the reliability of diffusion path prediction were tested. Specific experimental locations covered production workshops, warehouses, and control rooms in an industrial park, with 10 monitoring points selected. Data was collected at a frequency of 10 frames per second at each monitoring point, ultimately generating approximately 5.2TB of experimental data.

[0147] Table 1. Performance Comparison of the Invention and Existing Technologies in Early Fire Detection

[0148] Test metrics This invention Existing technology Increase Recognition accuracy (%) 98.5 91.2 +7.3 False alarm rate (%) 4.7 12.3 -7.6 Classification accuracy (%) 97.8 89.5 +8.3 Diffusion path prediction error (meters) 1.2 3.6 -2.4 Data compensation accuracy (%) 95.4 78.6 +16.8 Average delay time (seconds) 0.8 1.7 -0.9

[0149] The experimental results in the table above show that the recognition accuracy of this invention reaches 98.5%, an improvement of 7.3 percentage points compared to the 91.2% of the prior art. This result indicates that this invention can more accurately capture the smoke characteristics in the early stages of a fire, significantly reducing the risk of missed detections. Furthermore, regarding the false alarm rate, this invention, through dynamic background templates and feature optimization algorithms, reduces the false alarm rate from 12.3% of the prior art to 4.7%, reducing the impact of environmental background interference on fire monitoring, and performing particularly well in complex dynamic scenarios.

[0150] Regarding classification accuracy, the present invention achieves 97.8%, an improvement of 8.3 percentage points compared to the prior art's 89.5%. This indicates that the present invention has higher reliability in distinguishing fire states (such as normal, early smoke stage, and fire source development), providing a more reliable basis for fire emergency decision-making. Simultaneously, in terms of diffusion path prediction performance, the average error distance of the present invention is only 1.2 meters, while the prior art is 3.6 meters, significantly reducing prediction error. This improvement effectively enhances the accuracy of fire source diffusion path prediction, enabling fire monitoring systems to provide more precise references for the deployment of fire-fighting measures.

[0151] To address the issue of missing data, this invention employs a context compensation strategy, improving the accuracy of data compensation to 95.4%, a 16.8 percentage point improvement compared to the 78.6% of traditional interpolation algorithms. This performance demonstrates that the invention can effectively compensate for feature information when dealing with incomplete modal data, thereby ensuring the integrity and consistency of multimodal data fusion. Furthermore, regarding system response time, the average latency from data acquisition to outputting the recognition result is 0.8 seconds, a reduction of 0.9 seconds compared to the 1.7 seconds of existing technologies. This significant performance improvement elevates the real-time capabilities of this invention in early fire detection to a new level.

[0152] The data analysis in this embodiment fully demonstrates that the present invention, through multimodal fusion, dynamic background processing, sparse feature enhancement, and dynamic feedback optimization, effectively addresses the shortcomings of traditional fire monitoring technologies in areas such as sparse smoke recognition, dynamic scene adaptation, data loss compensation, and real-time response. The present invention exhibits significant advantages in recognition accuracy, real-time performance, and adaptability, providing reliable technical support for early fire warning and possessing broad application prospects.

[0153] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multi-modal fusion-based early fire smoke image recognition method, characterized in that, The method comprises the following steps: S1, acquiring multi-modal data through a visible light camera, an infrared camera and a photoacoustic spectrum sensor, and performing time alignment on the multi-modal data based on a timestamp; S2, constructing a visible light branch, an infrared branch and a photoacoustic branch, respectively extracting smoke diffusion morphology, temperature gradient and gas component dynamic change characteristics from the multi-modal data, generating multi-modal features and storing them in a shared feature pool; S3, embedding a space-time tracking module in the shared feature pool, capturing the change law of the multi-modal features on the time axis and the spatial distribution, and generating a global feature representation with space-time correlation characteristics; S4, based on the global feature representation, generating a dynamic background template using a low-rank matrix decomposition algorithm, separating and removing background noise, and dynamically optimizing the multi-modal features in the shared feature pool; S5, fusing multi-modal features and compensating missing features in the shared feature pool through a lightweight self-attention mechanism, and generating a complete multi-modal feature vector; S6, based on the complete multi-modal feature vector, enhancing early smoke signals through a sparse feature enhancement algorithm, and applying a multi-task classification network for fire state classification and smoke diffusion path prediction; S7, according to the deviation of real-time environmental data, updating the shared feature pool and the space-time tracking module through a dynamic feedback mechanism, and optimizing the fire source positioning accuracy; The S4 specifically comprises: S41, based on the global feature representation, applying a low-rank matrix decomposition algorithm to decompose the global feature representation into a low-rank matrix L and a sparse matrix S; S42, optimizing the low-rank matrix L and the sparse matrix S through a kernel norm minimization method: min L,S ||L|| * +λ||S||1; Among them, ||L|| * Let ||S||1 denote the kernel norm of the low-rank matrix L, ||S||1 denote the L1 norm of the sparse matrix S, and λ denote the regularization parameter. S43, iteratively solving by using an alternating direction multiplier method, and the specific update rule is: Y k+1 = Y k + G(t,s) - L k+1 - S k+1 ; where L k+1 denotes the low-rank matrix in the k+1th iteration, S k+1 denotes the sparse matrix in the k+1th iteration, Y k+1 denotes the Lagrange multiplier in the k+1th iteration, ||·||F F denotes the Frobenius norm, S k denotes the sparse matrix in the kth iteration, Y k denotes the Lagrange multiplier in the kth iteration, p denotes the augmented Lagrange parameter, and G(t,s) denotes the global feature representation with spatiotemporal correlation. S44、After completing the iteration, obtain the optimized low-rank matrix L * and the optimized sparse matrix S * and remove the background noise from the global feature representation to obtain a denoised feature representation: G'(t,s) = G(t,s) - L * ; Wherein, G'(t,s) represents the denoising feature representation; S45, dynamically optimizing the denoising feature representation G'(t,s), and adjusting the weight of the multi-modal features in the shared feature pool by using a weighted mechanism based on a gradient boosting decision tree: wherein, denotes the updated weight of the i-th modality feature in the shared feature pool, denotes the original weight of the i-th modality feature in the shared feature pool, exp denotes the exponential function, and η denotes the learning rate, denotes the gradient of the loss function J; S46, updating the optimized denoising feature representation G'(t,s) back to the shared feature pool.

2. The method according to claim 1, wherein, The S2 specifically comprises: S21, acquiring visible light image data through a visible light camera, and constructing a visible light branch, the visible light image data is used to obtain the diffusion morphology characteristics of smoke, including the edge features, diffusion speed and diffusion direction of smoke; S22, acquiring infrared image data through an infrared camera, and constructing an infrared branch, the infrared image data is used to obtain the temperature gradient characteristics of the smoke area, including the temperature change of the smoke diffusion area and the intensity of the hot spot area; S23, acquiring gas component data through a photoacoustic spectrum sensor, and constructing a photoacoustic branch, the gas component data is used to obtain the dynamic change characteristics of the gas in the smoke; S24, standardizing the features extracted by each branch, generating multi-modal features and storing them in a shared feature pool.

3. The method according to claim 1, characterized in that, The S3 specifically comprises: S31, a space-time tracking module is constructed, the space-time tracking module is embedded in a shared feature pool, the space-time tracking module includes a time convolution network and a space-time attention mechanism, the time convolution network is used for capturing the change rule of the multi-modal feature in the time dimension, and the space-time attention mechanism is used for capturing the distribution characteristics of the multi-modal feature in the space dimension; S32, the time convolution network is applied to time modeling of the multi-modal feature, and a time feature representation at each time step is generated: T(t) = σ(W t X(t) + b t ); where T(t) denotes the time feature representation at time step t, σ denotes an activation function, W t denotes a time convolution kernel weight, * denotes a convolution operation, X(t) denotes a multi-modal feature input at time step t, b t denotes a bias vector; S33, the space-time attention mechanism is applied to weighted processing of the multi-modal feature in the space dimension, and a space feature representation at each space position is generated: wherein A(s) represents a spatial feature representation of the spatial position s, Q(s) represents a query matrix of the spatial position s, K(s) represents a key matrix of the spatial position s, V(s) represents a value matrix of the spatial position s, Softmax represents a normalization operation, d k denotes the dimension of the key vector; S34, the time feature representation T(t) and the space feature representation A(s) are fused to generate a global feature representation with space-time correlation: Wherein, G(t,s) represents a global feature representation with spatio-temporal correlation, and α, β and γ represent weight parameters, represents a tensor product operation for capturing the interaction relationship between time and space features.

4. The method according to claim 1, characterized in that, The S5 specifically includes: S51, a lightweight self-attention mechanism module is constructed, the lightweight self-attention mechanism module includes a multi-head attention unit and a feature reconstruction unit, and is used for real-time monitoring of a fire; S52, the extracted multi-modal feature vector is input into the multi-head attention unit in the lightweight self-attention mechanism module, mutual correlation analysis between multi-modal features is performed, the importance weight of each modal feature is dynamically calculated through the self-attention mechanism, and a weighted feature representation is generated; S53, the weighted feature representation is reconstructed by using the feature reconstruction unit, and an enhanced modal feature vector is generated, and the reconstruction process includes feature fusion and redundancy information suppression; S54, a self-attention compensation strategy based on context dependence is adopted, feature information of a missing modal is inferred and compensated by analyzing the correlation between existing modal features, and a complete multi-modal feature vector is generated: where F complete denotes the complete multimodal feature vector, ω denotes a nonlinear transformation function, denotes the concatenation operation, Q i denotes the query matrix of the i-th modality feature, K i denotes the key matrix of the i-th modality feature, V i denotes the value matrix of the i-th modality feature, n denotes the total number of modalities, M denotes the index set of missing modalities, θ j denotes the compensation weight factor of the j-th missing modality, C j (X j ) denotes the compensation feature generation function of the j-th missing modality.

5. The method according to claim 1, wherein The S6 specifically includes: S61, the complete multi-modal feature vector is taken as an input of a sparse feature enhancement algorithm, and a regularization parameter and a sparsity threshold value are initialized; S62, the complete multi-modal feature vector is processed by the sparse feature enhancement algorithm, each modal feature is represented as a linear combination of original features by using a sparse coding mechanism, an optimization target is to minimize reconstruction error, and sparsity of the modal feature representation is maintained at the same time; S63, a sparse coding matrix is solved by an iterative optimization method, and is combined with a feature dictionary to generate an enhanced sparse feature matrix; S64, the enhanced sparse feature matrix is input into a multi-task classification network, task 1 is fire state classification, task 2 is smoke diffusion path prediction, the multi-task classification network extracts common features through a shared layer, and meanwhile, feature representations of the two tasks are processed in task-specific layers, forming a parallel structure of classification and prediction; S65, a target function of multi-task classification is defined, and joint optimization is performed in combination with the sparse feature matrix, and the target function is represented as a weighted sum of two task loss functions, wherein a weight coefficient is dynamically adjusted according to task importance; S66, a fire state classification result and a smoke diffusion path prediction result are generated by the multi-task classification network.

6. The method according to claim 1, wherein The S7 specifically includes: S71, environmental data is collected in real time, and a deviation matrix E between multi-modal feature data of current fire monitoring and a prediction result is calculated; S72, based on the deviation matrix E, the feature weights in the shared feature pool are optimized and adjusted through a dynamic feedback mechanism; wherein, denotes the updated weight of the i-th modality feature in the shared feature pool, denotes the original weight of the i-th modality feature in the shared feature pool, and η1 denotes a learning rate, denotes the gradient of the bias matrix on the modality feature weight; S73, input the optimized shared feature pool into the embedded space-time tracking module to update and calculate the fire source position; S74, according to the optimized fire source position, the future position of the fire source is predicted in combination with the environmental features, and a trajectory prediction result is generated, including the future fire source position and the diffusion range; S75, the trajectory prediction result is evaluated, the error change trend is monitored in real time, whether the prediction accuracy reaches the expectation is judged through the change amplitude of the deviation matrix E, if the amplitude of the deviation matrix E exceeds the set threshold, the space-time tracking module parameters are further optimized to improve the positioning accuracy; S76, the optimized trajectory prediction result is fed back to the space-time tracking module of the shared feature pool, so as to improve the response speed and accuracy of the fire source position change in the next round of monitoring, and a dynamic optimization closed loop feedback mechanism is formed.

Citation Information

Patent Citations

  • Fire early warning method and system based on smoke image joint feature analysis

    CN119181197A

  • Fire recognition method and apparatus, and computer device and storage medium

    WO2022121129A1