Warehouse fire intelligent detection and fire extinguishing method based on multi-modal information fusion
Through intelligent detection and fire extinguishing methods of multimodal information fusion, traditional fire protection systems have solved the problems of lagging response, high false alarm rate and low fire extinguishing efficiency in fire detection and fire extinguishing, and achieved accurate detection and dynamic fire extinguishing of warehouse fires, improving fire extinguishing efficiency and safety.
Patent Information
- Application Number
- CN202510495269.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional fire protection systems have problems such as lagging response, high false alarm rate and low fire extinguishing efficiency in fire detection and fire extinguishing, especially in warehouse environments, which are difficult to achieve accurate and rapid fire extinguishing.
Intelligent fire detection and fire extinguishing method based on multimodal information fusion is adopted. By collecting visual images, thermal imaging images and smoke concentration time series, combining multimodal fire prediction model and flame target detection model, fire level prediction and fire source positioning are achieved, and water guns are dynamically dispatched to extinguish fire.
Accurate detection and positioning of fires is achieved, and water guns are dynamically dispatched to improve fire extinguishing efficiency, avoiding inaccurate fire judgments caused by the fusion of single modal information, reducing false alarm rates and waste of fire extinguishing agents, and effectively suppressing the spread of fire.
Smart Images

Figure CN120204670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine vision and fire protection, and specifically to an intelligent detection and extinguishing method for warehouse fires based on multi-modal information fusion. Background Art
[0002] Traditional fire protection systems mainly rely on single environmental perception devices such as smoke detectors and temperature sensors, combined with manual or semi-automatic fire extinguishing devices to achieve fire prevention and control. There are generally problems such as lagging response, high false alarm rate, and low fire extinguishing efficiency. The fire protection system based on smoke detectors issues an alarm only when smoke is detected, and it cannot accurately locate the specific location of the fire, making it difficult to take effective fire extinguishing measures. At the same time, it is easily interfered by non-fire factors such as dust and steam, resulting in possible false alarms due to smoke generated in non-fire situations. Although the fire protection system based on thermal imaging can identify the location of the fire according to temperature information, the fire spreads over time, and the temperature information cannot reflect the trend of fire spread, so it cannot suppress the spread of the fire. The fire protection system based on visual images identifies the flame through target detection and locates the fire location, but it can only be accurately identified when there is an open fire, and it cannot be effectively identified in the initial stage of the fire, showing lag.
[0003] Traditional fire protection systems usually deploy water guns in corridors or around warehouses, with a wide coverage area. After a fire occurs, all water guns participate in the fire extinguishing task, which is likely to cause waste of fire extinguishing agents and secondary water damage. For the fire protection task of warehouses, it is necessary to extinguish the fire accurately and quickly to prevent the fire from spreading and causing greater losses, and at the same time, it is necessary to avoid large-area spraying to prevent excessive damage to items. Therefore, there is an urgent need for a fire extinguishing method that integrates real-time monitoring, accurate positioning, and intelligent decision-making to improve the fire extinguishing efficiency and reduce disaster losses. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide an intelligent detection and extinguishing method for warehouse fires based on multi-modal information fusion.
[0005] The technical solution adopted by the present invention to solve the above technical problem is:
[0006] An intelligent detection and extinguishing method for warehouse fires based on multi-modal information fusion, characterized by including the following steps:
[0007] The first step: Collect visual images, thermal imaging images, and smoke concentration time series;
[0008] The second step: Input the smoke concentration time series, visual images, and thermal imaging images into a multi-modal fire prediction model to predict the fire level;
[0009] Extract the smoke concentration feature using the smoke concentration time series, extract the temperature feature from the thermal imaging image, and extract the flame feature from the visual image;
[0010] The three features of smoke concentration, temperature, and flame are enhanced respectively using a fusion method driven by the correlation relationship to obtain the enhanced smoke concentration feature, enhanced temperature feature, and enhanced flame feature; the enhanced temperature feature and flame feature are concatenated to obtain the shallow fusion feature; the enhanced smoke concentration feature and the shallow fusion feature are respectively passed through an MLP and then concatenated to obtain the middle fusion feature; at the same time, the enhanced smoke concentration feature, temperature feature, and flame feature are concatenated to obtain the enhanced fusion feature;
[0011] Adopt an attention gating mechanism to assign weights to each fusion feature; perform weighted fusion on each fusion feature according to the following formula to obtain the deep fusion feature;
[0012]
[0013] In the formula, C final represents the deep fusion feature, G early , G middle , G AF represent the weights of the shallow fusion feature C early , the middle fusion feature C middle and the enhanced fusion feature C AF respectively;
[0014] The deep fusion feature passes through the Softmax function to obtain the prediction score, and then the fire level is obtained;
[0015] Step 3: Input the visual image into the flame target detection model to predict the fire source and the fire range; according to the predicted flame bounding box, use the optical flow method to predict the flame spread direction;
[0016] Step 4: Dispatch the water guns according to the fire level, fire source, and flame spread direction;
[0017] Divide the internal space of the warehouse into four regions in a cross form, and a rotatable water gun is installed at the center position of each region; calculate the distance from each water gun to the fire source, and the water gun with the closest distance undertakes the main fire extinguishing task, and select a water gun in the flame spread direction to undertake the auxiliary fire extinguishing task; the fire level is divided into three levels: first, second, and third from high to low. For each level increase, an additional water gun with a relatively close distance undertakes the main fire extinguishing task; when the fire level is the first level, all four water guns are put into fire extinguishing.
[0018] Furthermore, the flame target detection model is obtained by replacing the C2f module of the YOLOv10 backbone network with the C2f_RepViTSimAM module one by one; the C2f_RepViTSimAM module is obtained by replacing the Bottleneck structure of the C2f module with the RepViTSimAM module; in the RepViTSimAM module, the input feature is split into three sub-features along the channel dimension. After the second sub-feature passes through a 3×3 depthwise separable convolution and the third sub-feature passes through a 1×1 depthwise separable convolution, they are concatenated with the first sub-feature in the channel dimension. The concatenated feature passes through the SimAM attention mechanism to obtain an attention feature; after the attention feature passes through a feed-forward neural network, the output feature of the RepViTSimAM module is obtained.
[0019] Furthermore, the smoke concentration time series is input into a long short-term memory network, and the hidden state of the last time step is used as the smoke concentration feature; the thermal imaging image is input into a ResNet50 network, and the output feature of the second stage of the ResNet50 network is used as the temperature feature; the visual image is input into the flame target detection model, and the output feature of the backbone network of the flame target detection model is used as the flame feature.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] 1. The present invention obtains fire information through visual images, thermal imaging images, and smoke sensing, fuses three-modal information of vision, temperature, and smoke, predicts the fire level, and avoids problems such as inaccurate fire determination that may be caused by single-modal or dual-modal information fusion. Precise fire source and fire range positioning through target detection provide a basis for precise fire extinguishing. At the same time, the flame spread direction can be determined. According to the fire level, fire source, and flame spread direction, that is, according to the specific situation of the fire, the water guns for performing the fire extinguishing task are determined. The water gun closest to the fire source is mainly used for fire extinguishing, and the water guns in the flame spread direction assist in fire extinguishing. This strategy effectively suppresses the spread of the fire while extinguishing the fire, realizing precise and effective automatic fire extinguishing, and solving the problem that existing fire extinguishing methods cannot dynamically dispatch water guns according to the fire situation.
[0022] The water guns are arranged at the central positions of each area of the warehouse, rather than around the warehouse. On the basis of ensuring that the entire internal space of the warehouse is covered by the spraying range of the water guns, the space covered by a single water gun is reduced, achieving precise aiming, solving the problem of unnecessary water loss that may be caused by the extensive spraying area of traditional fire extinguishing systems, especially for the surrounding goods that are not directly affected by the flame, and the existing fire extinguishing dead corners, ensuring that any fire source can be effectively handled.
[0023] 2. The multi-modal fire prediction model first extracts the smoke concentration, temperature, and flame characteristics, then processes these three characteristics to extract shallow fusion features, middle fusion features, and enhanced fusion features. Finally, these three fusion features are weighted and fused to obtain deep fusion features, and the deep fusion features are used for fire level prediction, achieving the full fusion of three-modal information and improving the prediction accuracy.
[0024] Use the C2f_RepViTSimAM module to replace the C2f module of the YOLOv10 backbone network one by one to obtain a flame target detection model, which can accurately locate the fire location. Use the RepViTSimAM module to replace the Bottleneck structure of the C2f module to obtain the C2f_RepViTSimAM module, which reduces the computational complexity and improves the target detection speed. The C2f_RepViTSimAM module introduces the SimAM attention mechanism in the RepViTSimAM module, enhancing the feature extraction ability. At the same time, no additional parameters are introduced, hardly increasing the computational burden while improving the model performance, which is very suitable for resource-constrained environments. Compared with other complex attention mechanisms, the SimAM attention mechanism does not require complex hyperparameter tuning and can be more easily integrated into the existing network architecture without spending a lot of time on parameter tuning. Brief Description of the Drawings
[0025] Figure 1 is the overall flowchart of the present invention;
[0026] Figure 2 is the prediction flowchart of the fire level of the present invention;
[0027] Figure 3 is the structural diagram of the flame target detection model of the present invention;
[0028] Figure 4 is the structural diagram of the C2f_RepViTSimAM module of the present invention. Detailed Embodiment
[0029] Specific embodiments are given below in conjunction with the accompanying drawings. The specific embodiments are only used to introduce the technical solutions of the present invention in detail and do not limit the protection scope of this application.
[0030] The present invention provides a warehouse fire intelligent detection and extinguishing method based on multi-modal information fusion, including the following steps:
[0031] The first step: Collect visual images, thermal imaging images, and smoke concentration time series;
[0032] Use a visual camera and an infrared thermal imaging camera to respectively collect the video streams inside the warehouse in real time, perform frame extraction on the video streams, for example, extract one frame every 5s, to obtain visual images and thermal imaging images; use a smoke sensor to collect the smoke concentration to obtain a smoke concentration time series; for example, detect the smoke concentration 4 times per second, and form a smoke concentration time series with a length of 20 every 5s.
[0033] Step 2: Input the smoke concentration time series, visual images, and thermal imaging images into a multi-modal fire prediction model to predict the fire level;
[0034] First, the smoke concentration time series undergoes feature extraction through a long short-term memory network, and the hidden state of the last time step is used as the smoke concentration feature F smoke ; the thermal imaging image is input into the ResNet50 network, and the output feature of the second stage is used as the temperature feature F temp ; the visual image is input into a flame target detection model, and the output feature of the backbone network is used as the flame feature F fire ;
[0035] Next, the three features of smoke concentration, temperature, and flame are enhanced respectively using an association relationship-driven fusion method (AF), and the features of the three modalities are unified into a semantically consistent association relationship space to obtain the enhanced smoke concentration feature C smoke , the enhanced temperature feature C temp and the enhanced flame feature C fire ;
[0036] Taking the smoke concentration feature as an example, perform a power mapping on the smoke concentration feature according to Equation (1) to expand the dimension of the smoke concentration feature to L dimensions, which is used to enhance the non-linear representation ability of the smoke concentration feature and improve the high-order relationship of the feature, so as to achieve the enhancement of the smoke concentration feature and obtain the smoke concentration features of each dimension;
[0037] [F smoke 1 ,...,F smoke i ,...,F smoke j ,...,F smoke L = PowerMapping(F smoke ,L) (1)
[0038] In the formula, F smoke i , F smoke j respectively represent the smoke concentration features of the i-th and j-th dimensions, and PowerMapping(·) represents the power mapping operation;
[0039] Calculate the correlation matrix between smoke concentration features in different dimensions according to Equation (2):
[0040]
[0041] where R smoke (i,j) represents the correlation matrix between the smoke concentration features F smoke i and F smoke j in the i-th and j-th dimensions, cov represents the covariance, respectively represent the standard deviations of the smoke concentration features F smoke i and F smoke j ;
[0042] Approximate and fuse the smoke concentration features and the correlation matrix in each dimension through the Taylor series:
[0043]
[0044] where C smoke represents the enhanced smoke concentration feature;
[0045] Similarly, obtain the enhanced temperature feature and flame feature.
[0046] Then, perform hierarchical fusion on the enhanced smoke concentration feature, temperature feature, and flame feature; since the temperature feature and the flame feature are highly similar, the enhanced temperature feature C temp and the flame feature C fire are directly concatenated to obtain the shallow fusion feature C early ; the enhanced smoke concentration feature C smoke and the shallow fusion feature C early are respectively processed by a multi-layer perceptron (MLP) for cross-modal feature processing, and the output features of the two MLPs are concatenated to obtain the middle fusion feature C middle ; according to Equation (4), the enhanced smoke concentration feature, temperature feature, and flame feature are concatenated to obtain the enhanced fusion feature C AF ;
[0047] C AF = Concat(C smoke , C temp , C fire ) (4)
[0048] Use the attention gating mechanism to dynamically allocate weights for different fusion features through learnable parameters, and the calculation formula is:
[0049] Gearly = σ(W early ·C early + b early )(5)
[0050] Wherein, G early represents the weight of the shallow fusion feature, W early and b final are learnable weights and biases, and σ represents the Sigmoid activation function;
[0051] Perform weighted fusion on each fusion feature according to Equation (6) to obtain the deep fusion feature C final ;
[0052]
[0053] Finally, pass the deep fusion feature C final through the Softmax function to obtain the prediction score, and then obtain the fire level;
[0054] Score = Softmax(W final ·C final + b final )(7)
[0055] Wherein, Score represents the prediction score, Softmax represents the Softmax function, and W final , b final represent learnable weights and biases.
[0056] Step 3: Input the visual image into the flame target detection model to obtain the predicted flame bounding box, and then the fire source and the ignition range can be obtained; according to the predicted flame bounding box, combine the optical flow method to calculate the motion vector of adjacent frames and predict the flame spread direction;
[0057] Improve the YOLOv10 network, and use the C2f_RepViTSimAM module to replace the C2f module of the YOLOv10 backbone network one by one to obtain the flame target detection model, see Figure 3 ; As Figure 4As shown in the figure, the Bottleneck structure of the C2f module is replaced by the RepViTSimAM module to obtain the C2f_RepViTSimAM module; the SimAM attention mechanism is introduced into the RepViTSimAM module of the C2f_RepViTSimAM module to enhance the feature extraction ability; in the C2f_RepViTSimAM module, the input visual image undergoes preliminary feature extraction through the first convolutional layer, and the extracted features are divided into two parts along the channel dimension through the Split operation. One part is processed through three RepViTSimAM modules, and the output features of the three RepViTSimAM modules are concatenated with the other part in the channel dimension. The concatenated features undergo feature extraction through the second convolutional layer to obtain the output features of the C2f_RepViTSimAM module.
[0058] In the RepViTSimAM module, the input features are divided into three sub-features along the channel dimension. The first sub-feature is directly passed to the concatenation operation as part of the direct connection path to retain the original information; the second sub-feature realizes in-depth mining of spatial information through a 3×3 depthwise separable convolution; the third sub-feature passes through a 1×1 depthwise separable convolution to adjust the number of channels and reduce the computational complexity; the features of the second sub-feature and the third sub-feature after the depthwise separable convolution are concatenated with the first sub-feature in the channel dimension to obtain the concatenated features; the concatenated features are weighted through the SimAM attention mechanism to strengthen important features while suppressing irrelevant information to obtain the attention features; the attention features undergo dimensional transformation through the feed-forward neural network FFN to obtain the output features of the RepViTSimAM module.
[0059] Step 4: Determine the fire extinguishing plan according to the fire level, the fire source, and the flame spread direction; the fire levels are divided into three levels: level one, level two, and level three from high to low;
[0060] The internal space of the warehouse is divided into four areas in a cross shape, and a rotatable water gun is installed at the center of each area; calculate the distance from each water gun to the fire source, and the water gun with the shortest distance undertakes the main fire extinguishing task. At the same time, a water gun in the direction of the flame spread is selected to undertake the auxiliary fire extinguishing task to inhibit the spread of the fire; for each level increase in the fire level, an additional water gun with a relatively short distance undertakes the main fire extinguishing task; when the fire level is level one, all four water guns are put into fire extinguishing, thus effectively controlling and extinguishing the fire.
[0061] The parts not described in this invention are applicable to the prior art.
Claims
1. A warehouse fire intelligent detection and extinguishing method based on multimodal information fusion, characterized in that: The following steps are involved: Step 1: Collect visual images, thermal images, and smoke concentration time series; Step 2: Input smoke concentration time series, visual images and thermal images into the multimodal fire prediction model to predict the fire level; The smoke density time series is used to extract smoke density features, the thermal imaging image is used to extract temperature features, and the visual image is used to extract flame features; The three features of smoke density, temperature and flame are enhanced respectively by using the fusion method driven by correlation to obtain enhanced smoke density feature, enhanced temperature feature and enhanced flame feature. The enhanced temperature features and flame features are spliced to obtain shallow fusion features; The enhanced smoke density features and shallow fusion features are respectively processed through MLP and then spliced to obtain the middle fusion features; at the same time, the enhanced smoke density features, temperature features and flame features are spliced to obtain the enhanced fusion features; The attention gating mechanism is used to assign weights to each fusion feature. The fusion features are weighted and fused according to the following formula to obtain the deep fusion feature. In the formula, C final represents the deep fusion feature, G early , G middle , G AF Represent the shallow fusion features C early , middle-level fusion features C middle And enhanced fusion feature C AF The weight of The deep fusion features are passed through the Softmax function to obtain the prediction score, and then the fire level is obtained; Step 3: Input the visual image into the flame target detection model to predict the fire source and the range of the fire; according to the predicted flame boundary box, the optical flow method is used to predict the flame spread direction; Step 4: Dispatch water guns according to the fire level, fire source and flame spread direction; The internal space of the warehouse is divided into four areas in the form of a cross, and a rotatable water gun is installed at the center of each area; the distance from each water gun to the fire source is calculated, and the water gun closest to it is responsible for the main fire-fighting task, and a water gun in the direction of flame spread is selected to undertake the auxiliary fire-fighting task; the fire level is divided into one, two, and three levels from high to low. With each increase in level, a water gun that is closer is added to undertake the main fire-fighting task; when the fire level is one, all four water guns are used for fire extinguishing.
2. The warehouse fire intelligent detection and extinguishing method based on multimodal information fusion according to claim 1 is characterized in that: The flame target detection model is obtained by replacing the C2f modules of the YOLOv10 backbone network one by one by the C2f_RepViTSimAM modules; the C2f_RepViTSimAM module is obtained by replacing the Bottleneck structure of the C2f module by the RepViTSimAM module; in the RepViTSimAM module, the input feature is divided into three sub-features along the channel dimension, the second sub-feature is subjected to 3×3 depth-separable convolution and the third sub-feature is subjected to 1×1 depth-separable convolution, and then is spliced with the first sub-feature in the channel dimension, and the spliced features are passed through the SimAM attention mechanism to obtain attention features; After the attention features pass through the feedforward neural network, the output features of the RepViTSimAM module are obtained.
3. The warehouse fire intelligent detection and extinguishing method based on multimodal information fusion according to claim 1 or 2 is characterized in that smoke The concentration time series is input into the long short-term memory network, and the hidden state of the last time step is used as the smoke concentration feature; the thermal imaging image is input into the ResNet50 network, and the output feature of the second stage of the ResNet50 network is used as the temperature feature; The visual image is input into the flame target detection model, and the output features of the flame target detection model backbone network are used as flame features.