Lightweight fire detection method based on bimodal image fusion
By combining a deep learning model with RGB visible light and thermal infrared images, the problem of low accuracy and large response delay in fire detection technology under complex environments is solved, achieving efficient flame detection and real-time early warning, which is suitable for embedded devices.
Patent Information
- Application Number
- CN202511114530.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-18
AI Technical Summary
Existing fire detection technologies suffer from low detection accuracy, large response delays, complex models, and low inference efficiency in complex environments, making it difficult to meet the real-time early warning requirements of embedded devices.
A deep learning-based fire detection model is adopted. Through feature extraction, feature fusion enhancement and adaptive detection output modules, combined with RGB visible light images and thermal infrared images, a lightweight ResNet-18 variant, MSR-CSSA module and dynamic weight generation strategy are used to achieve differentiated expression and complementarity of cross-modal information.
It improves the accuracy and real-time performance of fire detection, reduces the number of parameters, enhances the computational efficiency of the model, adapts to flame detection in complex environments, and meets the real-time early warning requirements of embedded devices.
Smart Images

Figure CN120976860A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and fire detection technology, specifically relating to a lightweight fire detection method based on dual-modal image fusion. Background Technology
[0002] Fire detection methods that rely on physical signals from devices such as ionization smoke sensors and thermocouples offer advantages such as low cost and ease of deployment. However, in industrial production environments, dust, high humidity, and frequent airflow can interfere with the normal operation of sensors, resulting in a false alarm rate as high as 30%-50%. Furthermore, these fire detection methods have a weak ability to identify smoldering fires, often missing the optimal time for initial fire response.
[0003] With the development of computer vision technology, image-based fire detection technology has emerged. Early traditional image processing methods identified flame features through color space conversion and contour extraction, achieving certain detection results in ideal environments. However, in real-world scenarios, flame color can be easily confused with red objects, and factors such as changes in lighting and smoke obstruction can severely affect detection accuracy. The false positive rate increases significantly at night or under complex lighting conditions, making it difficult to meet the needs of practical applications.
[0004] The introduction of deep learning technology has brought new breakthroughs to fire detection. Convolutional neural network models such as the YOLO series, ResNet, and EfficientDet can automatically extract deep features of flames, achieving end-to-end fire identification and localization, and performing excellently on standard datasets. However, these models still have limitations in practical applications. When facing tiny flame targets (flame pixels accounting for less than 1%), the false negative rate exceeds 25%; in low-visibility environments such as nighttime or dense smoke, the accuracy drops to 60%-70%. In addition, models based solely on visible light images cannot perceive the thermal radiation characteristics of flames, and their detection performance is significantly reduced in scenarios with smoke obscuring the image or drastic changes in lighting.
[0005] To address the aforementioned issues, dual-band detection technology fusing infrared and visible light images has become a research hotspot. However, current fusion methods have significant shortcomings. Pixel-level fusion methods (such as simple image overlay or channel stitching) do not fully consider the semantic and spatial resolution differences between visible light and infrared images, easily generating redundant features and affecting detection performance. Decision-level fusion methods (weighted averaging after separate detection) ignore the interaction and collaboration of cross-modal information, resulting in slow response to dynamic flame targets and an increase in detection latency of 1-2 seconds. Furthermore, flames in real-world scenarios exhibit large scale variations and blurred boundaries. Existing multi-scale feature processing structures (such as FPN and PANet) suffer from poor feature fusion performance when processing dual-band images due to a lack of design considerations for modal differences. Moreover, to meet the real-time warning requirements of edge devices, both model computation efficiency and inference speed must be considered. However, existing parallel dual-stream network structures have a large number of parameters, insufficient feature interaction, and inference speeds generally below 10 frames per second, making efficient deployment on embedded devices difficult.
[0006] In summary, existing fire detection technologies suffer from low detection accuracy, long response delays, difficulty adapting to scale changes, complex models, and low inference efficiency in complex environments. There is an urgent need for innovative technologies to meet the pressing demand for efficient fire early warning systems in the intelligent security field. Summary of the Invention
[0007] To address the aforementioned issues, this invention provides a lightweight fire detection method based on dual-modal image fusion, employing a deep learning-based fire detection model to achieve regional fire detection; the fire detection model includes a feature extraction module, a feature fusion enhancement module, and an adaptive detection output module;
[0008] The processing steps of the fire detection model include the following:
[0009] S1. Acquire RGB visible light images and Thermal infrared images of the same scene to form a set of samples;
[0010] S2. Input the sample into the feature extraction module to obtain RGB-specific feature maps and Thermal-specific feature maps; the feature extraction module includes a backbone network, an RGB feature branch, and a Thermal feature branch;
[0011] S3. Input the RGB-specific feature map and the Thermal-specific feature map into the feature fusion enhancement module to obtain a fused feature map; the feature fusion enhancement module includes the MSR-CSSA module;
[0012] S4. Input the fused feature map into the adaptive detection output module to obtain the prediction result; the adaptive detection output module includes a three-branch detection head and a dynamic weight generation module.
[0013] The beneficial effects of this invention are:
[0014] This invention proposes a lightweight fire detection method based on dual-modal image fusion, addressing the issues of low detection accuracy and poor real-time performance in complex environments. Performance breakthroughs are achieved through multi-module collaborative innovation. In the modal feature extraction module, a lightweight ResNet-18 variant with parameter sharing is used to extract general features, reducing the number of parameters by 35%. The RGB feature branch enhances texture feature extraction through residual blocks and the SE module, while the Thermal feature branch utilizes residual blocks and CBAM to strengthen heat source feature capture, outputting modal-specific feature maps to achieve differentiated expression and complementarity of dual-modal information. Cross-modal deep fusion is achieved through the MSR-CSSA module. In the channel dimension, the dominant channel is dynamically enhanced based on the efficient channel attention (ECA) mechanism, while in the spatial dimension, pooling and convolution are used to locate the flame region. A three-branch detection head is designed to process RGB, Thermal, and fused features respectively, and a dynamic weight generation strategy is introduced to dynamically weight the fused detection results according to environmental conditions, achieving scene-adaptive detection.
[0015] A multi-task loss function including FocalLoss, GIoULoss, and InfoNCELoss is constructed, combined with contrastive learning and knowledge distillation strategies. Contrastive learning forces semantic alignment of bimodal features, while knowledge distillation guides the fusion detector to learn visible light feature representations. Efficient network training is achieved by dynamically adjusting the loss weights. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the overall structure of the dual-modal fire detection method described in this embodiment of the invention.
[0017] Figure 2 Here is a structural diagram of the MSR-CSSA module;
[0018] Figure 3 This is a schematic diagram of the three detection heads and the dynamic weighting module. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Figure 1 This is an overall framework diagram of the fire detection model shown in some embodiments of the present invention.
[0021] Some embodiments of the present invention provide a lightweight fire detection method based on dual-modal image fusion, such as... Figure 1 As shown, a fire detection model based on deep learning is used to realize regional fire detection; the fire detection model includes a feature extraction module, a feature fusion enhancement module, and an adaptive detection output module;
[0022] The processing steps of the fire detection model include the following:
[0023] S1. Acquire RGB visible light images and thermal infrared images of the same scene to form a set of samples.
[0024] S2. Input the sample into the feature extraction module to obtain RGB-specific feature maps and Thermal-specific feature maps; the feature extraction module includes a backbone network, an RGB feature branch, and a Thermal feature branch.
[0025] In some embodiments, step S2 specifically includes:
[0026] S21. Input the RGB visible light image and the thermal infrared image into the backbone network to obtain a general feature map.
[0027] In some embodiments, a lightweight ResNet-18 variant is used as the backbone network, which shares parameters between the RGB feature branches and the Thermal feature branches. The backbone network includes a Conv1 layer, a MaxPool layer, and two residual blocks, wherein:
[0028] The Conv1 layer uses a 7×7 convolution kernel with a stride of 2 and padding of 3. It has 4 input channels (3 channels for the RGB image + 1 channel for the Thermal image) and 64 output channels. The convolution operation of the Conv1 layer can be represented as follows:
[0029]
[0030] Among them, F conv1 (m,n,c1) represents the value of the output feature of Conv1 layer at coordinate (m,n) in the c1th channel, W(i,j,c0,c1) represents the convolution kernel parameters, and I n b(c1) represents the input image of the Conv1 layer, and b(c1) represents the bias term.
[0031] The MaxPool layer uses a 3×3 pooling kernel with a stride of 2, and both the input and output channels have 64 channels.
[0032] Each residual block comprises two 3×3 convolutional layers, each with 64 input and 64 output channels; and each 3×3 convolutional layer has a stride of 1. The residual blocks of this invention employ identity mapping, meaning the input of the first 3×3 convolutional layer is directly skipped to the output of the second 3×3 convolutional layer. Since both convolutional layers have 64 input and 64 output channels and a stride of 1, the skip connections do not require dimensionality adjustment, meaning the number of channels remains constant.
[0033] In this embodiment, the input size of the Conv1 layer is 640×640×4, and the output size is 320×320×64; the MaxPool layer downsamples the output of the Conv1 layer, and its output size is 160×160×64; after passing through a residual block containing a convolution operation with a stride of 2, the final general feature map output by the backbone network has a size of 80×80×128.
[0034] S22. Input the general feature map into the RGB feature branch to obtain the RGB specific feature map.
[0035] In some embodiments, the RGB feature branch includes two residual blocks and one SE (Squeeze-and-Excitation) module. The general feature map is processed through two residual blocks to obtain feature map F. res , feature map F res The input to the SE module yields RGB-specific feature maps; the processing steps of the SE module include:
[0036] Extrusion operation: on feature map F res Perform global average pooling to obtain the channel descriptor z:
[0037]
[0038] Where z(c) represents the channel descriptor corresponding to the c-th channel; H and W are feature maps F res In this embodiment, H=80 and W=80; F res (i,j,c) represents the pixel value at coordinate (i,j) of the c-th channel of the feature map.
[0039] Activation operation: The channel descriptor z is passed through two fully connected layers (the first fully connected layer compresses the number of channels to 8, and the second fully connected layer restores the number of channels to 128) to learn the dependencies between channels and obtain the channel feature weights S.
[0040]
[0041] Where W1 and W2 are the weights of the fully connected layer, σ is the sigmoid activation function, and δ is the ReLU activation function.
[0042] Weighting operation: Weighting feature map F res Multiplying the feature by the channel feature weights yields the RGB-specific feature map F. RGB In this embodiment, the RGB-specific feature map F RGB The dimensions are 80×80×256.
[0043] S23. Input the general feature map into the Thermal feature branch to obtain the Thermal-specific feature map.
[0044] In some embodiments, the Thermal feature branch includes two residual blocks and one CBAM (Convolutional Block Attention Module). The general feature map is processed through two residual blocks to obtain the feature map F'. res , feature map F' res Inputting the CBAM module yields Thermal-specific feature maps; the CBAM module's processing includes:
[0045] Channel attention: for feature map F' res Global average pooling and global max pooling are performed respectively to obtain the feature map F. avg and feature map F max Feature map F is processed using a multilayer perceptron. avg and feature map F max Gain channel attention:
[0046]
[0047] MLP stands for Multilayer Perceptron, M c For channel attention.
[0048] Spatial attention: attention to feature map F' in the channel dimension res Perform average pooling and max pooling respectively to obtain the feature map. and feature map ; feature map and feature map Spatial attention is obtained after splicing through a 7×7 convolutional layer:
[0049]
[0050] Among them, M s For spatial attention, Conv 7×7 It is a 7×7 convolutional layer, and [;] represents the splicing operation.
[0051] Feature weighting: Combining channel attention, spatial attention, and feature map F' res Multiplying them together yields the Thermal-specific feature map F. Thermal The dimensions are 80×80×256.
[0052] S3. Input the RGB-specific feature map and the Thermal-specific feature map into the feature fusion enhancement module to obtain the fused feature map; the feature fusion enhancement module includes the MSR-CSSA module (i.e., the multi-scale residual dynamic fusion module).
[0053] In some embodiments, the RGB-specific feature map and the Thermal-specific feature map are concatenated along the channel dimension to obtain a cross-modal stitched image:
[0054]
[0055] Among them, F concat (i,j,k1) represents the pixel value at coordinate (i,j) in the k1th channel of the cross-modal stitched image, F RGB (i,j,k1) represents the pixel value at coordinate (i,j) of the k1th channel of the RGB-specific feature map, F Thermal (i,j,k1) represents the pixel value at coordinate (i,j) of the k1th channel of the Thermal-specific feature map.
[0056] The cross-modal stitched image is input into the MSR-CSSA module to obtain the fused feature map; the MSR-CSSA module includes a multi-scale residual path and a dynamic gating fusion module.
[0057] Figure 2 This is a structural diagram of the MSR-CSSA module shown in some embodiments of the present invention.
[0058] In some embodiments, such as Figure 2 As shown, the multi-scale residual path includes small-scale branches and large-scale branches;
[0059] In the small-scale branch, the cross-modal stitched image is passed through the first convolutional layer to obtain the first feature map. The first feature map is then concatenated with the residual of the cross-modal stitched image to obtain the small-scale output F. s The first convolutional layer uses a 3×3 convolutional kernel with a stride of 1, padding of 1, and a kernel count of 256.
[0060] In the large-scale branch, the cross-modal stitched image is passed through a second convolutional layer to obtain a second feature map. The second feature map is then concatenated with the residual of the cross-modal stitched image to obtain the large-scale output F. l The second convolutional layer uses a 7×7 convolutional kernel with a stride of 1, padding of 3, and a kernel count of 256.
[0061] Small-scale output F s With large-scale output F l The concatenation is performed along the channel dimension, and the concatenation result is mapped to the feature dimension dmodel = 1024 through a linear transformation to form a feature F adapted to the Transformer multi-head attention mechanism. merge .
[0062] In some embodiments, the dynamic gating fusion module employs a Transformer-based multi-head attention mechanism; the multi-head attention mechanism has 8 heads, d model =1024. Specifically, it includes:
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069] Among them, F Fusion This represents the fused feature map, which is 80×80×1024; W O , , , For the projection matrix, head i For the i-th attention head, Q i K i V i For the query, key, and value of the i-th attention head, d k =128 is the dimension of each attention head.
[0070] S4. Input the fused feature map into the adaptive detection output module to obtain the prediction result; the adaptive detection output module includes a three-branch detection head and a dynamic weight generation module.
[0071] Figure 3 This is a structural diagram of an adaptive detection output module according to some embodiments of the present invention.
[0072] In some embodiments, the three-branch detection head includes an RGB detection head, a Thermal detection head, and a fusion detection head;
[0073] The RGB detection head receives RGB-specific feature maps. These feature maps are first passed through two 3×3 convolutional layers. Each 3×3 convolutional layer has C_rgb kernels, a stride of 1, and padding of 1. Each 3×3 convolutional layer is followed by a BatchNormalization layer and a ReLU activation function. After this processing, the first processed feature is generated. This first processed feature is then input into two parallel 1×1 convolutional branches. One 1×1 convolutional branch has the number of kernels equal to the number of categories (corresponding to flame and non-flame) and outputs the category prediction result. The other 1×1 convolutional branch has 4 kernels and outputs the bounding box prediction result (the bounding box coordinates are normalized, i.e., [x_min, y_min, x_max, y_max], where x_min and x_max represent the minimum and maximum values of the horizontal coordinate, and y_min and y_max represent the minimum and maximum values of the vertical coordinate, all within the range [0,1]). The outputs of the two 1×1 convolutional branches are merged to form P. RGB ;
[0074] The thermal detection head receives a thermal-specific feature map, which is passed through two identical 3×3 convolutional layers. Each 3×3 convolutional layer has C_thermal kernels, a stride of 1, and padding of 1. Each 3×3 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function. This process generates a second processed feature map, which is then input into two parallel 1×1 convolutional branches. These branches output the class prediction result and the bounding box prediction result, respectively. The outputs of the two 1×1 convolutional branches are then combined to form P. Thermal ;
[0075] The fusion detection head receives the fused feature map output by the MSR-CSSA fusion module. The number of channels in this fused feature map is twice the maximum number of channels in the RGB-specific feature map and the Thermal-specific feature map, i.e., max(C_rgb, C_thermal)×2. The fused feature map is passed through two 3×3 convolutional layers. The number of kernels in each 3×3 convolutional layer is the same as the number of channels in the fused feature map, the stride is 1, and the padding is 1. Each 3×3 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function. After the above processing, a third processed feature is generated. The third processed feature is input into two parallel 1×1 convolutional branches, which output the class prediction result and the bounding box prediction result, respectively. The outputs of the two 1×1 convolutional branches are merged into P. Fusion .
[0076] The dynamic weight generation module receives the fused feature map, performs global average pooling on the fused feature map to obtain a 1×1×1024 compressed vector, and then performs a 1×1 convolution and a softmax function on the compressed vector to obtain the RGB weights ω. RGB Thermal weights ω Thermal and fusion weight ω Fusion ;
[0077] The final prediction result P is:
[0078] .
[0079] In some embodiments, the present invention first pre-trains the three-branch detection head using a dedicated supervision mechanism, and then performs joint optimization during the training process of the fire detection model.
[0080] The dedicated oversight mechanism includes:
[0081] For the RGB detection head, it is pre-trained using an RGB visible light image dataset, where the RGB visible light images are labeled with category labels and bounding box coordinates;
[0082] For the Thermal detection head, the Thermal infrared image dataset is used for pre-training, where the Thermal infrared images are labeled with category labels and bounding box coordinates;
[0083] For the fusion detection head, a joint image dataset is used for pre-training. The joint image dataset consists of images aligned with RGB visible light images and thermal infrared images. Each aligned image is labeled with a category label and bounding box coordinates.
[0084] In some embodiments, the joint optimization process of the fire detection model according to the present invention is as follows:
[0085] First, the process of acquiring the image dataset used in the fire detection model training includes:
[0086] S11. Using industrial-grade surveillance cameras (resolution no less than 1920×1080) and infrared thermal imagers (thermal sensitivity ≤50mK), RGB visible light images and thermal infrared images are simultaneously acquired in different indoor and outdoor scenarios. The bounding box coordinates and category labels (such as flame, non-flame) of the flame target are annotated for each image in VOC format. The scenarios cover daytime, nighttime, industrial plants, forests, etc. Environmental information (such as temperature and humidity) of the scene is recorded during the image acquisition process.
[0087] S12. Take the RGB visible light image and Thermal infrared image simultaneously acquired in the same scene as a pair of images, and perform standardization processing on each pair of images, uniformly scaling the RGB visible light image and Thermal infrared image to 640×640 pixels to obtain a sample set; the standardization process includes:
[0088] For RGB visible light images, a normalization formula is used to process them. The normalization formula is as follows:
[0089]
[0090] Where I(i,j,k) represents the pixel value at coordinates (i,j) in the k-th channel of the RGB visible light image, and k=1,2,3 correspond to the R channel, G channel, and B channel, respectively; norm (i,j,k)∈[0,1] represents the pixel value at coordinate (i,j) in the k-th channel of the normalized RGB visible light image;
[0091] For thermal infrared images, the temperature value of each pixel in the thermal infrared image is mapped to the grayscale range of [0, 255]. The mapping formula is as follows:
[0092]
[0093] Where T(i,j) represents the temperature value of the thermal infrared image at coordinate (i,j), and G(i,j) represents the temperature value of the mapped thermal infrared image at coordinate (i,j); T max T min These represent the maximum and minimum temperatures of the thermal infrared image, respectively.
[0094] S13. Divide all standardized samples into training, validation, and test sets in a ratio of 7:1:2. The training set is used to update model parameters, the validation set is used to tune hyperparameters, and the test set is used to evaluate the final performance of the model.
[0095] Secondly, the samples are input into the fire detection model, and the prediction results are output.
[0096] Then, the multi-task loss is calculated based on the prediction results, and the parameters of the fire detection model are jointly optimized based on the multi-task loss.
[0097] The multi-tasking loss L is as follows:
[0098]
[0099] Among them, L cls L represents the classification loss.reg L represents the regression loss. con L represents the comparative loss. dis λ1, λ2, λ3, and λ4 represent distillation losses, and λ4 represent weighting coefficients. In this embodiment, λ1=0.4, λ2=0.3, λ3=0.2, and λ4=0.1.
[0100] For classification loss L cls The calculation formula is as follows:
[0101]
[0102] Where p represents the category prediction result (i.e., the category prediction probability), and α and γ represent hyperparameters. In this embodiment, α = 0.25 and γ = 2.
[0103] For regression loss L reg The calculation formula is as follows:
[0104]
[0105] Among them, b pred b represents the bounding box prediction result. gt This represents the true bounding box. GIoU() represents the generalized intersection-union ratio, which is calculated as follows:
[0106]
[0107] IoU stands for Intersection over Union, which is the ratio of the area of the intersection of two bounding boxes to the area of their union. A and B represent the two bounding boxes to be compared, respectively; C represents the area of the smallest bounding rectangle that can contain both A and B.
[0108] For the contrastive loss, RGB-Thermal positive and negative sample pairs are constructed between the input (i.e., the cross-modal stitched image of the RGB-specific feature map and the Thermal-specific feature map) and the output (i.e., the fused feature map) of the MSR-CSSA module. The positive sample pair is constructed by extracting the feature vectors at the same spatial coordinates (m,n) in the RGB feature map and the Thermal feature map of the same sample (for example, the feature vector q_rgb at coordinates (m,n) in the RGB feature map is used as the query sample q, and the corresponding feature vector k_thermal at the same coordinates in the Thermal feature map is used as the positive sample). Negative sample pairs are generated using two strategies: cross-sample negative sampling and sample space negative sampling. (Cross-sample negative sampling randomly selects K feature vectors from the Thermal feature maps of other samples within the batch.) As negative samples, spatial negative sampling also selects feature vectors at different spatial coordinates from the Thermal feature map of the same sample as negative samples; contrastive loss Lcon The calculation formula is:
[0109]
[0110] Where sim() represents cosine similarity calculation, τ is the temperature hyperparameter, usually τ=0.1; q is the query sample, k + For positive samples, k i K represents the number of negative samples.
[0111] For the distillation loss, a knowledge distillation strategy is adopted, using the pre-trained RGB detector head as the teacher model and the fused detector head as the student model. During joint training, the parameters of the RGB detector head are fixed at the pre-trained values, and only the parameters of the fused detector head are optimized and updated. The pre-training process uses RGB single-modal data to train the RGB detector head separately to ensure stable flame recognition capabilities. The distillation loss L... dis The calculation formula is:
[0112]
[0113] In the formula, B is the batch size and N is the number of prediction results for each sample. This represents the probability distribution of the teacher model's prediction of the i-th sample (output via softmax). This represents the probability distribution of the student model for the same prediction result (i.e., the i-th prediction result of the b-th sample). The prediction probability includes both classification probability and bounding box probability. The classification probability corresponds to the softmax output of the flame / non-flame category, and the bounding box probability is the probability distribution obtained by converting the bounding box coordinates into a probability distribution using a Gaussian kernel.
[0114] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "setting," "connection," "fixing," "rotation," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Unless otherwise explicitly limited, those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0115] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A lightweight fire detection method based on dual-modal image fusion, characterized in that, A fire detection model based on deep learning is used to realize regional fire detection; the fire detection model includes a feature extraction module, a feature fusion enhancement module, and an adaptive detection output module; The processing steps of the fire detection model include the following: S1. Acquire RGB visible light images and Thermal infrared images of the same scene to form a set of samples; S2. Input the sample into the feature extraction module to obtain RGB-specific feature maps and Thermal-specific feature maps; the feature extraction module includes a backbone network, an RGB feature branch, and a Thermal feature branch; S3. Input the RGB-specific feature map and the Thermal-specific feature map into the feature fusion enhancement module to obtain a fused feature map; the feature fusion enhancement module includes the MSR-CSSA module; S4. Input the fused feature map into the adaptive detection output module to obtain the prediction result; the adaptive detection output module includes a three-branch detection head and a dynamic weight generation module.
2. The lightweight fire detection method based on dual-modal image fusion according to claim 1, characterized in that, In step S2, the RGB visible light image and the Thermal infrared image are input together into the backbone network to obtain a general feature map; the general feature map is then input into the RGB feature branch and the Thermal feature branch respectively to obtain RGB-specific feature maps and Thermal-specific feature maps; the backbone network includes a Conv1 layer, a MaxPool layer, and two residual blocks; the RGB feature branch includes two residual blocks and one SE module; the Thermal feature branch includes two residual blocks and one CBAM module; wherein: The Conv1 layer uses a 7×7 convolution kernel with a stride of 2, padding of 3, 4 input channels, and 64 output channels. The MaxPool layer uses a 3×3 pooling kernel with a stride of 2, and both the number of input channels and the number of output channels are 64. Each residual block consists of two 3×3 convolutional layers, each with 64 input and 64 output channels; the input of the first 3×3 convolutional layer is connected to the output of the second 3×3 convolutional layer in a skip connection.
3. The lightweight fire detection method based on dual-modal image fusion according to claim 2, characterized in that, Input the general feature map into the RGB feature branch: General The feature map F is obtained after passing through two residual blocks. res , feature map F res The input to the SE module yields RGB-specific feature maps; the processing steps of the SE module include: For feature map F res Perform global average pooling to obtain channel descriptors; The channel descriptors are passed through two fully connected layers to obtain the channel feature weights; Feature map F res Multiplying the feature by the channel feature weights yields the RGB-specific feature map.
4. A lightweight fire detection method based on dual-modal image fusion according to claim 2, characterized in that, Input the general feature map into the Thermal feature branch: the general feature map is processed through two residual blocks to obtain the feature map F'. res , feature map F' res Inputting the CBAM module yields Thermal-specific feature maps; The processing steps of the CBAM module include: For feature map F' res Global average pooling and global max pooling are performed respectively to obtain the feature map F. avg and feature map F max The feature map F is processed using a multilayer perceptron. avg and feature map F max Gain channel attention; Feature map F' in the channel dimension res Perform average pooling and max pooling respectively to obtain the feature map. and feature map ; feature map and feature map Spatial attention is obtained by using a 7×7 convolutional layer after splicing; Integrate channel attention, spatial attention, and feature map F' res Multiplying them together yields a Thermal-specific feature map.
5. A lightweight fire detection method based on dual-modal image fusion according to claim 1, characterized in that, In step S3, the RGB-specific feature map and the Thermal-specific feature map are concatenated along the channel dimension to obtain a cross-modal concatenation map. The cross-modal concatenation map is then input into the MSR-CSSA module to obtain a fused feature map. The MSR-CSSA module includes a multi-scale residual path and a dynamic gating fusion module.
6. A lightweight fire detection method based on dual-modal image fusion according to claim 5, characterized in that, The multi-scale residual path includes small-scale branches and large-scale branches; In the small-scale branch, the cross-modal stitched image is passed through the first convolutional layer to obtain the first feature map. The first feature map is then concatenated with the residual of the cross-modal stitched image to obtain the small-scale output. The first convolutional layer uses a 3×3 convolutional kernel with a stride of 1, padding of 1, and a kernel count of 256. In the large-scale branch, the cross-modal stitched image is passed through the second convolutional layer to obtain the second feature map. The second feature map is then concatenated with the residual of the cross-modal stitched image to obtain the large-scale output. The second convolutional layer uses a 7×7 convolutional kernel with a stride of 1, padding of 3, and a kernel count of 256. The small-scale output and the large-scale output are concatenated in the channel dimension before being output.
7. A lightweight fire detection method based on dual-modal image fusion according to claim 5, characterized in that, The dynamic gating fusion module adopts a Transformer-based multi-head attention mechanism; the number of heads in the multi-head attention mechanism is 8.
8. A lightweight fire detection method based on dual-modal image fusion according to claim 1, characterized in that, The three-branch detection head includes an RGB detection head, a Thermal detection head, and a fusion detection head, wherein: The RGB detection head receives RGB-specific feature maps, which are then passed through two 3×3 convolutional layers to generate the first processed feature. Each 3×3 convolutional layer has C_rgb kernels, a stride of 1, and padding of 1. Each 3×3 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function. The first processed feature is then input into two parallel 1×1 convolutional branches. One 1×1 convolutional branch has the number of kernels equal to the number of classes and outputs the class prediction result. The other 1×1 convolutional branch has 4 kernels and outputs the normalized bounding box coordinates as the bounding box prediction result. The outputs of the two 1×1 convolutional branches are merged into a single P. RGB ; The thermal detection head receives a thermal-specific feature map. This map is processed through two 3×3 convolutional layers to generate a second processing feature. Each 3×3 convolutional layer has C_thermal kernels, a stride of 1, and padding of 1. Each 3×3 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function. The second processing feature is then input into two parallel 1×1 convolutional branches, which output the class prediction result and the bounding box prediction result, respectively. The outputs of the two 1×1 convolutional branches are combined to form P. Thermal ; The fusion detection head receives the fused feature map and generates a third processing feature map by passing it through two 3×3 convolutional layers. Each 3×3 convolutional layer has the same number of kernels as the number of channels in the fused feature map, a stride of 1, and padding of 1. Each 3×3 convolutional layer is followed by a Batch Normalization layer and a ReLU activation function. The number of channels in the fused feature map is twice the maximum number of channels in both the RGB-specific and Thermal-specific feature maps. The third processing feature map is then input into two parallel 1×1 convolutional branches, which output the class prediction result and the bounding box prediction result, respectively. The outputs of the two 1×1 convolutional branches are merged into a single P. Fusion ; The dynamic weight generation module receives the fused feature map, performs global average pooling on the fused feature map to obtain a compressed vector, and then performs a 1×1 convolution and a softmax function on the compressed vector to obtain the RGB weights ω. RGB Thermal weights ω Thermal and fusion weight ω Fusion ; The final prediction result P is: 。 9. A lightweight fire detection method based on dual-modal image fusion according to claim 1, characterized in that, During training, multiple loss functions are used to jointly optimize the parameters of the fire detection model, including: Classification loss L cls : , Where p represents the category prediction result, and α and γ represent hyperparameters; Regression loss L reg : , Among them, b pred b represents the bounding box prediction result. gt Represents the true bounding box; GIoU() represents the generalized intersection-union ratio; Comparison loss L con : , Where sim() represents cosine similarity calculation, τ is the temperature hyperparameter, usually τ=0.1; q is the query sample, k + For positive samples, k i For negative samples, K is the number of negative samples; Distillation loss L dis The calculation formula is: , Where B represents the batch size and N represents the number of prediction results for each sample. Let represent the probability distribution of the teacher model for the i-th prediction result of the b-th sample. Let represent the probability distribution of the student model's prediction of the i-th sample.
Citation Information
Cited By
Fire point detection method and system based on cross-modal feature marshalling
CN121305474A
A fire point detection method and system based on cross-modal feature grouping
CN121305474B