Smoke Detection Method, System, Medium, Electronic Device and Smoke Detection Model
Through the design of multi-layer network structure and feature fusion module, the smoke detection model's attention to the smoke moving area is enhanced, and the problems of low accuracy and high false alarm rate in the prior art are solved, achieving the effect of rapid and accurate smoke positioning.
Patent Information
- Application Number
- CN202410187760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-02-20
AI Technical Summary
The existing smoke detection methods have low accuracy and are prone to missed and missed detection, especially in forest fire scenes, which are small and susceptible to disturbances.
A smoke detection model with a multi-layer network structure is adopted, including a feature extraction network and a target feature enhancement network, combined with a multi-scale feature fusion module, and the attention mechanism and improved inception module are used to improve the attention and feature extraction capabilities of smoke motion areas.
It improves the accuracy and reliability of smoke detection, reduces the false alarm rate, and can quickly and accurately locate the smoke occurrence location. It is suitable for outdoor macro stations, urban smart light poles and fire warnings in indoor scenes.
Smart Images

Figure CN118097547B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of physics, in particular to neural network data processing technology, and more particularly to a smoke detection method, system, medium, electronic device and smoke detection model. Background Art
[0002] Traditional smoke detection methods mainly use smoke sensors for detection. However, it is necessary for the smoke to reach a certain concentration to trigger the smoke sensor to alarm, and the detection distance of the sensor is limited, often resulting in missed detection and false detection. Since the shape of the smoke is not fixed and there are many smoke-like objects that are very similar in color or appearance to the smoke, the accuracy of the existing smoke detection algorithms is low. Summary of the Invention
[0003] The purpose of the present invention is to provide a smoke detection method, system, medium, electronic device and smoke detection model to solve the problems pointed out in the above background art.
[0004] In a first aspect, the present invention provides a smoke detection model applied to the smoke detection of a target area. The smoke detection model is a multi-layer network structure, and the smoke detection model includes: a feature extraction network and a target feature enhancement network; wherein, the feature extraction network is used to receive the monitoring data of the target area and to extract features from the monitoring data to obtain multi-scale features; the target feature enhancement network is connected to the feature extraction network, and the target feature enhancement network is used to receive the multi-scale features and to obtain target features based on the multi-scale features, so as to realize the smoke detection of the target area based on the target features.
[0005] In the present invention, aiming at the problem of low accuracy of the existing smoke detection algorithms, a new type of smoke detection model is proposed. Through the design of the target feature enhancement network, the smoke detection model can better focus on the smoke movement area, enhance its recognition ability for small targets, and thus improve the accuracy of smoke detection.
[0006] In an implementation manner of the first aspect, the target feature enhancement network is used to obtain target features based on the multi-scale features, including: the target feature enhancement network performs summation and averaging processing on all frames of the multi-scale features to obtain a background map; the target feature enhancement network subtracts the background map from each frame of the multi-scale features to obtain a foreground region detection map; the target feature enhancement network performs feature enhancement on the foreground region detection map to obtain the target features.
[0007] In an implementation of the first aspect, the target feature enhancement network enhances the features of the foreground region detection map to obtain the target features, including: the target feature enhancement network performs convolution and non-linear processing on the foreground region detection map to obtain a processed feature map; the target feature enhancement network adds an attention mechanism to the processed feature map to obtain the target features.
[0008] In this implementation, by adding an attention mechanism to the smoke detection algorithm, the attention of the smoke detection model to small targets is improved.
[0009] In an implementation of the first aspect, the feature extraction network is used to obtain multiple multi-scale features corresponding to different layers of the multi-layer network structure respectively, so as to obtain multiple target features through the target feature enhancement network based on the multiple multi-scale features; the smoke detection model further includes: a multi-scale feature fusion module; the multi-scale feature fusion module is connected to the target feature enhancement network, and the multi-scale feature fusion module is used to perform feature fusion on the multiple target features in two paths, from low layer to high layer and from high layer to low layer, to obtain an output feature, so as to realize smoke detection of the target region based on the output feature.
[0010] In this implementation, by adding a multi-scale feature fusion module to the smoke detection model, the smoke detection model extracts richer multi-scale features, fully fuses the semantic information of the high-level feature map and the spatial information of the low-level feature map, and reduces the false alarm rate caused by objects with colors similar to smoke.
[0011] In an implementation of the first aspect, the target features are visualized.
[0012] In this implementation, by visualizing the target features, it is convenient for researchers to perform interpretable analysis on the results identified by the smoke detection model, and the accuracy of smoke recognition is improved.
[0013] In an implementation of the first aspect, the multi-layer network structure has multiple different layers using improved inception modules; the convolutional kernel dimensions of the improved inception modules are 1×1×1, 3×3×3, 3×3×3, 5×5×5.
[0014] In this implementation, an improved Inception module is provided. Compared with the traditional Inception module, the convolutional kernel dimensions of the improved Inception module are changed from the original 1×1×1, 3×3×3, 3×3×3, 3×3×3 to 1×1×1, 3×3×3, 3×3×3, 5×5×5. As a result, it is possible to obtain features of different receptive field regions, better judge the irregular characteristics of smoke, capture the global information of smoke, thereby better improving the smoke detection rate and reducing the false alarm rate.
[0015] In a second aspect, the present invention provides a smoke detection method implemented based on the above-mentioned smoke detection model. The smoke detection method includes: obtaining monitoring data of a target area; inputting the monitoring data into the smoke detection model so that the smoke detection model outputs target features to implement smoke detection of the target area based on the target features.
[0016] In an implementation of the second aspect, before the step of inputting the monitoring data into the smoke detection model, the smoke detection method further includes: obtaining a smoke monitoring data set; training the smoke detection model based on the smoke monitoring data set to obtain a trained smoke detection model; the smoke monitoring data set includes a plurality of smoke type monitoring data and a plurality of fog type monitoring data; inputting the monitoring data into the smoke detection model includes: inputting the monitoring data into the trained smoke detection model.
[0017] In this implementation, a real smoke monitoring data set is provided. The monitoring data in the smoke monitoring data set are all obtained by a monitoring system, which can reflect the smoke characteristics generated in the target area under different scenarios. This smoke monitoring data set is of great significance for studying smoke recognition in real scenarios.
[0018] In a third aspect, the present invention provides a smoke detection system implemented based on the above-mentioned smoke detection model. The smoke detection system includes: an acquisition module for obtaining monitoring data of a target area; a detection module for inputting the monitoring data into the smoke detection model so that the smoke detection model outputs target features to implement smoke detection of the target area based on the target features.
[0019] In a fourth aspect, the present invention provides an electronic device. The electronic device includes: a processor and a memory; the memory is used for storing a computer program; the processor is used for executing the computer program stored in the memory so that the electronic device executes the above-mentioned smoke detection method.
[0020] In a fifth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by an electronic device, the above-mentioned smoke detection method is implemented.
[0021] As described above, the smoke detection method, system, medium, electronic device and smoke detection model of the present invention have the following beneficial effects:
[0022] (1) Compared with the prior art, the present invention proposes a smoke detection model based on deep learning. Through the design of the target feature enhancement network and the multi-scale feature fusion module in the smoke detection model, it effectively overcomes the problems that small targets are difficult to detect, the detection effect is unstable, the image resolution is low, and there are a large number of objects similar in color to smoke interfering in fire monitoring, thereby improving the accuracy and reliability of smoke detection.
[0023] (2) The present invention adopts innovative algorithms and technologies, combines advanced computer vision technologies and machine learning methods, and provides a smoke detection method; at the same time, the present invention also considers the real-time performance and efficiency of the smoke detection algorithm to ensure that smoke detection and early warning can be carried out quickly in practical applications; through this smoke detection method, it can be quickly detected whether there is smoke generated and the specific location where the smoke occurs can be accurately located; by applying it to scenarios such as outdoor macro base stations, urban smart lamp posts or indoors for fire early warning, the probability of fire occurrence can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It shows a schematic structural diagram of the smoke detection model described in the embodiment of the present invention.
[0025] Figure 2 It shows a flowchart of the target feature enhancement network described in the embodiment of the present invention for obtaining target features based on multi-scale features.
[0026] Figure 3 It shows a flowchart of the target feature enhancement network described in the embodiment of the present invention for enhancing the features of the foreground region detection map to obtain target features.
[0027] Figure 4 It shows a network architecture diagram of the smoke detection model described in the embodiment of the present invention.
[0028] Figure 5 It shows a network structure diagram of the improved inception described in the embodiment of the present invention.
[0029] Figure 6 It shows a network structure diagram of the SFM module described in the embodiment of the present invention.
[0030] Figure 7 It shows a flowchart of the smoke detection method described in an embodiment of the present invention.
[0031] Figure 8It shows a flowchart of the smoke detection method according to another embodiment of the present invention.
[0032] Figure 9 It shows a schematic structural diagram of the smoke detection system according to an embodiment of the present invention. Detailed implementation manners
[0033] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0034] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The drawings only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0035] Forest fires have always been one of the major environmental disasters with a catastrophic impact on forest wealth. When a forest fire occurs, the fire situation quickly gets out of control, and extinguishing them requires a huge amount of energy, time, and resources. Forest fires can have a great negative impact on natural resources and human life. Accurate and rapid forest fire detection is crucial for reducing losses. Since smoke is always the first sign seen when a forest fire occurs, smoke detection in the forest is of great significance for early warning.
[0036] Existing forest fire smoke detection faces challenges such as small detection targets and interference from a large amount of fog. Therefore, the existing forest fire smoke detection has a low accuracy rate and a high false alarm rate.
[0037] Forest fire smoke mainly has three characteristics: physical, visual, and dynamic. Among them, the detection of physical characteristics mostly uses sensor detection, which requires installing a large number of sensors. The detection distance is limited, and the sensors are easily damaged and aged due to the environment, resulting in a large number of missed detections and false detections. The detection of the other two characteristics can use traditional image processing, machine learning, and deep learning methods for smoke recognition.
[0038] In traditional image processing methods, due to the obvious motion characteristics of smoke, background subtraction has been widely used for preprocessing in smoke recognition to extract smoke candidate regions; in machine learning methods, SVM (full English name: Support Vector Machine, Chinese translation: Support Vector Machine) has been very frequently used to distinguish smoke frames and non-smoke frames. The main characteristics of smoke are visual, temporal, spatial, and frequency characteristics. Usually, detecting a single characteristic alone cannot achieve an ideal effect, and several methods are often used together. Traditional image processing methods and machine learning methods are highly dependent on manually selected features, so it is difficult to apply them to actual environments.
[0039] To address the deficiencies in traditional algorithms, recent researchers have adopted deep learning algorithms. They have respectively established deep network models for self-built datasets from image recognition and video recognition, and proposed a video-based smoke recognition network model for the temporal characteristics between frames in smoke videos, which has improved the accuracy of wildfire smoke detection to a certain extent. Different from the wildfire smoke detection scenario, in the forest fire smoke detection scenario, smoke always appears far from the camera, and the area of smoke usually only occupies a small part of the video frame; at the same time, fog is a natural phenomenon in the forest and is very similar to smoke visually. Therefore, directly applying the wildfire smoke detection model to forest fire smoke detection will result in false alarms and missed detections.
[0040] Refer to Figures 1 to 9 To address the deficiencies in the above deep learning algorithms, the following embodiments of the present invention provide a smoke detection method, system, medium, electronic device, and smoke detection model. Compared with the prior art, the present invention proposes a deep learning-based smoke detection model. Through the design of the target feature enhancement network and multi-scale feature fusion module in this smoke detection model, it effectively overcomes problems such as small targets being difficult to detect, unstable detection effects, low image resolution, and interference from a large number of objects with colors similar to smoke in fire monitoring, thereby improving the accuracy and reliability of smoke detection; the present invention adopts innovative algorithms and technologies, combines advanced computer vision technologies and machine learning methods, and provides a smoke detection method; at the same time, the present invention also considers the real-time performance and efficiency of the smoke detection algorithm to ensure that smoke detection and early warning can be carried out quickly in practical applications; through this smoke detection method, it is possible to quickly detect whether there is smoke generated and accurately locate the specific location where the smoke occurs; by applying it to scenarios such as outdoor macro stations, urban smart lamp posts, or indoors for fire early warning, the probability of fire occurrence can be reduced.
[0041] Such as Figure 1As shown, in one embodiment, the present invention provides a smoke detection model, which is applied to smoke detection in a target area (such as a target forest); specifically, the smoke detection model is a multi-layer network structure, and the smoke detection model includes: a feature extraction network 11 and a target feature enhancement network 12.
[0042] Among them, the feature extraction network 11 is used to receive the monitoring data of the target area and extract features from the monitoring data to obtain multi-scale features; the target feature enhancement network 12 is connected to the feature extraction network 11, and the target feature enhancement network 12 is used to receive the multi-scale features, and based on the multi-scale features to obtain target features, and based on the target features to realize smoke detection in the target area.
[0043] In one embodiment, the monitoring data is a monitoring video.
[0044] As Figure 2 shown, in one embodiment, the target feature enhancement network 12 is used to obtain target features based on the multi-scale features, including:
[0045] Step S11, the target feature enhancement network performs summation and averaging processing on all frames of the multi-scale features to obtain a background map.
[0046] Step S12, the target feature enhancement network subtracts the background map from each frame of the multi-scale features to obtain a foreground area detection map.
[0047] Step S13, the target feature enhancement network performs feature enhancement on the foreground area detection map to obtain the target features.
[0048] As Figure 3 shown, in one embodiment, the target feature enhancement network 12 performs feature enhancement on the foreground area detection map to obtain the target features, including:
[0049] Step S131, the target feature enhancement network performs convolution and non-linear processing on the foreground area detection map to obtain a processed feature map.
[0050] Step S132, the target feature enhancement network adds an attention mechanism to the processed feature map to obtain the target features.
[0051] It should be noted that in smoke recognition, the neural network structure usually adopts an attention mechanism to enable the network to better locate the information of interest and suppress useless information. The attention mechanism is mainly divided into three types: spatial attention model, channel attention model, and spatial and channel hybrid attention model. Among them, the spatial attention model is mainly used to weight the feature maps at different spatial positions in the convolutional neural network so that the network can better focus on the spatial positions most useful for the current task. Its representative model is the spatial transformer network. Channel attention is mainly used to strengthen the feature representations of different channels in the convolutional neural network, extract the importance or correlation information of different channels, so that the network can better focus on the channel information most useful for the current task. Its representative network is SE-Net. The spatial and channel hybrid attention model combines the spatial attention model and the channel attention model, and can achieve better results compared with the spatial transformer network and SE-Net. Its representative network is CBAM.
[0052] It should be noted that SE-Net, the full English name: Squeezeexcitation Network, is a module that distributes attention on the channels of a convolutional feature map and can be embedded into other network structures. It is a conventional technical means in the field. CBAM, the full English name: Convolutional BlockAttention Module, Chinese translation: Convolutional block attention module, is a lightweight attention mechanism for convolutional neural networks, and it is also a conventional technical means in the field.
[0053] In one embodiment, multiple multi-scale features corresponding to different layers of the multi-layer network structure are obtained through the feature extraction network, so as to obtain multiple target features through the target feature enhancement network based on the multiple multi-scale features.
[0054] In this embodiment, the smoke detection model further includes: a multi-scale feature fusion module 13.
[0055] Specifically, the multi-scale feature fusion module 13 is connected to the target feature enhancement network 12. The multi-scale feature fusion module is used to perform feature fusion on the multiple target features along two paths, from low layer to high layer and from high layer to low layer, to obtain output features, so as to implement smoke detection for the target area based on the output features.
[0056] It should be noted that the convolutional neural network extracts the features of the target in a layer-by-layer abstraction manner. The high-level feature map has a relatively large receptive field and strong semantic information representation ability, which is beneficial to accurately detecting the target. However, the resolution of the feature map is low and the geometric information representation ability is weak. The low-level feature map has a relatively small receptive field, strong geometric detail information representation ability, high resolution, and is beneficial to determining the object position information. However, the semantic information representation ability is weak. Therefore, in object detection, these features are often combined to improve the detection effect.
[0057] There are two common types of multi-scale feature fusion networks. The first is the parallel multi-branch network, and the second is the serial skip connection network. Both perform feature extraction under different receptive fields. In the object detection task, the representative network framework in the parallel multi-branch network is the inception network. The basic module inception in this network includes four parallel branch structures, namely 1×1 convolution, 3×3 convolution, 5×5 convolution, and 3×3 max pooling. Finally, the four channels are combined. In the serial multi-scale feature structure in object detection, FPN, PANet, NAS-FPN, and BiFPN are representatives. Feature combination needs to be achieved through skip connections. Among them, BiFPN adopts two paths, top-down and bottom-up, and can optimize multi-scale feature fusion in a more intuitive and interpretable way.
[0058] It should be noted that FPN, PANet, NAS-FPN, and BiFPN all adopt conventional technical means in the field. Among them, FPN is the most basic feature fusion structure, and it is also the simplest and most effective. It combines deep features and shallow features through a top-down operation. PANet combines the top-down and bottom-up structures, with stronger feature fusion ability. However, it also brings the defect of a large number of parameters and a large amount of computation. NAS-FPN is an FPN structure obtained by neural network architecture search. BiFPN, the full English name is: Bi-directional Feature Pyramid Network, that is, a bidirectional feature pyramid structure. It is a variant of the FPN structure and has better feature extraction ability compared to models such as FPN, PAN, and NAS-FPN.
[0059] In one embodiment, the target features are visualized.
[0060] In one embodiment, Grad-Cam is used to visualize the target features.
[0061] In one embodiment, multiple different layers of the multi-layer network structure (corresponding to L3, L4, L10, and L12 of RGB-I3D in the following embodiments) adopt improved Inception modules; the convolutional kernel dimensions of the improved Inception modules are 1×1×1, 3×3×3, 3×3×3, and 5×5×5.
[0062] In one embodiment, the 5×5×5 convolutional kernel in the improved Inception module is decomposed into two 3×3×3 convolutional kernels to reduce network calculation parameters and improve network calculation efficiency.
[0063] The smoke detection model of the present invention will be further explained and illustrated through specific embodiments below.
[0064] As Figure 4 shown, in one embodiment, the present invention provides a smoke detection model; specifically, the smoke detection model uses an RGB-I3D with an improved Inception module (Inception-v1) as the backbone network, and by improving the existing RGB-I3D network, an SFM (corresponding to the above-mentioned target feature enhancement network 12) and an MFM (corresponding to the above-mentioned multi-scale feature fusion module 13) are introduced.
[0065] It should be noted that SFM, the full English name: Small Object Feature Enhancement Module, is translated into Chinese as: Small Target Feature Enhancement Module; MFM, the full English name: Multi-scale Feature Fusion Module, is translated into Chinese as: Multi-scale Feature Fusion Module.
[0066] It should be noted that I3D uses Inception modules for 3D expansion, which is a conventional technical means in the field, so its working principle will not be elaborated in detail here.
[0067] As Figure 4 shown, the smoke detection model mainly includes three parts: a feature extraction network 41 (corresponding to the above-mentioned feature extraction network 11 and target feature enhancement network 12), a feature enhancement network 42 (corresponding to the above-mentioned multi-scale feature fusion module 13), and a prediction network 43.
[0068] (1) The explanation of the feature extraction network 41 is as follows:
[0069] The RGB-I3D of the backbone network takes RGB frames as input and expands on Inception-v1. To better retain the time information of smoke movement, the convolutional kernels corresponding to the first two max pooling layers (maxpooling) are 1×3×3, and the stride is 1×2×2. The subsequent kernels and strides are normal expansions. Traditional convolutional layers are selected at the front of RGB-I3D, the middle layer uses the Inception modular design, and finally, average pooling is used instead of the fully connected layer, greatly reducing the model parameters.
[0070] It should be noted that the Inception module is a sparse network structure that can generate dense data, which can not only improve the performance of the neural network but also ensure the efficiency of computing resources. In the smoke detection task, since smoke images have different scales, and in traditional neural networks, only one operation, i.e., a single convolution or pooling, is used in one layer, the extracted features are often too monotonous. Therefore, to improve the smoke detection rate, the present invention uses RGB-I3D with Inception-v1 as the backbone network and uses the Inception-v1 module to extract features with different receptive fields, so that the extracted features are richer. At the same time, 1×1×1 convolutional kernels are added after 3×3×3 and maxpooling respectively to reduce the feature dimension, thus obtaining better performance. The specific network structure of the Inception-v1 module is as Figure 5 shown.
[0071] It should be noted that Figure 5 Previous layer in it represents the convolutional input of the previous layer; Filter concatenation represents the final output of the Inception-v1 module.
[0072] Since the location where forest fires occur is usually far from the surveillance video, the smoke imaging is small and the movement is slow. Directly applying the RGB-I3D network for feature extraction will result in a high false alarm rate. Therefore, the present invention introduces SFM to strengthen the feature extraction of small targets. The network architecture of SFM is as Figure 6 shown. SFM uses the RGB-I3D network and motion estimation to extract the foreground region and uses the channel attention mechanism and spatial attention mechanism for enhancement, greatly improving the small target recognition ability.
[0073] Considering the complexity of network end-to-end processing and calculation and better feature representation ability, the present invention introduces SFM in the middle layer of the RGB-I3D network. For a given input video f in ∈R T×C×H×W , first, a series of convolutions of RGB-I3D are performed to obtain multi-scale features f mid ∈RT×C×H×W , where T represents time, C represents the number of channels, H and W represent height and width respectively, and the multi-scale feature f mid and the input video f in The relationship between them is expressed by the following formula:
[0074] f mid = conv(f in , W θ );
[0075] where conv represents the RGB-I3D convolutional block, and W θ represents the learnable parameters of the RGB-I3D convolutional network.
[0076] Then, sum and average all the frames in the intermediate layer (the multi-scale feature f mid ) at all times t within the corresponding time T to simulate the background image f back , and finally subtract the background image from each frame to obtain the foreground region detection image f fore of the moving target, which is specifically expressed by the following formula:
[0077]
[0078] f fore = f mid - f back ;
[0079] Since direct subtraction may cause problems such as difficult model convergence and data overflow, it is necessary to perform convolution and non-linear processing on the detected foreground region to enhance the feature representation ability, which is specifically expressed by the following formula:
[0080]
[0081] where δ(.) represents the ReLu function, represents the learnable parameters of the convolution.
[0082] To enable the convolutional neural network to better focus on the moving regions in the forest surveillance video, the present invention performs SFM processing on the extracted foreground region, so that the subsequent network can better extract the global and local features of the target, which is specifically expressed by the following formula:
[0083]
[0084]
[0085] f out = f mid + M s M C f' fore ;
[0086] Among them, M C ∈R T×C×1×1 , M s ∈R T×1×H×W , f′ fore ∈R T×C×H×W , f out ∈R T×C×H×W , respectively represent the channel attention coefficient, the spatial attention coefficient, the foreground feature map (corresponding to the above-mentioned processed feature map), and the output feature map (corresponding to the above-mentioned target feature).
[0087] It should be noted that sum represents the summation function; max represents the maximum value function; avg represents the average value function.
[0088] (2) The explanation of the feature enhancement network 42 is as follows:
[0089] Since smoke is a non-rigid object, smoke target detection is multi-scale target detection. At the same time, there are many interference items for smoke detection in the forest environment, such as fog, clouds, light, and other objects with the same color as smoke. Therefore, in order to reduce the false alarm rate and improve the accuracy of smoke detection, this embodiment proposes multi-scale fusion based on MFM; specifically, the RGB-I3D different-scale feature modules L3, L4, L10, and L12 are subjected to feature fusion in two paths, from low layer to high layer and from high layer to low layer, so as to enhance the network feature extraction ability.
[0090] Since smoke moves slowly and has an irregular shape, this embodiment improves the L3, L4, L10, and L12 modules as Figure 5 shown, where the time dimension of the convolution kernel is set to 1, 3, and 5 respectively, so as to better extract the slow movement information of smoke. At the same time, the spatial dimension is set to 1×1, 3×3, and 5×5 respectively, so as to obtain the features of different receptive field regions, better judge the irregular characteristics of smoke, and thus better improve the smoke detection rate and reduce the false alarm rate.
[0091] In order to reduce the network calculation parameters and improve the network calculation efficiency, the improved inception module also decomposes the 5×5×5 convolution kernel into two 3×3×3 convolution kernels.
[0092] On the other hand, since the low-scale feature map has better spatial information and poor semantic information, and the high-scale feature map has better semantic information and poor spatial information, cross-feature fusion of different modules can better express the feature space information and semantic information. The multi-scale feature fusion network architecture (i.e., the feature enhancement network 42) proposed in this embodiment is as Figure 4As shown, the network architecture not only includes paths from top to bottom (corresponding to the above, from high-level to low-level) and from bottom to top (corresponding to the above, from low-level to high-level), but also includes the fusion of a single top-level feature map and a single bottom-level feature map. And a weighted fusion method is adopted to more comprehensively improve the feature expression ability. The specific formula is as follows:
[0093]
[0094]
[0095] Among them, w1, w2, w′1, w′2, w′3, w′4 represent network learnable parameters, and ε represents a very small value to prevent the denominator from being 0. represents the network multi-scale features, represents the network input features, represents the network output features.
[0096] It should be noted that Resize is a function specifically used to resize images, and it uses conventional technical means in the field.
[0097] (3) The explanation of the prediction network 43 is as follows:
[0098] By repeating multiple MFM modules, rich semantic information and spatial information can be obtained. Since the feature map output by the last layer of the network has good semantic information, in order to verify the network classification result, in this embodiment, Grad-Cam is used to visualize the last layer feature map output by the network, so as to better judge the specific features of network classification.
[0099] It should be noted that the basic principle of Grad-Cam is: the gradient of the output values of different categories of the network before softmax (classification function) is solved for the last layer feature map of the network respectively, and then the average of the gradient values of different channels is calculated to obtain the contribution degree of each channel to the classification result, and finally weighted average is performed to obtain Grad-Cam, denoted as It can be expressed as the following formula:
[0100]
[0101]
[0102] Among them, y c represents the score predicted for category c by the network before passing through the softmax activation, represents the data at the position of ij in the k-th channel of the feature map f, Z represents the product of the height and width of the feature map f, f represents the feature map output by the last convolution of RGB-I3D, k represents the k-th channel in the feature layer f, c represents the category, fk Represents the data of channel k in the feature map f, Represents for f k The weight.
[0103] Since the obtained Grad-Cam is a grayscale map of a smaller scale, operations such as scale scaling, colorization, and superposition are required to obtain a heatmap of the same size as the original Figure 1 Specifically expressed by the following formula:
[0104]
[0105] Wherein, Represents a heatmap of the same size as the original obtained through operations such as scale scaling, colorization, and superposition, that is, the finally obtained Grad-Cam heatmap; Scale is a scaling function in the R language, used to unify the observed values of different variables to a specific range (usually [-1, 1]); The Color function is a library function in the C language, which can change the foreground color and background color of the text and supports multiple color option pairs, and all of them adopt conventional technical means in the field. Figure 1 The finally obtained Grad-Cam heatmap can better judge the specific features of network classification, so as to conduct an interpretive analysis of smoke classification.
[0106] The finally obtained Grad-Cam heatmap can better judge the specific features of network classification, so as to conduct an interpretive analysis of smoke classification.
[0107] Next, the smoke detection method in the embodiments of the present invention will be described in detail with reference to the accompanying drawings in the embodiments of the present invention.
[0108] As Figure 7 shown, in one embodiment, the present invention provides a smoke detection method implemented based on the above-mentioned smoke detection model, and the smoke detection method includes:
[0109] Step S21, obtaining monitoring data of the target area.
[0110] Step S24, inputting the monitoring data into the smoke detection model, so that the smoke detection model outputs target features, and realizing smoke detection of the target area based on the target features.
[0111] As Figure 8 shown, in one embodiment, before the step of inputting the monitoring data into the smoke detection model, the smoke detection method further includes:
[0112] Step S22, obtaining a smoke monitoring data set.
[0113] Step S23, training the smoke detection model based on the smoke monitoring data set to obtain a trained smoke detection model.
[0114] It should be noted that the smoke monitoring dataset at least includes a smoke monitoring data group and a fog monitoring data group; among them, the smoke monitoring data group includes multiple smoke monitoring data; the fog monitoring data group includes multiple fog monitoring data.
[0115] In this embodiment, inputting the monitoring data into the smoke detection model includes: inputting the monitoring data into the trained smoke detection model.
[0116] As Figure 8 shown, in one embodiment, the smoke detection method includes:
[0117] Step S25: Input the monitoring data into the trained smoke detection model, so that the trained smoke detection model outputs target features, and based on the target features, perform smoke detection on the target area.
[0118] It should be noted that this step S25 corresponds to step S24 in the above embodiment.
[0119] The following further explains the training of the smoke detection model in step S23 through specific embodiments.
[0120] In one embodiment, build an experimental platform; specifically, this experimental platform is a personal desktop computer, the experimental environment is the ubuntu system, the processor is an AMD Ryzen 9 5900X 12-Core Processor, the graphics card is an NVIDIA GeForce RTX 3090 GPU, and the PyTorch framework is used.
[0121] During the training process, the initial learning rate is 0.1, the learning rate change boundaries (Milestones) are (500, 1500), and the decay weight is 10 -6 , the video shape input to the network is [40, 3, 36, 224, 224], corresponding to Batch-Size, the number of input channels, the number of frames, and the height and width of the image respectively.
[0122] In this embodiment, SAN-SD, EFFNet, VSSNet, M-3DFCN, and RGB-I3D are selected for comparison of the smoke recognition and detection effects.
[0123] It should be noted that SAN-SD, EFFNet, VSSNet, M-3DFCN, and RGB-I3D all adopt conventional technical means in the field, so their working principles will not be elaborated in detail here.
[0124] I. Smoke dataset
[0125] In this embodiment, a real forest fire smoke video dataset is built (since relevant resources cannot be found elsewhere), named the Forest Smoke dataset. This dataset is collected from multiple cameras in forest areas. Most of the smoke images in this dataset are at a relatively long distance, and there are many interfering objects similar to smoke, which can prove the effectiveness and feasibility of the smoke detection algorithm of the present invention. This dataset includes 2 categories and 13,140 video clips, among which there are 6,570 smoke class videos and 6,570 fog class videos. The collected videos belong to the RGB color space, with pixels of 1920×1080. In this embodiment, 9,198 videos are used for training, 1,314 videos are used for verification, and 2,628 videos are used to test the experimental effect.
[0126] By manually annotating the video data in the dataset, the smoke characteristics generated by forest fires in different scenarios can be reflected. This dataset is of great significance for studying the recognition of forest fire smoke in real scenarios.
[0127] Due to different acquisition environments, the intensity of light and the performance of the equipment are different, resulting in the lack of contrast and noise in the collected data. Therefore, during the training process, standard data augmentation is applied, including operations such as horizontal flipping, random resizing and cropping, perspective transformation, area erasing, and color jitter. To ensure the consistency of variables, the same image preprocessing is performed in all comparative experiments. The complex background makes this dataset more challenging, so this dataset has specific research significance and application value.
[0128] II. Evaluation Criteria
[0129] The evaluation metrics used in this experiment are precision P, recall R, the harmonic mean F1 of precision and recall, false positive rate FPR, and accuracy ACC, which are specifically expressed as the following formulas:
[0130]
[0131]
[0132]
[0133]
[0134]
[0135] Among them, TP represents the number of positive classes predicted as positive classes, FN represents the number of positive classes predicted as negative classes, FP represents the number of negative classes predicted as positive classes, and TN represents the number of negative classes predicted as negative classes.
[0136] III. Comparative Experiments
[0137] To better evaluate the effectiveness of the smoke detection algorithm of the present invention, comprehensive experiments were conducted on the forest smoke dataset. To better understand the performance of the smoke detection method of the present invention, it was compared with classical and state-of-the-art video classification methods, including SAN-SD, EFFNet, VSSNet, M-3DFCN, RGB-I3D. To ensure the fairness of the evaluation results, all experiments were implemented on the same platform. Except for the smoke detection algorithm of the present invention, all other algorithms were fine-tuned on the original pre-trained models. Table 1 shows the video classification results of these different methods.
[0138] Table 1 Comparison Experiment Results
[0139]
[0140] Among them, FLOPs represents the number of floating-point operations per second, which is understood as the computing speed; the F1 score, also known as the F1-Score, is a metric used to evaluate the performance of binary classification models.
[0141] As shown in Table 1, RGB-I3D performs better than these benchmark methods on the test set. The combination of SFM and MFM greatly improves the smoke recognition ability of RGB-I3D. The advantages of the algorithm proposed in this paper on forest smoke confirm that SFM and MFM improve the smoke recognition performance. In addition, SFM and MFM do not impose an additional large computational load. The SFM and MFM proposed in this paper overcome the problems that small targets are difficult to detect, the detection effect is unstable, the image resolution is low, and there are a large number of objects similar in color to smoke interfering in the forest environment.
[0142] Although SFM can enhance small target regions, it is impossible to identify the target region for slowly moving smoke through the motion detection algorithm. In addition, although MFM can improve the accuracy of smoke detection and exclude the interference of similar objects, it is a challenging task to extract the features of thin smoke without a clear contour. Even humans need to observe carefully to distinguish these subtle differences.
[0143] IV. Ablation Experiments
[0144] Through ablation studies, the contribution of each module of the algorithm proposed in this paper was evaluated. The ablation analysis of the algorithm proposed in this paper on ForestSmoke is shown in Table 2. In the ablation experiment analysis, 7 model variants were tested to explore the influence of different sub-models and tasks.
[0145] Table 2 Ablation Experiment Results
[0146]
[0147] Among them, RGB-I3D is a general smoke video classification model, which is considered as the basic network framework in these ablation experiments. RGB-I3D+SFM-front adds SFM to the first half of RGB-I3D to predict the smoke classification result. RGB-I3D+SFM-mid adds SFM to the middle part of RGB-I3D to predict the smoke classification result. RGB-I3D+SFM-back adds SFM to the second half of RGB-I3D to predict the smoke classification result. RGB-I3D+MFM-3 adds MFM to L10, L12, and L4 based on RGB-I3D to perform enhanced fusion of features at different scales. RGB-I3D+MFM-4 adds MFM to L10, L12, L4, and L3 based on RGB-I3D to perform enhanced fusion of features at different scales. RGB-I3D+MFM-5 adds MFM to L10, L12, L4, L3, and L1 based on RGB-I3D to perform enhanced fusion of features at different scales. RGB-I3D+SFM-mid+MFM-3 adds MFM to L10, L12, and L4 based on RGB-I3D+SFM-mid to perform enhanced fusion of features at different scales. RGB-I3D+SFM-mid+MFM-4 adds MFM to L10, L12, L4, and L3 based on RGB-I3D+SFM-mid to perform enhanced fusion of features at different scales. RGB-I3D+SFM-mid+MFM-5 adds MFM to L10, L12, L4, L3, and L1 based on RGB-I3D+SFM-mid to perform enhanced fusion of features at different scales.
[0148] As shown in Table 2, the experimental results show that adding SFM-mid to RGB-I3D leads to a 0.03 increase in accuracy and no change in FPR, indicating that the addition of SFM-mid enhances the small target recognition ability and improves the accuracy of smoke recognition. Under the same other conditions, adding MFM-4 to RGB-I3D leads to a 0.02 increase in accuracy and a 0.03 decrease in the FPR metric, indicating that the addition of MFM-4 enhances the fusion of features at different scales and reduces the false alarm rate of objects with colors similar to smoke. Finally, adding both SFM-mid and MFM-4 to RGB-I3D results in a 0.07 increase in accuracy and a 0.03 decrease in FPR, indicating that the fusion of SFM and MFM can enhance the small target recognition ability and reduce the interference of similar objects, obtaining a better recognition effect.
[0149] V. Detection Results
[0150] Visualize the prediction results of the smoke detection algorithm of the present invention. Select a set of videos of the early stage of common forest fires in a real scenario and transmit the videos to all comparison models for classification and recognition, so as to accurately identify the smoke in various complex forest environments.
[0151] It should be noted that in order to solve the deficiencies in the deep learning algorithm, the present invention proposes a forest fire smoke recognition based on a small target feature enhancement network. The network framework takes RGB-I3D as the basic network framework and makes major modifications to the network according to the complex smoke detection scenario in the forest. The key ideas of this framework are as follows: 1) Add a spatial attention mechanism and a channel attention mechanism to the motion estimation algorithm to enhance the network's attention to small target features; 2) At the same time, use parallel multi-branch and serial skip-layer connections to make the multi-scale features extracted by the convolutional neural network diverse; 3) In order to improve the network detection accuracy, use Grad-Cam to visualize the network detection target; specifically,
[0152] (1) Propose a small target feature enhancement module of SFM, so that the neural network can better focus on the smoke movement area, enhance the network's ability to recognize small targets, and improve the smoke detection accuracy.
[0153] (2) Propose a multi-scale feature fusion module of MFM, and use the advantages of parallel multi-branch networks and serial skip-layer connection networks at the same time to make the convolutional neural network extract richer multi-scale features, fully fuse the semantic information of the high-level feature map and the spatial information of the low-level feature map, and reduce the false alarm rate caused by objects with colors similar to smoke.
[0154] (3) Use Grad-Cam to visualize the output of the target detection, so that researchers can perform interpretable analysis on the network recognition and improve the smoke recognition accuracy.
[0155] (4) Build a real forest fire smoke video dataset by oneself, which can reflect the smoke characteristics generated by forest fires in different scenarios and is of great significance for studying the recognition of forest fire smoke in real scenarios.
[0156] There are challenges in forest fire smoke recognition, such as small imaging targets being difficult to detect and a high false alarm rate caused by many interfering objects. In response to these challenges, the present invention proposes a video smoke recognition algorithm based on a small target feature enhancement network, whose main purpose is to quickly and accurately determine whether there is smoke in a forest surveillance video and the specific location of the smoke area. The algorithm proposed in the present invention introduces two modules, SFM and MFM, into the RGB-I3D algorithm. Among them, the SFM module adds an attention mechanism in background subtraction to improve the convolutional neural network's attention to small moving targets and enhance the recognition ability of small target smoke; the MFM module simultaneously adopts parallel multi-branch and serial skip-layer connections to make the convolutional neural network extract diverse multi-scale features, improve the network's discrimination of fog, and reduce the false alarm rate. The algorithm proposed in the present invention is compared with five other state-of-the-art methods on a self-built real dataset. The experimental results show that the proposed method is effective and feasible, and the smoke recognition accuracy reaches 97%, and the false alarm rate reaches 5%.
[0157] The protection scope of the smoke detection method described in the embodiments of the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.
[0158] The embodiments of the present invention also provide an electronic device, which includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the above-mentioned smoke detection method.
[0159] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by an electronic device, it implements the above-mentioned smoke detection method.
[0160] Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0161] An embodiment of the present invention further provides a smoke detection system. The smoke detection system can implement the smoke detection method of the present invention. However, the implementation devices of the smoke detection method of the present invention include, but are not limited to, the structures of the smoke detection systems listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.
[0162] As Figure 9 shown, this embodiment provides a smoke detection system implemented based on the above smoke detection model. The smoke detection system includes:
[0163] An acquisition module 91, configured to acquire monitoring data of a target area.
[0164] A detection module 92, configured to input the monitoring data into the smoke detection model, so that the smoke detection model outputs target features, and based on the target features, implement smoke detection of the target area.
[0165] It should be noted that the structures and principles of the acquisition module 91 and the detection module 92 correspond one by one to the steps (step S21 and step S24) in the above smoke detection method. The specific working principles can also refer to the introduction of the smoke detection method in the foregoing embodiments, and thus will not be elaborated here.
[0166] In several embodiments provided by the present invention, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical or other forms.
[0167] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present invention. For example, in each embodiment of the present invention, the functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.
[0168] Those of ordinary skill in the art should further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0169] The descriptions of the corresponding processes or structures in the above-mentioned various drawings have their own emphases. For the parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.
[0170] The above embodiments are only illustrative of the principles and effects of the present invention and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A smoke detection model is applied to the smoke detection of a target area, and is characterized in that, The smoke detection model is a multi-layer network structure, and the smoke detection model includes: a feature extraction network and a target feature enhancement network; where The feature extraction network is used to receive the monitoring data of the target area, and is used to extract features from the monitoring data to obtain multi-scale features; The target feature enhancement network is connected to the feature extraction network. The target feature enhancement network is used to receive the multi-scale features, and is used to obtain target features based on the multi-scale features, so as to realize smoke detection of the target area based on the target features; the target feature enhancement network is used to obtain target features based on the multi-scale features, including: The target feature enhancement network performs summation and averaging processing on all frames of the multi-scale features to obtain a background map; The target feature enhancement network subtracts the background map from each frame of the multi-scale features to obtain a foreground area detection map; The target feature enhancement network enhances the features of the foreground area detection map to obtain the target features; the target feature enhancement network enhances the features of the foreground area detection map to obtain the target features, including: The target feature enhancement network performs convolution and non-linear processing on the foreground area detection map to obtain a processed feature map; The target feature enhancement network adds an attention mechanism to the processed feature map to obtain the target features.
2. The smoke detection model according to claim 1, wherein Multiple multi-scale features corresponding to different layers of the multi-layer network structure are obtained through the feature extraction network, so as to obtain multiple target features based on the multiple multi-scale features through the target feature enhancement network; The smoke detection model further includes: a multi-scale feature fusion module; The multi-scale feature fusion module is connected to the target feature enhancement network. The multi-scale feature fusion module is used to perform feature fusion on multiple target features along two paths, from low layer to high layer and from high layer to low layer, to obtain output features, so as to realize smoke detection of the target area based on the output features.
3. The smoke detection model according to claim 1, wherein Visualize the target features; and / or Multiple different layers of the multi-layer network structure adopt improved inception modules; the convolutional kernel dimensions of the improved inception modules are 1×1×1, 3×3×3, 3×3×3, 5×5×5.
4. A smoke detection method implemented based on the smoke detection model described in any one of claims 1 to 3, characterized in that, The smoke detection method includes: Obtain the monitoring data of the target area; Input the monitoring data into the smoke detection model, so that the smoke detection model outputs target features, so as to realize smoke detection of the target area based on the target features.
5. The smoke detection method according to claim 4, characterized in that Before the step of inputting the monitoring data into the smoke detection model, the smoke detection method further includes: Obtain a smoke monitoring data set; Train the smoke detection model based on the smoke monitoring data set to obtain a trained smoke detection model; the smoke monitoring data set includes a plurality of smoke type monitoring data and a plurality of fog type monitoring data; Inputting the monitoring data into the smoke detection model includes: inputting the monitoring data into the trained smoke detection model.
6. A smoke detection system implemented based on the smoke detection model according to any one of claims 1 to 3, characterized in that, The smoke detection system includes: An acquisition module, configured to acquire monitoring data of a target area; A detection module, configured to input the monitoring data into the smoke detection model, so that the smoke detection model outputs target features, and based on the target features, perform smoke detection on the target area.
7. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, so that the electronic device executes the smoke detection method described in claim 4 or 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the electronic device, it implements the smoke detection method described in claim 4 or 5.
Citation Information
Patent Citations
Smoke image segmentation and recognition method based on dictionary and BP neural network
CN110415260A
Outdoor fire smoke image detection method based on Recursive BIFPN network
CN115690564A
Detection model training and fire detection method and system based on global attention
CN116824346A
Cited By
A large commercial complex fire hazard cross-region cooperative supervision method and system
CN122530941A