A smoke recognition method for fire video based on dual-channel temporal and spatial domains

By using a 3D convolutional neural network with dual-channel time and space domains and residual attention blocks in smoke recognition, the static and dynamic characteristics of smoke are extracted and adaptively fusion is carried out, and the problems of low accuracy and poor robustness of smoke recognition in the prior art are solved, achieving the recognition effect of high accuracy and low false alarm rate.

CN114580541BActive Publication Date: 2025-05-23ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210215812.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-05-23
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The existing smoke identification methods have low accuracy and poor robustness, and cannot effectively deal with the variable characteristics of smoke.

Method used

The improved 3D convolutional neural network and residual attention block (RAB) based on space-time and space domain dual channels are used to extract the static and dynamic features of smoke respectively, and fuse them through adaptive feature fusion method for real-time identification and early warning of smoke.

Benefits of technology

The high accuracy and low false alarm rate of smoke are recognized, with an accuracy rate of 98.87% and a false alarm rate of 1.06%, meeting the real-time and accuracy requirements of smoke recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580541B_ABST
    Figure CN114580541B_ABST
Patent Text Reader

Abstract

The present invention proposes a smoke recognition method based on dual-channel fire video in time and space domain, and the steps are: collect and make a smoke data set containing cloud and fog interference; build a static feature extraction network and a dynamic feature extraction network, fuse and connect the static feature extraction network and the dynamic feature extraction network, and build a video smoke recognition network model; use the smoke data set to train the video smoke recognition network model to obtain an optimized video smoke recognition network model; use the optimized network model to process the smoke video collected in real time, the static feature extraction network extracts the static features of the image in the spatial domain, the dynamic feature extraction network extracts the dynamic features of the video in the time domain, the static features and the dynamic features are fused to generate smoke features, and the smoke features are identified to determine whether smoke exists. The present invention has higher accuracy and recall rate, lower false alarm rate, and can effectively identify and warn smoke in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smoke recognition, and in particular to a smoke recognition method based on dual-channel fire videos in time and space domains. Background Art

[0002] Existing smoke detection devices based on traditional physical sensors are widely used due to their low cost and simple installation. Simply installing it in a fixed position can basically meet the standard requirements for smoke detection. However, this type of detection device also has certain limitations: on the one hand, they must be installed near the fire source, and the triggering of the alarm requires sufficient smoke concentration or air temperature, which will greatly affect the real-time performance of smoke detection. On the other hand, since the detector is in direct contact with dust and smoke all year round, the harsh environment causes the sensor to easily malfunction and lose the detection effect. Therefore, physical sensor-based smoke detection methods are difficult to meet the requirements of today's industrial processes and smoke safety warnings, and a higher performance and higher accuracy smoke detection method is urgently needed.

[0003] The smoke detection algorithm based on traditional image processing overcomes the shortcomings of the physical sensor method. First, the smoke video in the scene is obtained using a video surveillance device; then the color, texture, shape and other features of the smoke in the video image are extracted through an artificially designed algorithm; finally, the smoke classifier is trained to identify the smoke and issue an early warning. Appana et al. converted the smoke image from the RGB color space to the HSV color space, and thresholded its saturation and brightness to achieve the segmentation of the smoke area. Peng et al. achieved smoke detection by extracting the minimum enclosing rectangle of the moving area in the image as the shape feature of the smoke. However, this type of method requires a lot of time and effort on feature extraction and feature selection, which increases the complexity of the algorithm. At the same time, due to the changeable characteristics of smoke, this type of detection algorithm also has a high false alarm rate.

[0004] Smoke detection in the early stages of a fire plays an important role in fire early warning. In recent years, video smoke recognition algorithms based on convolutional neural networks (CNNs) have been widely proposed and applied in various fields of industry. However, due to the changeable characteristics of smoke, such algorithms still face problems such as low accuracy, poor robustness, and insufficient data sets.

[0005] Compared with the manually designed feature extraction algorithm, the convolutional neural network can automatically learn the basic information and more complex pixel information in the smoke image, which not only reduces the complexity of the feature extraction process, but also overcomes the changeable characteristics of smoke. Gu et al. used a deep dual-channel neural network (DCNN) to fuse the basic information and detailed information of smoke for smoke classification, with a good smoke recognition rate. Zhang et al. used transfer learning to extract the generalized features of smoke, which accelerated the convergence speed and recognition speed during network training. Cao et al. proposed a feature foreground model to construct an enhanced feature foreground network for smoke source prediction and detection, which improved the network's learning ability for smoke. He et al. combined the attention mechanism, feature level and decision level fusion modules to achieve the recognition ability of small smoke targets. However, it is difficult to make further breakthroughs in accuracy and false alarm rate by only using the static features of smoke. Studies have found that the diffusion characteristics of smoke are the key features in smoke recognition, which can significantly improve the accuracy of the model. Therefore, how to effectively extract the motion features in smoke has become the main problem faced by smoke recognition. The existing motion feature extraction methods are mainly divided into two categories: one is based on traditional image processing methods, and the other is the method using 3D CNN (convolutional neural network). Traditional image processing methods use frame difference method, background difference method, optical flow method or ViBe (Visual Background Extractor) algorithm to calculate the binary image of the moving object as the motion feature. Liu et al. introduced the concept of visual change image to describe the diffusion characteristics of smoke, and proposed a two-stage smoke detection algorithm combining deep normalization network and support vector machine (SVM), which reduced the false alarm rate caused by objects such as clouds and fog, but the visual change image still needs to be calculated by manually designed complex algorithms. Gao et al. used the MSER (Maximally Stable Extrernal Regions) algorithm to make up for the blank area problem generated by the ViBe algorithm in short-distance conditions, and generated a more complete candidate smoke region representing the motion characteristics of smoke, which improved the accuracy of smoke recognition and reduced the missed detection rate, but the algorithm cost increased by 20 times under the same hardware conditions. Sheng et al. extracted all statistical images in the time domain, frequency domain and time-frequency domain of the image as the static and motion features of smoke, and input them into the deep belief network for smoke recognition, which improved the speed and accuracy of smoke detection in complex environments. The above algorithm processes are complex and easily affected by image quality and complex environment. For example, when there are non-smoke moving objects in the scene, the extracted motion features often contain interference from other moving objects, reducing the accuracy of recognition. The 3DCNN-based method can automatically learn the motion features in video images without using prior knowledge to design algorithms to screen the interference features of motion. It has the advantages of simple algorithm, easy implementation, and little interference from the environment.Therefore, 3D CNN has an inherent advantage in smoke recognition tasks with diffusion characteristics. However, there are very few literatures using 3D CNN in existing methods, and the accuracy of these methods needs to be improved.

[0006] Even though many scholars have continuously optimized the network model to improve the accuracy of the smoke recognition algorithm. However, the smoke recognition algorithm based on CNN still has many shortcomings: on the one hand, a large number of existing algorithms are simply the simple application of CNN in smoke image recognition, without considering the motion characteristics of smoke. Even though a small number of algorithms take into account the diffusion characteristics of smoke, they still use traditional image processing methods to extract motion features in video smoke. The algorithm is complex and loses a lot of high-frequency detail information in the smoke image. On the other hand, most of the existing methods are based on traditional deep learning models, and the network model is close to being outdated, and it is impossible to further improve the accuracy of smoke detection. The more novel network model with better detection effect has not been widely used. At the same time, the single scene of the training data set also leads to poor generalization ability and low robustness of the smoke recognition model. Summary of the invention

[0007] In view of the technical problems that the existing smoke recognition methods have low accuracy and poor robustness and cannot meet the changing characteristics of smoke, the present invention proposes a smoke recognition method based on dual-channel fire videos in the spatiotemporal domain. The dynamic and static features of smoke are extracted respectively based on an improved 3D convolutional neural network and residual attention block (RAB), and then fused, so that smoke can be recognized and warned in real time and effectively.

[0008] In order to achieve the above object, the technical solution of the present invention is implemented as follows: a method for smoke recognition based on dual-channel fire video in time and space domain, the steps of which are as follows:

[0009] Step 1: Collect and create a smoke dataset containing cloud and fog interference images and videos;

[0010] Step 2: Build a static feature extraction network and a dynamic feature extraction network, fuse and connect the static feature extraction network and the dynamic feature extraction network, and build a video smoke recognition network model;

[0011] Step 3: Use the smoke data set in step 1 to train the video smoke recognition network model constructed in step 2 to obtain an optimized video smoke recognition network model;

[0012] Step 4: Use the optimized video smoke recognition network model to process the real-time collected smoke video. The static feature extraction network extracts the static features in the image space domain, and the dynamic feature extraction network extracts the dynamic features of the video sequence in the time domain. The static features and dynamic features are fused to generate smoke features, and the smoke features are identified to determine whether there is smoke. If there is smoke, an alarm is issued.

[0013] The smoke dataset in step 1 includes smoke images or videos in forest, field, indoor, playground, construction site, city, and highway scenes. The smoke dataset includes images and videos of multiple positive samples and negative samples.

[0014] The video smoke recognition network model includes a static feature extraction network and a dynamic feature extraction network connected in parallel, both of which are connected to a fusion component, and the fusion component is connected to a fully connected unit; the fusion component adopts an adaptive fusion method and utilizes the learning ability of a neural network to reallocate weights for fusion features.

[0015] The fusion component includes a feature fusion unit, which combines the static features extracted by the static feature extraction network and the dynamic features extracted by the dynamic feature extraction network: a set of 1×(n+k) feature vectors I is obtained by using 3D global average pooling, and the feature vectors are transformed into The feature matrix I is processed by convolution to obtain the weight matrix II. The feature matrix II is converted into a 1×(n+k) weight vector II; the weight vector II is multiplied by the corresponding static features and dynamic features to obtain the fused features, which are input into the fully connected layer of the fully connected unit.

[0016] The fusion method of the video smoke recognition network model is:

[0017] Among them, the parameters and After back propagation and autonomous learning, F st and F dy Respectively represent the static features and dynamic features used for fusion, The fused features.

[0018] The static feature extraction network is built based on the residual attention module; the static feature extraction network connects 12 residual attention modules in sequence, and a pooling layer is connected after every two residual attention modules; the residual attention blocks shown all use a 3×3 convolution kernel with a step size of 1; the pooling layers all use a maximum pooling of 2×2 with a step size of 2; the static feature extraction network uses a ReLU nonlinear non-saturated activation function.

[0019] The residual attention block includes a channel attention unit, a spatial attention unit and a residual structure. The input feature map X is transmitted to the channel attention unit through the convolution operation in the trunk branch, and then the feature map obtained by the channel attention unit is processed. Transmitted to the spatial attention unit, the spatial attention unit processes the feature map Feature Map is the feature map after redistribution of weights, feature map Add the input feature map X in the jump branch to get the output feature map

[0020] The dynamic feature extraction network is a weighted 3D convolutional neural network, and the dynamic feature extraction network includes a feature extraction module and an attention module, the feature extraction module includes at least two feature extraction units connected in sequence, each feature extraction unit includes a 3D convolution layer and a pooling layer connected in sequence; the feature map processed by the feature extraction module is subjected to convolution operation to obtain feature maps F = {f 1 ,f 2 ,...,f k} and attention map A = {a 1 ,a 2 ,...,a k}, attention map A = {a 1 ,a 2 ,...,a k} k As weights and feature maps F = {f 1 ,f 2 ,...,f k The feature f in k Multiply them sequentially to get the dynamic characteristics and

[0021]

[0022] The feature extraction module of the dynamic feature extraction network includes 5 feature extraction units. The 3D convolution layers in the feature extraction units all use 3D convolution kernels of size 3×3×3, stride of 1×1×1, and padding attribute padding=1; pooling layer II is 2×2×2 3D maximum pooling.

[0023] Compared with the prior art, the present invention has the following beneficial effects: using RAB to build a static feature extraction network to enhance the static features of smoke in the spatial domain of the image; and in view of the diffusion characteristics of smoke, a weighted 3D convolutional neural network (W-3DCNN) is proposed to extract the dynamic features of smoke in the time domain; at the same time, an adaptive feature fusion method is used to fuse the static and dynamic features in the smoke for the final smoke recognition. In the experimental stage, the proposed recognition model was evaluated on a custom data set. The evaluation results show that the accuracy and false alarm rate of the present invention are 98.87% and 1.06% respectively, which provides a practical and more accurate method for smoke recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0025] Figure 1 It is a schematic diagram of the process of the present invention.

[0026] Figure 2 It is a structural schematic diagram of the static feature extraction network of the present invention.

[0027] Figure 3 For the present invention Figure 2 Schematic diagram of the residual attention block in [5].

[0028] Figure 4 It is a structural schematic diagram of the dynamic feature extraction network of the present invention.

[0029] Figure 5 It is a structural schematic diagram of the video smoke recognition network model of the present invention.

[0030] Figure 6 These are some samples in the experiment of the present invention, where (a) is a missed detection sample and (b) is a false detection sample. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] like Figure 1 As shown, a smoke recognition method based on dual-channel fire video in time and space domain, the steps are as follows:

[0033] Step 1: Collect and create a smoke dataset containing cloud and fog interference images.

[0034] In order to enhance the robustness of the constructed smoke recognition network model, smoke datasets in multiple scenes including forests, fields, indoors, playgrounds, construction sites, cities, and highways were collected and produced for training and testing the network model. The images or videos in the smoke dataset are interfered by clouds and fog. The smoke dataset contains a variety of positive and negative sample image data and video data, which solves the problem of insufficient smoke datasets. Positive samples are images with smoke, and negative samples are images without smoke.

[0035] Step 2: Build a static feature extraction network and a dynamic feature extraction network, fuse and connect the static feature extraction network and the dynamic feature extraction network, and build a video smoke recognition network model.

[0036] The video smoke recognition network model is a dual-channel network based on static and dynamic features of smoke, which can detect smoke images under smoke-like interference. The static feature extraction network is based on residual attention blocks, which improves the learning ability of the network and solves the problem that the high-frequency information and low-frequency information in the smoke image are treated equally, resulting in the network being unable to fully exert its performance.

[0037] The static feature extraction network is used to extract deeper static features of smoke and improve the accuracy of smoke recognition. The color, texture, edge and other features in the smoke image are the main features for effective smoke recognition. In order to extract more significant static smoke features, the present invention builds a static feature extraction network based on RAB, such as Figure 2 As shown in the figure, the static feature extraction network is connected to 12 residual attention modules in sequence, and every two residual attention modules are connected to 1 maximum pooling. The last residual attention module is connected to the fully connected layers fc1 and fc2 in sequence.

[0038] The static feature extraction network is composed of a stack of RAB layers and pooling layers. The fully connected layer at the end of the network summarizes the smoke features and uses them for classification. On the one hand, the RAB layer uses channel attention to find more powerful channel expressions in the smoke feature map, and then enhances the expression ability of these channels; on the other hand, spatial attention can enhance more effective information in the smoke features and suppress invalid information, allowing the network to focus more on the smoke object. The introduction of the pooling layer can effectively reduce the size of the smoke feature map, play a role in dimensionality reduction and removal of redundant information. The original feature map is redistributed by weight, and the feature map after the redistribution of weights is the feature map with attention. This process adds spatial attention. And when the image undergoes a series of rigid transformations, the pooling layer ensures that the network can still extract effective features. When the input image of the overall model changes, the model can still recognize smoke. For example: when the input is an ordinary smoke image, the model can recognize the smoke in the image, but if the image is rigidly transformed, that is, flipped, rotated or scaled, the model can still recognize it. That is, a simple transformation of the model input will not affect the model results.

[0039] The input of the network is an RGB image of 256×256 pixels. RAB uses a convolution kernel of 3×3 size. When stride=1, the RAB layer does not change the size of the input image. The pooling layer uses a global maximum pooling of size 2×2. Each time the network passes through a pooling layer with a stride of 2, the image size will be reduced to half of the original size, so the number of convolution kernels of RAB connected after the pooling layer with a stride of 2 will increase to twice the original size. The size of the feature map after pooling will become smaller, which means that the information in the feature map will be lost. In order to compensate for the information loss caused by the smaller feature map, the number of convolution kernels of RAB connected later will increase. The increase in the number of convolution kernels will increase the number of feature maps, and the increase in the number of feature maps will compensate for the loss caused by the reduction in the size of the feature map. When the neural network is back-propagating, the gradient will decay each time it is passed through a layer until it almost disappears, making the training network converge more and more slowly. In order to overcome the vanishing gradient and speed up the convergence of the network, the static feature extraction network uses the ReLU nonlinear non-saturated activation function.

[0040] For high-resolution images, the depth of CNN plays a crucial role, but networks that are too deep are difficult to train. Low-resolution images and features contain rich low-frequency information, but they are treated equally between channels, which hinders the network's representation ability. The Residual Attention Block (RAB) combines channel attention, spatial attention, and residual structure to find a more powerful static representation capability in the smoke feature map, enhance more effective information in the smoke features, and suppress invalid information, further improving the accuracy of smoke recognition.

[0041] The residual attention block (RAB) is used to extract the deeper pixel features of smoke. RAB is constructed by stacking multiple layers of attention modules and multiple skip connections, such as Figure 3 As shown in the figure, channel attention represents the spatial attention layer, and spatial attention represents the channel attention layer, which allows the model to pay more attention to the smoke image in the image and ignore some unimportant areas. conv represents the convolution operation, and the input feature X is transmitted to the spatial attention layer through convolution to obtain the feature feature Get features through channel attention layer By inputting feature X through residual structure and feature Add the output feature F st(X). Each RAB is divided into two parts: a mask branch and a trunk branch. The trunk branch focuses on learning high-frequency information in the image and extracts feature maps after feature refinement through the continuous arrangement of attention modules. The mask branch is the main method to implement the residual block, which speeds up network training while allowing rich low-frequency information to be directly propagated through multiple skip connections.

[0042] The residual attention block is calculated as:

[0043]

[0044] in, Represents the feature map after the weights are redistributed. The method of rewriting the weight distribution is to multiply the feature map and the weight matrix. X represents the output feature of the previous layer of the network, which is passed to the next layer through the trunk branch and the jump branch. F st (X) represents the output features of the residual attention block.

[0045] On the one hand, the color, texture, edge and other features of smoke have poor anti-interference ability and are easily confused with clouds, shadows, fog and other types of smoke, which reduces the recognition accuracy; on the other hand, the shape and color of smoke change continuously from generation to diffusion, and ordinary neural networks cannot perceive such diverse changes in time sequence. The input of 3D neural convolutional network (CNN) is a continuous video sequence. By sliding and performing convolution operations on the three dimensions of video frame width, frame height and frame number, it can effectively extract relevant information of moving objects in the video, so that they can be flexibly used in various complex tasks. By analyzing the false detection and missed detection samples, a smoke dynamic feature extraction network based on 3D CNN is proposed. In this network, the motion characteristics of smoke are used to improve the recognition ability of similar interference targets such as smoke, clouds and fog, and solve the problem of high complexity and weak anti-interference ability of motion feature extraction in traditional image processing algorithms.

[0046] In order to further improve the accuracy and robustness of smoke detection, a 3D neural convolutional network is used to extract the motion features of smoke diffusion, so as to distinguish it from the smoke-like features; at the same time, an attention mechanism is added to the network to learn the corresponding feature weights for the smoke features at different periods, in order to increase the generalization ability of smoke recognition, identify the smoke features at different periods, and detect fire as early as possible. The network structure of the dynamic feature extraction network is shown in Figure 2. Figure 4As shown in the figure. The dynamic feature extraction network includes a feature extraction module and an attention module. The feature extraction module includes 5 feature extraction units connected in sequence, and each feature extraction unit includes a 3D convolution layer and a pooling layer. The feature map processed by the feature extraction module is convolved to obtain a feature map and input to the attention module. The attention module processes the convolved feature map into an attention map through a convolution operation. The features in the attention map and the feature map are multiplied in sequence to obtain a feature map, which is finally processed through the fully connected layer II to generate dynamic features.

[0047] The dynamic feature extraction network consists of a feature extraction module and an attention module. The network end uses a fully connected network to summarize the dynamic features of smoke for classification. A 25×3×256×256 video sequence is used as the input of the network. The network uses a 3D convolution kernel of size 3×3×3 in the feature extraction module, with a stride of 1×1×1 and a padding attribute of 1. The pooling layer is a 2×2×2 3D maximum pooling. Similarly, when the pooling layer reduces the input image size to half of the original size, the number of convolution kernels increases to twice the original size.

[0048] The dynamic feature extraction network extracts feature maps F = {f 1 ,f 2 ,...,f k} and attention map A = {a 1 ,a 2 ,...,a k}, and then one by one feature map f k And the corresponding attention map a k Multiply to generate the final feature map As shown in formula (2). The purpose is to expect each weight Figure 1 On the one hand, it can represent the different importance of each channel of the 3D feature map, and on the other hand, it can enhance the target area of ​​interest in each 3D feature map.

[0049]

[0050] The dynamic feature extraction network adds 3D attention connection to the 3D convolution to obtain the weighted 3D convolutional neural network (W-3D CNN), which is used to more fully and efficiently extract the smoke motion information between adjacent frames in the video image. Using W-3DCNN to extract the smoke motion information between adjacent frames in the video image can more effectively eliminate the interference of smoke-like factors such as clouds, fog, and shadows, and reduce the false alarm rate.

[0051] Commonly used feature fusion methods include concat fusion, add fusion, and max fusion. Add fusion and max fusion are based on the premise that the scale of feature maps is the same, which greatly limits the flexibility of the neural network output feature maps. Concat fusion only simply splices the feature maps and ignores the correlation between feature maps. This also leads to unsatisfactory effect of feature fusion in some scenarios. To solve the above problems, the present invention proposes an adaptive fusion method. Utilizing the learning ability of neural networks, weights are reallocated to feature maps, so that the more important feature maps are determined by model learning. This not only reduces manual intervention, but also does not require the fused feature maps to have the same resolution, thereby increasing the flexibility of the network. First, the fully connected layers of the static feature extraction network and the dynamic feature extraction network are removed respectively, and then the extracted features are input into the fusion component to perform feature fusion. The overall structure of the network is as follows: Figure 5 shown.

[0052] The fusion component performs 3D global average pooling on n groups of static feature maps and k groups of dynamic features to obtain a 1×(k+n) real number matrix. In order to reduce the number of parameters in the network, a convolutional network is used to obtain the corresponding weight matrix. Reshape means converting the 1×(k+n) real number matrix into The characteristic matrix of The feature matrix is ​​converted into a 1×(k+n) weight matrix. Finally, the fused features are input to the terminal fully connected layer for smoke recognition. The fusion process is:

[0053]

[0054] Among them, the parameters and It is obtained by the network through back propagation autonomous learning, F st and F dy Respectively represent the static features and dynamic features used for fusion, The fused features.

[0055] The adaptive feature fusion method fuses the static features of the smoke image with the dynamic features in the smoke video through automatic selection by the network to generate features for smoke recognition, which has higher flexibility.

[0056] Step 3: Train the video smoke recognition network model constructed in step 2 using the smoke data set in step 1 to obtain an optimized video smoke recognition network model.

[0057] The specific training method is the gradient descent method to update the training parameters.

[0058] Step 4: Use the optimized video smoke recognition network model to process the images in the real-time collected smoke video. The static feature extraction network extracts static features in the spatial domain, and the dynamic feature extraction network extracts dynamic features in the time domain. The static features and dynamic features are fused to generate smoke features, and the smoke features are identified to determine whether there is smoke. If there is smoke, an alarm is triggered.

[0059] The adaptive feature fusion method fuses the static features of smoke images with the dynamic features in smoke videos to generate summary features for smoke recognition.

[0060] Since the fused feature map is a multi-dimensional matrix, the input of the fully connected layer is a 1-dimensional vector. The common practice is to stretch the multi-dimensional matrix into a vector as the input of the fully connected layer. The fully connected layer will output a vector [y1, y2] containing two elements. y1 and y2 are decimals from 0 to 1, representing the probability of smoke or no smoke, respectively. For example, y1 represents the probability of smoke and y2 represents the probability of no smoke. If y1 is greater than y2 ([0.8, 0.2]), it means that the recognition result is smoke. If y2 is greater than y1 ([0.3, 0.7]), it means that the recognition result is no smoke. Smoke type: There are only two results, smoke or no smoke.

[0061] In order to verify the effectiveness of the RAB model proposed in the present invention, the static feature extraction network was trained and tested on the dataset set1. Set1 is the image in the dataset, and the following set2 is the video in the dataset. For further comparison, the experiment selected nine mainstream and advanced network models for comparative experiments with the static feature extraction network. Specifically including: AlexNet, ZF-Net, GoogLeNet, VGG16, ResNet, DenseNet, SE-Net, SqueezeNet, MobileNet. The 10 models in this experiment were trained and tested on the dataset set1, and the specific experimental results are shown in Table 1. As can be seen from Table 1, the RAB model achieved excellent results in all evaluation indicators.

[0062] Table 1 Comparison of different smoke image recognition algorithms

[0063]

[0064] The experimental data showed that the static feature extraction network with the RAB module achieved the highest accuracy of 96.61%, the recall rate of 95.69%, and the lowest false alarm rate of 4.45%. Compared with the SE-Net structure ranked second, the accuracy rate was 1.46% higher, the recall rate was 0.84% ​​higher, and the false alarm rate was 2.0% lower, indicating that the RAB module is more sensitive to smoke images and the system has a higher credibility when it comes to smoke alarms.

[0065] By training and testing the RAB module multiple times and analyzing the false positives and missed positives in each experiment, we found that 68% to 75% of the missed positives and false positives were interference samples caused by clouds, fog, and pure color backgrounds. Figure 6 Therefore, in order to further reduce the false alarm rate and improve the accuracy and robustness of smoke recognition, W-3D CNN is used to extract the dynamic features of smoke, and the fused features are used for smoke recognition. The effectiveness of the overall method is verified by experimental data.

[0066] In order to verify the overall performance of the video smoke recognition network proposed in the present invention, the optimized fusion smoke recognition network model of the present invention is compared with the better performing VGG16, ResNet, DenseNet, SE-Net, method [1] - literature Real-time video-based smoke detection with high accuracy and efficiency, method [2] - literature Visual Smoke Detection Based on Ensemble Deep CNNs using the dataset set2. The experimental results are shown in Table 2.

[0067] Table 2 Comprehensive comparison between the present invention and mainstream networks

[0068]

[0069] From the data in Table 2, it can be seen that VGGNet has the lowest accuracy (ACC) and recall rate (Recall), and the highest false alarm rate (FAR); ResNet and SE-Net perform better than VGGNet, but worse than other methods. This also shows that the general recognition model based on deep learning cannot show excellent performance in complex tasks (such as video smoke recognition). Methods [1] and [2] use traditional manual algorithms to extract smoke motion features. Although they can effectively extract low-frequency information in smoke, they ignore some high-frequency information and therefore do not show the best performance. The 3D convolution and residual attention network proposed in this invention can not only better extract static information with stronger representation ability in smoke, but also can simultaneously extract high-frequency information and low-frequency information in smoke motion features through the improved 3D CNN, and redistribute weights to them, which can effectively reduce the interference of smoke-like targets. Its accuracy is 98.73% and recall is 98.24%, both reaching the highest level; the false alarm rate is reduced to the lowest 1.06%. Compared with the performance of the static feature extraction network in Table 2, after adding the W-3D CNN network, the accuracy rate increased by 2.12%, the recall rate increased by 2.55%, and the false alarm rate decreased by 3.39%. At the same time, after multiple experiments, the present invention can achieve an average detection rate of 48 frames per second, which can meet the real-time detection of common surveillance videos of 25 to 30 frames per second.

[0070] In summary, the method proposed in the present invention is significantly ahead of other methods in three indicators and can effectively identify and warn smoke in real time. Through a large number of experiments and comparisons, it is verified that the method of the present invention has higher detection accuracy (98.73%), recall rate (98.24%) and lower false alarm rate (1.06%) than the existing methods.

[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A smoke recognition method based on dual-channel fire video in spatiotemporal domain. It is characterized in that The steps are as follows: Step 1: Collect and create a smoke dataset containing cloud and fog interference images and videos; Step 2: Build a static feature extraction network and a dynamic feature extraction network, fuse and connect the static feature extraction network and the dynamic feature extraction network, and build a video smoke recognition network model; The static feature extraction network is built based on the residual attention module; the static feature extraction network is connected to 12 residual attention modules in sequence, and a pooling layer is connected after every 2 residual attention modules; the residual attention blocks shown all use a 3×3 convolution kernel with a step size of 1; the pooling layers all use a maximum pooling of 2×2 with a step size of 2; the static feature extraction network uses a ReLU nonlinear non-saturated activation function; Step 3: Use the smoke data set in step 1 to train the video smoke recognition network model constructed in step 2 to obtain an optimized video smoke recognition network model; Step 4: Use the optimized video smoke recognition network model to process the real-time collected smoke video. The static feature extraction network extracts the static features in the image space domain, and the dynamic feature extraction network extracts the dynamic features of the video sequence in the time domain. The static features and dynamic features are fused to generate smoke features, and the smoke features are identified to determine whether there is smoke. If there is smoke, an alarm is issued.

2. According to the method for smoke recognition based on dual-channel fire video in time and space domains in claim 1, It is characterized in that The smoke dataset in step 1 includes smoke images or videos in forest, field, indoor, playground, construction site, city, and highway scenes. The smoke dataset includes images and videos of multiple positive samples and negative samples.

3. The smoke recognition method based on dual-channel fire video in spatiotemporal domain according to claim 1 or 2, It is characterized in that The video smoke recognition network model includes a static feature extraction network and a dynamic feature extraction network connected in parallel, both of which are connected to a fusion component, and the fusion component is connected to a fully connected unit; the fusion component adopts an adaptive fusion method and utilizes the learning ability of a neural network to reallocate weights for fusion features.

4. The method for fire video smoke recognition based on dual-channel time-space domain according to claim 3, It is characterized in that The fusion component includes a feature fusion unit, which combines the static features extracted by the static feature extraction network and the dynamic features extracted by the dynamic feature extraction network: a set of 1×(n+k) feature vectors I is obtained by using 3D global average pooling, and the feature vectors are transformed into The feature matrix I is processed by convolution to obtain the weight matrix II. The feature matrix II is converted into a 1×(n+k) weight vector II; the weight vector II is multiplied by the corresponding static features and dynamic features to obtain the fused features, which are input into the fully connected layer of the fully connected unit.

5. The method for smoke recognition based on dual-channel fire video in spatiotemporal domain according to claim 4, It is characterized in that The fusion method of the video smoke recognition network model is: Among them, the parameters and After back propagation and autonomous learning, F st and F dy Respectively represent the static features and dynamic features used for fusion, The fused features.

6. The method for fire video smoke recognition based on dual-channel time-space domain according to any one of claims 1, 4 or 5, It is characterized in that The residual attention block includes a channel attention unit, a spatial attention unit, and a residual structure. The input feature map X is transmitted to the channel attention unit through the convolution operation in the backbone branch, and then the feature map obtained after being processed by the channel attention unit is transmitted to the spatial attention unit, and the feature map obtained after being processed by the spatial attention unit The feature map is the feature map after reassigning weights. The feature map is added to the input feature map X in the skip branch to obtain the output feature map 7. The method for fire video smoke recognition based on dual-channel time-space domain according to claim 6, It is characterized in that The dynamic feature extraction network is a weighted 3D convolutional neural network, and the dynamic feature extraction network includes a feature extraction module and an attention module, the feature extraction module includes at least two feature extraction units connected in sequence, each feature extraction unit includes a 3D convolution layer and a pooling layer connected in sequence; the feature map processed by the feature extraction module is subjected to convolution operation to obtain feature maps F = {f 1 ,f 2 ,...,f k } and attention map A = {a 1 ,a 2 ,...,a k }, attention map A = {a 1 ,a 2 ,...,a k } k As weights and feature maps F = {f 1 ,f 2 ,...,f k The feature f in k Multiply them sequentially to get the dynamic feature F dy .

8. The method for fire video smoke recognition based on dual-channel time-space domain according to claim 7, It is characterized in that The feature extraction module of the dynamic feature extraction network includes 5 feature extraction units. The 3D convolution layers in the feature extraction units all use 3D convolution kernels of size 3×3×3, stride of 1×1×1, and padding attribute padding=1; pooling layer II is 2×2×2 3D maximum pooling.

Citation Information

Patent Citations

  • Image classification method based on perception loss and matching attention mechanism

    CN108647736A

  • Behavior recognition technical method based on deep learning

    CN110188637A