Flame area detection method in surveillance video based on lightweight two-stream convolutional network

By adopting a lightweight dual-stream convolutional network in flame area detection, combining the spatial and temporal characteristics of the video block, the existing detection network model is solved, and the flame area detection effect is achieved with high precision and high real-time.

CN113221793BActive Publication Date: 2025-05-09NANJING FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110562235.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-24
Publication Date
2025-05-09
Estimated Expiration
2041-05-24

AI Technical Summary

Technical Problem

The existing flame area detection network model has the problems of complex network structure, difficult training, and slow detection speed, and is difficult to apply to real-time detection tasks. The flame objects are not included in the mainstream public dataset and have diverse shapes, resulting in misjudgment when directly migrating the network model on the public dataset for flame area detection.

Method used

A lightweight dual-stream convolution network is adopted. By dividing the video to be detected into video blocks, extracting the intermediate frame of the video block as the spatial stream input, and calculating the differential image as the time stream input. The convolution feature maps output by the two branch networks are merged and fused, and the fused feature maps are analyzed using a 3-layer 1×1 convolutional layer. Finally, the Softmax classifier is used to determine whether the 2n×2n pixel block is a flame area.

Benefits of technology

High accuracy and real-time performance of flame area detection are achieved. By converting pixel-level segmentation problems into classification problems, using a lighter network structure, the detection speed is improved, and the receptive field is adaptively adjusted through the SK-Shuffle convolution module, improving the accuracy of detection of smaller flame areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113221793B_ABST
    Figure CN113221793B_ABST
Patent Text Reader

Abstract

A method for detecting flame regions in surveillance videos based on a lightweight two-stream convolutional network. First, the video to be detected is segmented into several video blocks, and the first frame, middle frame, and last frame are extracted from a single video block to calculate two difference images. Then, the middle frame of the video block and the two difference images are respectively input into the spatial stream branch and temporal stream branch of the convolutional network to obtain two feature maps with both the length and width reduced to 1 / 2<supgt;n< / supgt> of the original image. Next, the convolutional feature maps output by the two branches are merged, and a 3-layer 1×1 convolutional layer is used to analyze the feature vectors on each channel of the fused feature map to obtain a region determination feature map. Finally, a Softmax classifier is applied to each 1×1×2 element on the region determination feature map to determine whether the 2<supgt;n< / supgt>×2<supgt;n< / supgt> pixel block corresponding to this element on the original image is a flame region. The fire surveillance video flame region detection model constructed by this method has high detection accuracy and fast detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a flame area detection method, in particular to a monitoring video flame area detection method based on a lightweight dual-stream convolutional network. Background Art

[0002] At present, there are few network models for video-based flame area detection, and current research proposes more flame image detection methods based on convolutional networks. However, sometimes firefighters are mainly concerned about the spread of flames or the burning area in the flame image. At this time, the detection model is needed to extract the flame area rather than just find the flame image. There are currently some deep network models for target detection and image segmentation, but most of these networks are designed for public datasets of multiple objects. Therefore, there are problems such as complex network structure, difficult training, and slow detection speed, which makes it difficult to apply them to real-time detection tasks. In addition, flame objects are not included in the current mainstream public datasets, and the shapes of flames are diverse. Relying only on static spatial features, even the human eye will have misjudgments. These factors are not conducive to directly migrating network models on public datasets for flame area detection. Summary of the invention

[0003] The purpose of the present invention is to provide a flame area detection method based on fixed-point video monitoring, which can quickly extract and analyze the spatial and temporal features of the monitoring video, thereby maintaining high precision and high real-time performance in the flame area detection task.

[0004] In order to solve the above technical problems, the present invention provides a monitoring video flame area detection method based on a lightweight two-stream convolutional network, the steps comprising:

[0005] 1) First, the video to be detected is divided into several video blocks, the first frame, the middle frame and the last frame are extracted from a single video block, and two differential images are calculated;

[0006] 2) Then, the middle frame of the video block is used as the spatial stream input of the flame area detection network, and the two differential images are used as the temporal stream input of the flame area detection network. The convolutional feature maps output by the two branch networks have the same size, and the length and width are reduced to 1 / 2 of the original image. n , n is a natural number;

[0007] 3) Then, the convolutional feature maps output by the two branches of the flame area detection network are merged and fused;

[0008] Use three 1×1 convolutional layers to analyze the feature vectors on each channel of the fused feature map and obtain the region determination feature map;

[0009] 4) Finally, the Softmax classifier is used to determine the element on the regional feature map to determine whether the element corresponds to the 2 n ×2 n Whether the pixel block is a flame area;

[0010] In the step 2), the network structure of each branch in the flame area detection network includes: 2 layers of standard convolution, 8 layers of SK-Shuffle convolution and 4 layers of maximum pooling layers with a step size of 2;

[0011] Here, we take n as 4. The SK-Shuffle convolution enables the network to adaptively adjust the receptive field size for detection targets of different shooting distances and sizes, thereby improving detection accuracy. The network contains 8 layers of SK-Shuffle convolution to facilitate setting the adaptive adjustment range of the receptive field. The 4-layer maximum pooling layer with a step size of 2 enables the length and width of the network output feature map to be downsampled to 1 / 16 of the input original image.

[0012] The structure of the SK-Shuffle convolution is to replace the depth convolution in the shuffleNet V2 convolution with the SK depth convolution. The SK depth convolution is obtained by replacing the convolution operation on each branch of the SK convolution with the depth convolution.

[0013] SK convolution includes Split operation, Fuse operation and Select operation;

[0014] The Split operation includes: for any given feature map, convolution is performed using multiple convolution kernels of different sizes to obtain feature maps on multiple branches, each of which carries receptive field information of different sizes;

[0015] The Fuse operation includes: first, adding the elements of the corresponding positions of the feature maps on the multiple branches separated by the Split operation, and then using global average pooling to integrate the global information; finally, sending the output after global average pooling to a fully connected layer;

[0016] The Select operation includes: first, according to the number of branches included in the Split operation, the corresponding number of full connection and Softmax operations are performed on the basis of the output of the fully connected layer of the Fuse operation, so as to obtain the adaptive weight parameters of each channel of the convolution feature map on each branch, and finally, the values ​​of the corresponding channels of the convolution feature map on each branch are weighted and summed in combination with the adaptive weight parameters to obtain the output feature map of the SK convolution.

[0017] Specifically:

[0018] In the step 1), the two differential images are calculated using a three-frame difference method;

[0019] Since the three-frame difference method requires the calculation of frame difference images every two frames, a total of 7 frames are needed for 2 frame difference images. Therefore, each video block includes 7 frames of images, and the middle frame is extracted from a single video block as the input of the spatial stream; with the middle frame as the center, the three-frame difference method is used to calculate 2 differential images as the input of the temporal stream.

[0020] In step 3), three 1×1 convolutional layers are used to replace the fully connected layers of the traditional classification network to further extract features from the fused feature map, and obtain a region determination feature map to support input images of any size.

[0021] In steps 2) and 4), n is 4:

[0022] In step 4), a Softmax classifier is used for each 1×1×2 element on the regional judgment feature map to determine whether the element corresponds to 2 on the original image. n ×2 n Whether the pixel block is a flame area.

[0023] The beneficial effects of the present invention are:

[0024] (1) The present invention creates a lightweight convolutional network (i.e., a flame area detection network), the detection effect of which is equivalent to dividing the video frame into two n ×2 n Pixel blocks (such as Figure 4 As shown, n is 4), and each 2 n ×2 n Whether the pixel block is a flame area.

[0025] Compared with the existing area detection network, the invention transforms the pixel-level segmentation problem into a classification problem, so that the flame area detection can be completed with a lighter network structure and faster speed;

[0026] (2) The present invention creates an SK-Shuffle convolution that integrates the SK convolution and ShuffleNet V2 convolution modules in a convolutional network, so that the network can adaptively adjust the size of the receptive field and suppress the increase of network parameters, which can improve the network's detection accuracy for smaller flame areas or flame areas with a long shooting distance.

[0027] The present invention can be used for fire monitoring, and the fire monitoring video flame area detection model constructed by the detection method has high detection accuracy and fast detection speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of the flame area detection method created by the present invention;

[0029] Figure 2Schematic diagram of the structure of the SK convolution module.

[0030] Figure 3 Schematic diagram of SK-Shuffle convolution.

[0031] Figure 4 A map of the locations of network input and output. DETAILED DESCRIPTION

[0032] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0033] like Figure 1 In this example, a flame area detection method for surveillance video based on a lightweight two-stream convolutional network is described, and the detection steps include:

[0034] Step 1: Divide the video to be detected into several video blocks, each of which includes 7 frames of images. Extract the first frame, the middle frame and the last frame from a single video block, and calculate two differential images using the three-frame difference method.

[0035] Step 2: The middle frame of the video block obtained in step 1 is used as the spatial stream input of the flame area detection network, and the two differential images are used as the temporal stream input of the flame area detection network. The convolution feature maps output by the two branch networks have the same size, and the length and width are reduced to 1 / 2 of the original image. n ;

[0036] Step 3: merge the convolution feature maps output by the two convolutional networks, use three 1×1 convolutional layers to analyze the feature vectors on each channel of the merged feature map and obtain the region determination feature map;

[0037] Step 4: Use the Softmax classifier for each 1×1×2 element on the region determination feature map to determine whether the element corresponds to 2 on the original image. n ×2 n Pixel blocks (such as Figure 4 As shown, n is 4) whether it is a flame area.

[0038] The overall structure of the flame area detection network is as follows: Figure 1 As shown in the figure, in this example, the flame area detection network includes two layers of standard convolution, 8 layers of SK-Shuffle convolution (SK convolution module and shuffleNet V2 module fusion) (as shown in Figure 3 ), and 4 max pooling layers with a stride of 2.

[0039] The structure diagram of the SK convolution module is as follows Figure 2 As shown, it includes 3 operations: Split, Fuse and Select.

[0040] The Split operation involves convolution of kernels of different sizes in multiple paths. Convolution kernels of different sizes correspond to neuron receptive fields of different sizes, such as Figure 2 In the above example, for any given feature map X, two convolution kernels of different sizes are used to obtain U 1 and U 2 , these two branches carry receptive field information of different sizes respectively.

[0041] The first step of the Fuse operation is to merge the feature maps (U 1 and U 2 ) are added to the corresponding positions of:

[0042] U=U 1 +U 2 (1)

[0043] Then, global average pooling is used on the added result U to integrate global information:

[0044]

[0045] Among them, U c is the cth channel of feature map U, W and H are the length and width of feature map U respectively, s c It is the value after global pooling of the cth (1≤c≤C)th channel of the feature map U, where C is the number of channels of U.

[0046] like Figure 2 As shown, the final step of the Fuse operation is to convert the global average pooled output s = [s 1 ,…s c ,…s C ], and feed it into a fully connected layer:

[0047]

[0048] Among them, δ represents the ReLU function, β is batch normalization, is the weight parameter of the fully connected layer, d is the dimension of the middle hidden layer of the fully connected layer,

[0049] The purpose of the Select operation is to use the output of the fully connected layer to select the weight parameters of the corresponding channels of the feature maps on multiple branches. First, according to the number of branches included in the Split operation, the corresponding number of full connection and Softmax operations are performed on the basis of the output of the fully connected layer shown in formula (3), as shown in Figure 2 The Split operation in includes two branches, and the output z of the first fully connected layer is subjected to two full connection and Softmax operations respectively:

[0050]

[0051] in, Respectively Figure 2 The weight parameters of the fully connected layers distributed on the two branches, a c , b c represents the weight parameter of the cth channel of the convolution feature map on the two branches, A c , B c ∈R 1×d Represent the c-th channel of A and B respectively.

[0052] The final feature map V is obtained by formula (5):

[0053] V c =a c ·U 1c +b c ·U 2c , a c +b c =1 (5)

[0054] Among them, V c is the cth channel of feature map V, U 1c and U 2c They are feature maps U 1 and U 2 The cth channel of .

[0055] After the network is constructed using SK convolution, the convolution kernel size K of each SK convolution layer is l It cannot be regarded as a fixed value, but will adjust adaptively with the input:

[0056] K l =∑ i ( k i∑ c w′ ci ) (6)

[0057] Among them, k i is the size of the convolution kernel on the i-th branch in the SK convolution, w′ ci Represents the weight of the cth channel of the convolution kernel on the i-th branch. It is precisely because the SK convolution can adaptively calculate the weight w′ through the fully connected layer ci , which makes the size of each convolution kernel change adaptively, thus affecting the size of the receptive field.

[0058] In order to reduce the parameters of SK convolution and improve the network detection speed, the present invention creates SK-Shuffle convolution which is a fusion of SK convolution module and shuffleNet V2 module. The structure of SK-Shuffle convolution is Figure 3It is shown in the figure that the depth convolution in the shuffleNet V2 convolution module is replaced by SK depth convolution. SK depth convolution means that the convolution operation on each branch of SK convolution is depth convolution.

[0059] In this example, it is assumed that the size of the middle frame of the input video block is W×H×3, and the size of the two differential images combined is W×H×2. The input, output and convolution kernel parameters of each layer of the flame area detection network are shown in Table 1.

[0060] Table 1 Structure and processing flow of flame area detection network

[0061]

[0062] The corresponding schematic diagram of the network output result described in step 4 and the 16×16 pixel block on the original image is as follows Figure 4 shown.

[0063] Through experimental comparison on the same data set, it is concluded that the detection speed of the network model proposed in the present invention is 7 times that of the FCN model and 2.5 times that of the DeepLab V3+ model; although the network model proposed in the present invention uses 16×16 pixel blocks as detection units and has a lower resolution than pixel-level segmentation networks such as FCN and DeepLab V3+, the outline of the flame area is not clear against most backgrounds, and the network model proposed in the present invention uses a dual-stream structure to integrate the spatiotemporal features of the video, so the detection accuracy is higher than that of FCN and DeepLab V3+.

Claims

1. A surveillance video flame area detection method based on a lightweight two-stream convolutional network, characterized by the following steps include: 1) First, the video to be detected is divided into several video blocks, the first frame, the middle frame and the last frame are extracted from a single video block, and two differential images are calculated; 2) Then, the middle frame of the video block is used as the spatial stream input of the flame area detection network, and the two differential images are used as the temporal stream input of the flame area detection network. The convolutional feature maps output by the two branch networks have the same size, and the length and width are reduced to 1 / 2 of the original image. n , n is a natural number; 3) Then, the convolutional feature maps output by the two branches of the flame area detection network are merged and fused; Use three 1×1 convolutional layers to analyze the feature vectors on each channel of the fused feature map and obtain the region determination feature map; 4) Finally, the Softmax classifier is used to determine the element on the regional feature map to determine whether the element corresponds to the 2 n ×2 n Whether the pixel block is a flame area; In the step 2), the network structure of each branch in the flame area detection network includes: 2 layers of standard convolution, 8 layers of SK-Shuffle convolution and 4 layers of maximum pooling layers with a step size of 2; The structure of the SK-Shuffle convolution is to replace the depth convolution in the shuffleNet V2 convolution with the SK depth convolution. The SK depth convolution is obtained by replacing the convolution operation on each branch of the SK convolution with the depth convolution. SK convolution includes split operation, fuse operation and select operation; The Split operation includes: for any given feature map, convolution is performed using multiple convolution kernels of different sizes to obtain feature maps on multiple branches, each of which carries receptive field information of different sizes; The Fuse operation includes: first, adding the elements of the corresponding positions of the feature maps on the multiple branches separated by the Split operation, and then using global average pooling to integrate the global information; finally, sending the output after global average pooling to a fully connected layer; The Select operation includes: first, according to the number of branches included in the Split operation, the corresponding number of full connection and Softmax operations are performed on the basis of the output of the fully connected layer of the Fuse operation, so as to obtain the adaptive weight parameters of each channel of the convolution feature map on each branch, and finally, the values ​​of the corresponding channels of the convolution feature map on each branch are weighted and summed in combination with the adaptive weight parameters to obtain the output feature map of the SK convolution.

2. The method for monitoring video flame area detection based on a lightweight two-stream convolutional network according to claim 1 is characterized in that In the step 1), the two differential images are calculated using a three-frame difference method; Since the three-frame difference method requires the calculation of frame difference images every two frames, a total of 7 frames are needed for 2 frame difference images. Therefore, each video block includes 7 frames of images, and the middle frame is extracted from a single video block as the input of the spatial stream; with the middle frame as the center, the three-frame difference method is used to calculate 2 differential images as the input of the temporal stream.

3. The method for flame area detection in surveillance video based on a lightweight two-stream convolutional network according to claim 1 is characterized in that In the steps 2) and 4), n is 4.

4. The method for monitoring video flame area detection based on a lightweight two-stream convolutional network according to claim 1 is characterized in that In step 3), three 1×1 convolutional layers are used to replace the fully connected layers of the traditional classification network to further extract features from the fused feature map and obtain a region determination feature map to support input images of any size.

5. The method for flame area detection in surveillance video based on a lightweight dual-stream convolutional network according to claim 1 or 3, characterized in that In step 4), a Softmax classifier is used for each 1×1×2 element on the region determination feature map to determine whether the element corresponds to 2 on the original image. n ×2 n Whether the pixel block is a flame area.

Citation Information

Patent Citations

  • Flame target detection method based on digital image and convolution features

    CN110751089A

  • Salient object detection method and system for weak supervision-based spatio-temporal cascade neural network

    WO2019136591A1