A Video Smoke Detection Method for an End-to-End 3D Convolutional Object Detection Network
Through the end-to-end three-dimensional convolution target detection network, combined with the static and dynamic characteristics of smoke, the problem of high false alarm rate in video smoke detection is solved, and the effect of accurately identifying and positioning smoke is achieved.
Patent Information
- Application Number
- CN202210109359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Existing video smoke detection methods cannot effectively combine the static and dynamic characteristics of smoke, resulting in high false alarm rates and the inability to accurately identify and locate smoke.
The end-to-end three-dimensional convolution object detection network is adopted, including three-dimensional convolutional layer, cross-stage local residual network module, pyramid pooling module, path aggregation network module and tensor decoder. By performing data augmentation and feature fusion of video frame sequences, the static and dynamic features of smoke are extracted.
It improves the reliability of video smoke detection, can accurately identify and locate smoke, and reduces false alarm rate.
Smart Images

Figure CN114550032B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of video fire detection and deep learning, and particularly to a video smoke detection method and system based on an end-to-end three-dimensional convolutional object detection network. Background Art
[0002] Currently, there are mainly the following four types of research on video smoke detection using deep learning methods: (1) Independently detecting each video frame and achieving real-time detection with a high detection speed. This method completely fails to utilize the temporal information contained between consecutive video frames, so there are inevitably serious false negatives and false positives. (2) Using traditional motion detection algorithms to extract motion regions and then using DCNN to detect the motion regions. This method only utilizes shallow temporal information. Although it can eliminate false positives caused by some static objects, it is powerless against interference or false negatives of moving objects. (3) First, independently detecting each video frame, and then using a temporal network to make a judgment when a suspected target is detected. Although this method extracts effective deep dynamic features, its role is only to verify the detection results of the object detection network, which can eliminate false positives but is helpless against false negatives. In addition, this algorithm is not in an end-to-end form, so the running speed is often slow. (4) Using a temporal network to construct a classifier for video segments. This method can fully extract the motion features and static features contained in the video, but these features are only used for classification and do not locate the smoke target. Therefore, how to effectively extract the static and dynamic features of smoke and improve the reliability of smoke detection has become an urgent problem to be solved. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a video smoke detection method and system based on an end-to-end three-dimensional convolutional object detection network.
[0004] The technical solution of the present invention is as follows: A video smoke detection method based on an end-to-end three-dimensional convolutional object detection network includes:
[0005] Step S1: Obtaining video frames from multiple smoke videos, grouping them, constructing a video frame sequence, performing data augmentation on it, and constructing an augmented data set;
[0006] Step S2: Inputting the augmented data set into a three-dimensional convolutional smoke detection network, where the smoke detection network includes: a three-dimensional convolutional layer, a cross-stage local residual network module, a pyramid pooling module, a path aggregation network module, and a tensor decoder; outputting a smoke recognition result and localization.
[0007] Compared with the prior art, the present invention has the following advantages:
[0008] The present invention discloses a video smoke detection method based on an end-to-end three-dimensional convolutional object detection network, which can effectively extract the static and dynamic features of smoke. The combination of dynamic features and static features can effectively improve the reliability of the video smoke detection algorithm, thereby accurately identifying and locating the smoke in the video frame. This method can be applied in the field of video fire detection, has high application value, and provides a new method to solve the problem of high false alarm rate that plagues current video fire detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of a video smoke detection method based on an end-to-end three-dimensional convolutional object detection network in an embodiment of the present invention;
[0010] Figure 2 It is a schematic structural diagram of an end-to-end three-dimensional convolutional object detection network in an embodiment of the present invention;
[0011] Figure 3 It is a schematic structural diagram of a cross-stage partial residual network module in an embodiment of the present invention;
[0012] Figure 4 It is a structural block diagram of a video smoke detection system based on an end-to-end three-dimensional convolutional object detection network in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The present invention provides a video smoke detection method based on an end-to-end three-dimensional convolutional object detection network, which can effectively extract the static and dynamic features of smoke. The combination of dynamic features and static features can effectively improve the reliability of the video smoke detection algorithm, thereby accurately identifying and locating the smoke in the video frame.
[0014] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention through specific embodiments and in conjunction with the accompanying drawings.
[0015] Embodiment 1
[0016] As Figure 1 shown, a video smoke detection method based on an end-to-end three-dimensional convolutional object detection network provided by an embodiment of the present invention includes the following steps:
[0017] Step S1: Obtain video frames from multiple smoke videos, group them to construct a video frame sequence, perform data augmentation on it, and construct an augmented data set;
[0018] Step S2: Input the augmented data set into a three-dimensional convolutional smoke detection network, where the smoke detection network includes: a three-dimensional convolutional layer, a cross-stage partial residual network module, a pyramid pooling module, a path aggregation network module, and a tensor decoder; output the smoke recognition result and location.
[0019] In one embodiment, step S1: Obtain video frames from multiple smoke videos, group them, construct a video frame sequence, perform data augmentation on it, and construct an augmented dataset, specifically including:
[0020] Step S11: Obtain multiple smoke videos, extract images at a fixed frame interval, construct a video frame sequence for every 100 images extracted, and label each image containing smoke among them;
[0021] In the embodiment of the present invention, 44 segments of videos meeting the requirements are obtained from the publicly available fire smoke video image database, among which 32 segments of videos have smoke and are used to produce positive samples, and 12 segments of videos have no smoke and are used to produce negative samples. In addition, 28 segments of videos are shot as supplements, among which 21 segments of videos have smoke, and 7 segments of videos are backgrounds and interferences such as pedestrians. The above videos include three scenarios: indoor, outdoor close range, and outdoor long range.
[0022] Extract images from the above videos at a fixed interval of 3 frames per second. Every 100 pictures form a video frame sequence, and finally 147 video frame sequences are obtained, with a total of 14,700 pictures. Among them, there are 115 video frame sequences of positive samples and 32 video frame sequences of negative samples. Each positive sample has a corresponding label file, and the label file is made by the labelImg software. In the embodiment of the present invention, 132 video sequences, a total of 13,200 pictures, are randomly selected as the training set, and the remaining 15 sequences are used as the validation set.
[0023] Step S12: Generate two random numbers a and b, such that 0 < a < 132 and 0 < b < 89, and let i = a × 100 + b; Extract 12 consecutive images starting from the i-th picture in the dataset, that is, the b-th picture in the a-th video frame sequence, to obtain a new video frame sequence;
[0024] Since 12 pictures need to be continuously read and input into the network during subsequent network training, in order to ensure that the 12 pictures read each time come from the same video sequence, the embodiment of the present invention designs a reading rule for video sequences. When reading training data, first generate two random integers a and b, where the value range of a is between 0 and 132, and the value range of b is between 0 and 89. Let the variable i be equal to a × 100 + b, and then start reading from the i-th picture, and read 12 pictures in sequence. After the picture reading is completed, the label file corresponding to the last picture is read in as the label for this training to calculate the loss value. The same rule is adopted when reading validation data, and the value range of a becomes between 0 and 15.
[0025] Step S13: Select 4 new video frame sequences. After performing enhancement transformations on the images therein, sequentially downsample and splice the images in the 4 video frame sequences in the same manner to obtain an enhanced video frame sequence, thereby constructing an enhanced dataset; the enhancement transformations include: flipping, cropping, size transformation, translation transformation, and color gamut distortion.
[0026] Common data augmentation methods for object detection include image flipping, cropping, size transformation, translation transformation, rotation transformation, color gamut distortion, etc. Existing data augmentation algorithms perform random transformations on each image and then input it into the neural network for training. However, when the training samples are video sequences, the same transformation should be ensured for each image in the sequence. Therefore, existing data augmentation algorithms cannot be directly applied to the present invention.
[0027] The embodiment of the present invention designs a data augmentation algorithm for the video frame sequence of smoke. This algorithm includes functions such as image flipping, cropping, size transformation, translation transformation, and color gamut distortion. Since the movement of smoke has a certain directionality, the embodiment of the present invention does not include rotation transformation. In addition, since the smoke detection network in the embodiment of the present invention takes multiple consecutive video frame sequences as input, the computing device has a large pressure, so it is difficult to perform batch training, resulting in a relatively long training cycle and poor model robustness. To solve this problem, the data augmentation algorithm designed by the present invention can process four video frame sequences at a time. After the four sequences are respectively transformed, the images in the four sequences are sequentially downsampled and spliced in the same manner, and finally a new sequence is formed. This method can calculate the data of four images for each iterative training, enriching the background of the detected object. In addition, the spliced images contain the downsampled smoke targets, and these images will improve the sensitivity of the detection model to small targets, making it more suitable for early fire detection.
[0028] In one embodiment, as Figure 2 shown in the structural schematic diagram of the three-dimensional convolutional smoke detection network, the above step S2: Input the enhanced dataset into the three-dimensional convolutional smoke detection network. The smoke detection network includes: three-dimensional convolutional layers, cross-stage local residual network modules, pyramid pooling modules, path aggregation network modules, and tensor decoders; output the smoke recognition result and localization, specifically including:
[0029] Step S21: The three-dimensional convolutional layers include 4 three-dimensional convolutional layers. The input is 12×416×416×3, where 12 is 12 images in an enhanced video frame sequence, 416×416 is the image size, and 3 is the RGB channels of the image. After passing through each three-dimensional convolutional layer, the length and width are reduced by half, and the output is 1×52×52×128;
[0030] Step S22: The cross-stage partial residual network module is composed of 11 cascaded two-dimensional depthwise separable convolutions. The output of the three-dimensional convolutional layer is selected at three scales of 52×52, 26×26, and 13×13 for output decoding;
[0031] As Figure 3 shown in the schematic diagram of the cross-stage partial residual network module structure, the embodiment of the present invention introduces a residual structure and a cross-stage partial network structure. The residual structure is used to solve the problem of gradient disappearance when the network is too deep, and the cross-stage partial network structure can enhance the learning ability of the network. In order to enable the three-dimensional convolutional smoke detection network to effectively detect smoke targets of different sizes, the embodiment of the present invention selects three scales of 52×52, 26×26, and 13×13 for output.
[0032] Step S23: The pyramid pooling module performs enhanced regional feature processing on the 3 outputs of the cross-stage partial residual network module respectively;
[0033] Step S24: The output of the pyramid pooling module is subjected to feature fusion between scales via a path aggregation network module to obtain three feature tensors with sizes of 52×52×18, 26×26×18, and 13×13×18;
[0034] The embodiment of the present invention uses a path aggregation network for feature fusion between scales. The small-scale feature tensor is upsampled and fused with the medium-scale and large-scale feature tensors in turn, and then the large-scale feature is downsampled and fused with the medium-scale and small-scale feature tensors in turn, so as to shorten the information path and use the precise positioning signals existing in the large-scale feature tensor to enhance the feature pyramid.
[0035] Step S25: The three feature tensors are respectively input into the corresponding tensor decoder for decoding, and finally the smoke recognition result and positioning are output.
[0036] The three feature tensors with sizes of 52×52×18, 26×26×18, and 13×13×18 obtained by the path aggregation network module are respectively input into the corresponding yolo_head for tensor decoding to obtain the smoke recognition result and positioning. The working principle of yolo_head is as follows. Taking the tensor with a size of 52×52×18 as an example: First, the original picture is divided into a 52×52 grid. Each cell will predict 3 potential targets, and each target corresponds to 6 parameters, namely 4 bounding box parameters, 1 confidence, and 1 class probability value.
[0037] The present invention discloses a video smoke detection method based on an end-to-end three-dimensional convolutional object detection network, which can effectively extract the static and dynamic features of smoke. The combination of dynamic features and static features can effectively improve the reliability of the video smoke detection algorithm, thereby accurately identifying and locating the smoke in the video frame. This method can be applied in the field of video fire detection, has high application value, and provides a new method to solve the problem of high false alarm rate that plagues current video fire detection.
[0038] Embodiment 2
[0039] As Figure 4 shown, an embodiment of the present invention provides a video smoke detection system based on an end-to-end three-dimensional convolutional object detection network, including the following modules:
[0040] The dataset construction module 31 obtains video frames from multiple smoke videos, groups them, constructs a video frame sequence, performs data augmentation on it, and constructs an augmented dataset;
[0041] The detection network construction and training module 32 is used to input the augmented dataset into the three-dimensional convolutional smoke detection network. The smoke detection network includes: a three-dimensional convolutional layer, a cross-stage local residual network module, a pyramid pooling module, a path aggregation network module, and a tensor decoder; and outputs the smoke recognition result and location.
[0042] The above embodiments are provided only for the purpose of describing the present invention, and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. All equivalent substitutions and modifications made without departing from the spirit and principles of the present invention shall be covered within the scope of the present invention.
Claims
1. A video smoke detection method based on an end-to-end three-dimensional convolutional object detection network, characterized in that Including: Step S1: Obtain video frames from multiple smoke videos, group them, construct a video frame sequence, perform data augmentation on it, and construct an augmented dataset, specifically including: Step S11: Obtain multiple smoke videos, extract images at fixed intervals, construct a video frame sequence for every 100 images extracted, and label each image containing smoke among them; Step S12: Generate two random numbers a and b, such that 0 < a < 132 and 0 < b < 89, and let i = a × 100 + b; start from the i-th image in the augmented dataset, that is, the b-th image in the a-th video frame sequence, and extract 12 consecutive images to obtain a new video frame sequence; Step S13: Select 4 of the new video frame sequences, perform augmentation transformations on the images therein, and then sequentially shrink and splice the images in the 4 video frame sequences in the same way to obtain an augmented video frame sequence, thereby constructing an augmented dataset; the augmentation transformations include: flipping, cropping, size transformation, translation transformation, and color gamut distortion; Step S2: Input the augmented dataset into a 3D convolutional smoke detection network, which includes: a 3D convolutional layer, a cross-stage local residual network module, a pyramid pooling module, a path aggregation network module, and a tensor decoder; output the smoke recognition result and location, specifically including: Step S21: The 3D convolutional layer includes 4 3D convolutional layers, with an input of 12×416×416×3, where 12 is the 12 images in an augmented video frame sequence, 416×416 is the image size, and 3 is the RGB channel of the image. After passing through each 3D convolutional layer, both the length and width are halved, and the output is 1×52×52×128; Step S22: The cross-stage local residual network module is composed of 11 cascaded 2D depthwise separable convolutions, and the output of the 3D convolutional layer is selected for output decoding at three scales of 52×52, 26×26, and 13×13; Step S23: The pyramid pooling module performs enhanced regional feature processing on the 3 outputs corresponding to the cross-stage local residual network module respectively; Step S24: The output of the pyramid pooling module undergoes feature fusion between scales via the path aggregation network module to obtain 3 feature tensors of sizes 52×52×18, 26×26×18, and 13×13×18; Step S25: Input the 3 feature tensors into the corresponding tensor decoders for decoding, and finally output the smoke recognition result and location.
2. A video smoke detection system based on an end-to-end 3D convolutional object detection network, characterized in that, Including the following modules: A dataset construction module that obtains video frames from multiple smoke videos, groups them, constructs a video frame sequence, performs data augmentation on it, and constructs an augmented dataset, specifically including: Step S11: Obtain multiple smoke videos, extract images at fixed intervals, construct a video frame sequence for every 100 images extracted, and label each image containing smoke among them; Step S12: Generate two random numbers a and b, where 0 < a < 132 and 0 < b < 89, and let i = a × 100 + b; starting from the i-th picture in the enhanced dataset, that is, the b-th picture in the a-th video frame sequence, extract 12 consecutive images to obtain a new video frame sequence; Step S13: Select 4 of the new video frame sequences. After performing enhancement transformations on the images therein, sequentially stitch the images in the 4 video frame sequences in the same manner after shrinking to obtain an enhanced video frame sequence, thereby constructing an enhanced dataset; the enhancement transformations include: flipping, cropping, size transformation, translation transformation, and color gamut distortion; The detection network construction and training module is used to input the enhanced dataset into a three-dimensional convolutional smoke detection network. The smoke detection network includes: a three-dimensional convolutional layer, a cross-stage local residual network module, a pyramid pooling module, a path aggregation network module, and a tensor decoder; output the smoke recognition result and localization, specifically including: Step S21: The three-dimensional convolutional layer includes 4 three-dimensional convolutional layers. The input is 12×416×416×3, where 12 is the 12 images in an enhanced video frame sequence, 416×416 is the image size, and 3 is the RGB channel of the image. After passing through each three-dimensional convolutional layer, both the length and width are reduced by half, and the output is 1×52×52×128; Step S22: The cross-stage local residual network module is composed of 11 two-dimensional depthwise separable convolutions in series. Select three scales of 52×52, 26×26, and 13×13 from the output of the three-dimensional convolutional layer for output decoding; Step S23: The pyramid pooling module performs enhanced regional feature processing on the 3 outputs corresponding to the cross-stage local residual network module respectively; Step S24: The output of the pyramid pooling module is subjected to feature fusion between scales via the path aggregation network module to obtain 3 feature tensors with sizes of 52×52×18, 26×26×18, and 13×13×18; Step S25: Input the 3 feature tensors into the corresponding tensor decoders for decoding, and finally output the smoke recognition result and localization.
Citation Information
Patent Citations
Smoke detection method and system based on pseudo 3D convolutional neural network
CN111553403A
Target detection method, medium and system
CN112183578A
Multi-video-frame black smoke diesel vehicle detection method and system based on space-time optical flow network
CN113221976A