A smoke detection method based on suspected smoke proposal area and deep learning

Through smoke detection methods based on suspected smoke proposal areas and deep learning, the problem of low accuracy of feature repetition extraction and detection in the prior art is solved, and efficient and robust smoke detection is achieved.

CN114005090BActive Publication Date: 2025-05-23SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111342954.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-05-23
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

Existing smoke detection methods are prone to repeated extraction of features, waste of computing resources and low detection accuracy.

Method used

A smoke detection method based on suspected smoke proposal areas and deep learning is proposed. By constructing a deep learning neural network model, the suspected smoke proposal areas are used for screening and feature extraction, and deep feature extraction is performed directly to improve detection efficiency.

Benefits of technology

It improves the accuracy and efficiency of smoke detection, reduces the probability of missed reports, and can monitor wildfire smoke more robustly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005090B_ABST
    Figure CN114005090B_ABST
Patent Text Reader

Abstract

The present invention proposes a smoke detection method based on suspected smoke proposal area and deep learning, which relates to the technical field of smoke detection. First, RGB picture samples are collected, and then a deep learning neural network model is constructed. Before the picture samples are input into the deep learning neural network model, the entire sample picture is used to propose a suspected smoke area, and all suspected smoke proposal areas are screened. The candidate area covers the entire picture and the scales of the candidate area are distributed from small areas to large areas, which reduces the probability of missed reports. A bounding box bbox is generated based on the screened suspected smoke proposal area and its value is calculated. The value of the bounding box bbox is used as a spatial embedding vector input of a dynamic embedding layer DEN to improve the deep learning neural network model. Finally, the deep learning neural network model is trained, and the obtained model can detect wildfire smoke more robustly and monitor wildfire smoke more robustly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smoke detection, and more specifically, to a smoke detection method based on suspected smoke proposal area and deep learning. Background Art

[0002] Fire poses a great threat to human life and property safety. Early detection and timely handling of fire are of great significance to protecting life and property safety. Therefore, smoke and fire detection is one of the current research hotspots in the industry.

[0003] At present, smoke detection methods are mainly divided into physics-based and computer vision-based methods. The physics-based method refers to smoke alarms through smoke sensors. Its advantages are accuracy and timeliness, but its disadvantages are that it is only applicable to a small area and can only be used in narrow indoor spaces. For wildfire prevention and control, it requires the deployment of extremely dense equipment, which is costly and unfriendly to the ecological environment. Computer vision-based smoke detection methods are mainly divided into traditional methods and deep learning methods. Among them, traditional methods are mainly based on prior knowledge of human observation and manually extract features of smoke. These features are mainly divided into color, texture, motion and energy. The traditional method is followed by a discriminator. The advantage of this method is that the amount of calculation is small, and the classifier training does not require very large data and calculation. The disadvantage is that even if multiple features are combined, the accuracy of feature discrimination is not high.

[0004] The deep learning-based method regards smoke detection as a single target detection problem. The neural network model is trained using smoke images and their annotated data. It requires convolution of thousands of anchors one by one, which consumes a lot of computing resources. The network model training is completely dependent on data, and model fitting cannot be achieved in the case of small samples. At present, the combination of traditional methods and deep learning methods is also popular. For example, on November 5, 2019, a Chinese invention patent (publication number: CN110415260A) disclosed a smoke image segmentation and recognition method based on dictionary and BP neural network. In this scheme, the input image is first enhanced, and the enhanced image is compared with the smoke dictionary matrix to calculate the suspected smoke area. Then, the obtained suspected smoke image is subjected to multi-scale wavelet decomposition, and the statistics of the grayscale co-occurrence matrix of each layer of the image obtained after multi-scale wavelet decomposition are calculated. The statistics of the grayscale co-occurrence matrix of each layer of the image calculated are used to construct the texture feature vector. The texture feature vector is then input into the smoke recognition BP neural network to obtain the smoke recognition result. However, this scheme needs to build a background model first when extracting the suspected smoke area, which means that the suspected smoke area cannot be calculated from a single picture. This scheme only extracts a few areas, which often do not cover the entire picture. Once the candidate area misses the real smoke area, it will be missed, resulting in low smoke detection accuracy. Moreover, each candidate area must be manually extracted with a priori features, which is inefficient. If there is overlap between areas, it will lead to repeated extraction of features in some areas, wasting computing resources. Then the feature vectors corresponding to each area must be sent to the BP neural network for multiple tests. Summary of the invention

[0005] In order to solve the problem that the current traditional smoke detection method based on the combination of smoke feature extraction and deep learning is prone to repeated feature extraction, waste of computing resources and low detection accuracy, the present invention proposes a smoke detection method based on suspected smoke proposal area and deep learning. It only needs to send the original image directly to the neural network for one deep feature extraction, thereby improving the efficiency of smoke detection. Based on the deep learning neural network model, it can monitor wildfire smoke more robustly.

[0006] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0007] The present invention proposes a smoke detection method based on suspected smoke proposal area and deep learning, the method comprising the following steps:

[0008] S1. Collect RGB image samples to be detected;

[0009] S2. Construct a deep learning neural network model, wherein the deep learning neural network model includes a feature extraction layer backbone, a dynamic embedding layer DEN, a ROI Pooling layer and a detection head part head connected in sequence;

[0010] S3. grayscale the RGB image sample of step S1, and then perform grayscale correlation sampling on the suspected smoke proposal area to obtain n1 suspected smoke proposal areas;

[0011] S4. Convert the RGB image samples in step S1 into HSV channel image samples, expand the smoke positive samples, and then sample the suspected smoke proposal areas to obtain n2 suspected smoke proposal areas;

[0012] S5. Set a threshold value τ for the number of smoke proposal areas, and determine whether n1+n2>τ holds. If so, execute step S6; otherwise, expand the number of positive samples containing smoke, and then sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area, and execute step S6;

[0013] S6. Screen all suspected smoke proposal areas, generate bounding boxes bbox based on the screened suspected smoke proposal areas and calculate the value of the bounding box bbox, and use the value of the bounding box bbox as the spatial embedding vector input of the dynamic embedding layer DEN;

[0014] S7. Perform data enhancement operation on the RGB image samples in step S1, input the data-enhanced RGB image samples into the feature extraction layer backbone, and combine with the spatial embedding vector input into the dynamic embedding layer DEN in S6 to train the deep learning neural network model;

[0015] S8. Use the trained deep learning neural network model for smoke detection.

[0016] In this technical solution, RGB image samples to be detected are first collected, and then a deep learning neural network model is constructed. Before the image samples are input into the deep learning network model, the entire sample image is used to propose suspected smoke areas, and all suspected smoke proposal areas are screened. The candidate area covers the entire image and the scales of the candidate area are distributed from small areas to large areas, which reduces the probability of missed reports. A bounding box bbox is generated based on the screened suspected smoke proposal area and the value of the bounding box bbox is calculated. The value of the bounding box bbox is used as the spatial embedding vector input of the dynamic embedding layer DEN to improve the deep learning network model. Finally, the deep learning network model is trained. It only needs to directly input the original image into the deep learning network model for one deep feature extraction, without extracting features multiple times, and can monitor wildfire smoke more robustly.

[0017] Preferably, the feature extraction layer backbone described in step S2 adopts the backbone part of yolov5, and finally obtains ROI feature vectors of three scales, specifically including the Focus module, the first CBL module, the CSP_1 module, the second CBL module, the first CSP_3 module, the third CBL module, the second CSP_3 module, the fourth CBL module, the SPP module, the third CSP_3 module and the fifth CBL module connected in sequence; the tail end of the fifth CBL module is connected to the dynamic embedding layer DEN, the third CBL module, the fourth CBL module and the fifth CBL module amplify the resolution by upsampling, and the latter module of the third CBL module, the fourth CBL module and the fifth CBL module is cascaded with the former module for Concat feature fusion to obtain the ROI feature vector, and input it to the dynamic embedding layer DEN; the value of the bounding box bbox is input into the dynamic embedding layer DEN as the spatial embedding vector, and the tail end of the dynamic embedding layer DEN is connected to the ROI Pooling layer, and the ROI The Pooling layer pools the ROI feature vector; the detection head part head includes a first CSP_0 module, a second CSP_0 module and a convolution module connected in sequence, and outputs a prediction value corresponding to the pooled ROI feature vector.

[0018] Here, the feature extraction layer backbone adopts the CSPDarknet53 network structure, and the first CBL module, the second CBL module, the third CBL module, the fourth CBL module, and the fifth CBL module are used as basic convolution modules. The ordinary convolution layer Conv is modified to the depth-separable convolution structure DW, and the basic residual module is modified to the hourglass network structure (HourglassRes). It is only necessary to directly input the original image into the deep learning network model for one-time deep feature extraction without extracting features multiple times. The residual branches of the CSP_1 module, the first CSP_3 module, the second CSP_3 module, and the third CSP_3 module add the SE attention module.

[0019] Preferably, the process of sampling the grayscale correlation suspected smoke proposal area in step S3 is:

[0020] S31. After grayscale processing of the RGB image sample, a grayscale image is obtained, and the grayscale image is subjected to multiple threshold segmentation to obtain a plurality of threshold segmentation images, and then the threshold segmentation images are subjected to morphological corrosion and expansion processing to eliminate the fragmented areas and obtain images of different grayscale levels;

[0021] S32. Use the watershed algorithm to process each grayscale image to obtain different connected regions as suspected smoke proposal regions.

[0022] Preferably, the process of expanding the smoke positive samples and then sampling the suspected smoke proposal area in step S4 is as follows:

[0023] S41. After the RGB image samples are converted into HSV channel image samples, the H channel and the S channel are separated to obtain an H channel image and an S channel image, and the H channel image and the S channel image are pixel-inverted to obtain an HR image and an SR image;

[0024] S42. Perform pixel-by-pixel bit AND operation on the H channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area to obtain the HS image and the HSR image; perform pixel-by-pixel bit AND operation on the HR channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area to obtain the HRS image and the HRSR image;

[0025] S43. Perform adaptive threshold segmentation on the HS image, HSR image, HRS image, and HRSR image, calculate the four-connected components of all non-zero pixels in the image after adaptive threshold segmentation, and thus extract the connected domain as the suspected smoke proposal area.

[0026] Here, when grayscale correlation suspected smoke proposal area sampling is performed in step S3, areas with similar grayscale in the image are extracted as proposal areas. The union of these areas covers the entire image. The purpose is to perform grayscale correlation sampling on all positions of the image to prevent sample omission. These proposal areas contain smoke, but the vast majority of proposal areas do not contain smoke. Negative samples account for the majority, and when the smoke area is large, a complete smoke may be divided into multiple small proposal areas. The HSV channel of the image is used to expand the smoke positive samples to expand the smoke positive samples.

[0027] Preferably, the method for expanding the number of positive samples containing smoke in step S5 is a dark channel prior method, and a dark channel image is obtained by using the dark channel prior method, and the expression is:

[0028]

[0029] Among them, J dark represents the dark channel image, J channel represents one of the RGB channels, Ω(x) represents the domain of pixel x; J dark Adaptive threshold segmentation is performed, and then the four-connected components are extracted to sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area.

[0030] Preferably, the process of screening all suspected smoke proposed areas in step S6 is:

[0031] S61. Perform local binary pattern processing on the grayscale image obtained after grayscale processing of the RGB image sample using the LBP algorithm to obtain a texture image;

[0032] S62. Perform a one-dimensional discrete wavelet transform on each row of the texture image to obtain a low-frequency component L and a high-frequency component H of the texture image in the horizontal direction; perform a one-dimensional discrete wavelet transform on L in the vertical direction to obtain a low-frequency component LL and a high-frequency component LH of L in the vertical direction; perform a one-dimensional discrete wavelet transform on H in the vertical direction to obtain a low-frequency component HL and a high-frequency component HH of H in the vertical direction;

[0033] S63. Take the square sum of LH, HL and HH to obtain the local wavelet energy map E, which is expressed as:

[0034]

[0035] Among them, i and j represent the pixel value of the i-th row and the pixel value of the i-th row of the texture image respectively;

[0036] S64. Generate a mask using all the suspected smoke proposal areas, and use the mask to cut out the corresponding area on the wavelet energy map E. The cutout is based on the expression:

[0037]

[0038] Among them, E a (i, j) represents the intercepted area;

[0039] S65. Calculate the regional average energy P, sort the suspected smoke proposed regions from small to large according to the value of the regional average energy P, and select the required number of suspected smoke proposed regions; the regional average energy P formula is:

[0040]

[0041] Among them, Area Ψ(x) represents the area of ​​the suspected smoke proposal region; Ψ(x) refers to the suspected smoke proposal region in the wavelet energy map E.

[0042] Preferably, the process of performing local binary pattern processing on the grayscale image obtained after grayscale processing of the RGB picture sample using the LBP algorithm in step S61 is:

[0043] In an area with a radius of R, the central pixel value of the area is used as the threshold, and the grayscale values ​​of the adjacent N pixels are compared with the central pixel value of the area. If it is greater than the central pixel value, the pixel value is set to 1, otherwise, the pixel value is set to 0, and then the N pixel values ​​are integrated into binary form. The number represented by the binary code is the LBP value of the central pixel, and the expression is:

[0044]

[0045] Among them, (x c ,y c ) represents the center pixel of the neighborhood, i c Represents the gray value of the center pixel, i S Represents the gray value of the neighborhood pixels.

[0046] Preferably, in step S6, the process of generating a bounding box bbox based on the screened suspected smoke proposal area and calculating the value of the bounding box bbox is:

[0047] By traversing all non-zero pixels in the suspected smoke proposal area, extract the top pixel (x 1 ,y 1 ), the leftmost pixel (x 2 ,y 2 ), the rightmost pixel (x 3 ,y 3 ) and the bottom pixel (x 4 ,y 4 ), the upper left corner coordinate of the bounding box bbox (x top-left ,y top-left )=(x 1 ,y 2 )、lower right corner coordinate (x bottom-right ,y bottom-right )=(x 4 ,y 3 );

[0048] The coordinate values ​​of the upper left corner and the lower right corner of each bounding box are not less than 0, and the horizontal coordinate values ​​of the upper left corner and the lower right corner are both less than h, and the vertical coordinate values ​​of the upper left corner and the lower right corner are both less than w, where h represents the image height and w represents the image width;

[0049] Set the number threshold of bounding box bbox to N bbox , if the number of bounding boxes is less than N bbox , then copy the current bounding box bbox in the image, do random scaling, and fill it into the original bounding box bbox until the number of bounding boxes bbox reaches N bbox Otherwise, no processing is done.

[0050] Preferably, the data enhancement operations performed on the RGB image samples include: random cropping, random horizontal flipping, random scaling, random brightness adjustment, and random contrast adjustment.

[0051] Preferably, when training the deep learning neural network model in step S7, the Adam gradient descent method is used for end-to-end training, and the loss function L used by the deep learning neural network model is:

[0052] L=L conf +L SmoothL1 +L CIOU

[0053] Among them, L conf represents the confidence loss; L SmoothL1 Represents the loss of the bounding box bbox regression value, using the SmoothL1 paradigm, L CIoU Represents the CIOU loss, which is the CIOU loss between the bounding box bbox position obtained by combining the bounding box bbox regression value predicted by the deep learning neural network model and the spatial embedding vector and the true bounding box bbox; the SmoothL1 paradigm expression is:

[0054]

[0055] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0056] The present invention proposes a smoke detection method based on suspected smoke proposal area and deep learning. First, RGB image samples to be detected are collected, and then a deep learning neural network model is constructed. Before the image samples are input into the deep learning network model, the entire sample image is used to propose suspected smoke areas, and all suspected smoke proposal areas are screened. The candidate area covers the entire image and the scales of the candidate area are distributed from small areas to large areas, thereby reducing the probability of missed reports. A bounding box bbox is generated based on the screened suspected smoke proposal area and the value of the bounding box bbox is calculated. The value of the bounding box bbox is used as a spatial embedding vector input of a dynamic embedding layer DEN to improve the deep learning neural network model. Finally, the deep learning neural network model is trained. It only needs to directly input the original image into the deep learning neural network model for one deep feature extraction, without extracting multiple features. The candidate area is multi-scale, and sampling is adaptively performed at different feature layers according to the size of the candidate area to achieve a multi-scale training effect, so that wildfire smoke can be detected more robustly and wildfire smoke can be monitored more robustly. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram showing a flow chart of a smoke detection method based on suspected smoke proposal area and deep learning proposed in an embodiment of the present invention;

[0058] Figure 2 A diagram showing the overall structure of a deep learning neural network model proposed in an embodiment of the present invention;

[0059] Figure 3 A structural diagram showing an improvement of any one of the first CBL module to the fifth CBL module in the deep learning neural network model;

[0060] Figure 4 A diagram showing the structure of HourglassRes in a deep learning neural network model;

[0061] Figure 5 The improved structure diagram of the CSP_N (N can be 0, 1, 3) module in the deep learning neural network model;

[0062] Figure 6 A schematic diagram showing a Focus module, a deep separable convolution DW module, and a spatial attention mechanism module SE in a deep learning neural network model proposed in an embodiment of the present invention;

[0063] Figure 7 A schematic diagram showing the effect of the dynamic embedding layer DEN proposed in the present invention sampling CBL modules of different sizes for ROIs of different sizes;

[0064] Figure 8 Schematic diagram of input RGB picture;

[0065] Fig. 9 express Figure 8 The H, S, HR, SR, HS, HSR, HRS, and HRSR visualization effects of the input RGB image during the HSV conversion process are shown;

[0066] Fig.10 express Figure 8 Schematic diagram of a single suspected smoke area in the input RGB image shown;

[0067] Fig.11 Indicates based on Fig.10 Schematic diagram of the bounding box bbox generated by the suspected smoke area;

[0068] Fig.12 A schematic diagram showing the distribution of the bounding boxes bbox corresponding to all suspected smoke areas after screening. DETAILED DESCRIPTION

[0069] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0070] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size;

[0071] It is understandable to those skilled in the art that descriptions of certain well-known contents in the drawings may be omitted.

[0072] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent.

[0073] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0074] This embodiment proposes a smoke detection method based on suspected smoke proposal area and deep learning. The flowchart of the method is as follows: Figure 1 As shown, the overall process of the smoke detection method includes the following steps S1 to S8, specifically including:

[0075] S1. Collect RGB image samples to be detected;

[0076] S2. Build a deep learning neural network model, which includes a feature extraction layer backbone, a dynamic embedding layer DEN, a ROI Pooling layer and a detection head part head connected in sequence;

[0077] S3. grayscale the RGB image sample of step S1, and then perform grayscale correlation sampling on the suspected smoke proposal area to obtain n1 suspected smoke proposal areas;

[0078] S4. Convert the RGB image samples in step S1 into HSV channel image samples, perform smoke positive sample expansion, and then perform sampling of suspected smoke proposal regions to obtain n2 suspected smoke proposal regions;

[0079] S5. Set a threshold value τ for the number of smoke proposal areas, and determine whether n1+n2>τ holds. If so, execute step S6; otherwise, expand the number of positive samples containing smoke, and then sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area, and execute step S6;

[0080] S6. Screen all suspected smoke proposal areas, generate bounding boxes bbox based on the screened suspected smoke proposal areas and calculate the value of the bounding box bbox, and use the value of the bounding box bbox as the spatial embedding vector input of the dynamic embedding layer DEN;

[0081] S7. Perform data enhancement operation on the RGB image samples in step S1, input the data-enhanced RGB image samples into the feature extraction layer backbone, and combine with the spatial embedding vector input into the dynamic embedding layer DEN in S6 to train the deep learning neural network model;

[0082] The data augmentation operations performed on RGB image samples include: random cropping, random horizontal flipping, random scaling, random brightness adjustment, and random contrast adjustment.

[0083] S8. Use the trained deep learning neural network model for smoke detection.

[0084] In this embodiment, the overall structure of the constructed deep learning neural network model is as follows: Figure 2 As shown in the figure, the backbone of the feature extraction layer adopts the backbone part of yolov5, and finally obtains ROI feature vectors of three scales, including the Focus module, the first CBL module, the CSP_1 module, the second CBL module, the first CSP_3 module, the third CBL module, the second CSP_3 module, the fourth CBL module, the SPP module, the third CSP_3 module and the fifth CBL module connected in sequence; the tail end of the fifth CBL module is connected to the dynamic embedding layer DEN, the third CBL module, the fourth CBL module and the fifth CBL module enlarge the resolution by upsampling, and the latter module of the third CBL module, the fourth CBL module and the fifth CBL module is cascaded with the previous module for Concat feature fusion to obtain the ROI feature vector, and input it to the dynamic embedding layer DEN; the value of the bounding box bbox is input into the dynamic embedding layer DEN as the spatial embedding vector, and the tail end of the dynamic embedding layer DEN is connected to the ROI Pooling layer, and the ROI The Pooling layer pools the ROI feature vector; the detection head part head includes the first CSP_0 module, the second CSP_0 module and the convolution module connected in sequence, and outputs the prediction value corresponding to the pooled ROI feature vector.

[0085] The feature extraction layer backbone adopts the CSPDarknet53 network structure, and the CBL module in the feature extraction layer backbone is used as the basic convolution module. Figure 3 , the following improvements are made to several CBL modules (the first CBL module to the fifth CBL module) in the feature extraction layer backbone: the ordinary convolution layer Conv in each CBL module is modified to a depth-separable convolution structure DW, and the basic residual module is modified to Figure 4 The hourglass network structure (HourglassRes) shown; see Figure 5 For the CSP_N (N can be 0, 1, 3) module, the SE attention module is added to the residual branch of the CSP_N (N can be 0, 1, 3) module. The overall diagram of the Focus module, the depth-separable convolution structure DW, and the SE attention module is as follows Figure 6 shown.

[0086] The feature extraction layer backbone receives input data for feature extraction, and the dynamic embedding layer DEN receives the feature map output by the third CBL module, the fourth CBL module, and the fifth CBL module in the feature extraction layer backbone and a spatial embedding vector (the value of the bounding box bbox described in step S6). The size of the feature map is (B, W i , H i , C), the vector size is (B,N,4), B is the input data batch size, (W i ,H i ) represents the width and height of the feature map output by the i-th CBL module, C represents the number of channels, N is the number of input embedding vectors, and 4 represents the amount of data for each spatial vector. The spatial embedding vector initialization data all comes from the suspected smoke proposal area (obtained in the above steps S3 to S6). The dynamic embedding layer DEN uses the spatial embedding vector and the (W, H) information of the feature map to dynamically generate ROI data. The ROI data is output through the ROI Pooling layer, and the output structure is (B, N, 7, 7, C). (7, 7) indicates that the ROI Pooling layer uniformly pools feature maps of different sizes into 7. Among them, the specific process of the dynamic embedding layer DEN is:

[0087] First, for the convenience of description, Figure 2 The feature maps directly output by the fifth CBL and the feature maps directly output by the fifth CBL are named after the two upsampled output feature maps and the feature maps directly output by the fifth CBL. The feature map output by the fifth CBL is called a small feature map, which is generally 20×20 in size, and each pixel occupies 0.05 of the feature map; after the small feature map is combined with the feature map output by the fourth CBL through upsampling, the final output result is called a medium feature map, which is generally 40×40 in size, and each pixel occupies 0.025 of the feature map; after the medium feature map is combined with the feature map output by the third CBL through upsampling, the output result is called a large feature map, which is 80×80 in size, and each pixel occupies 0.0125 of the feature map. Then, the DEN layer calculates the ratio μ of the width and height of each suspected smoke area bbox obtained by the traditional prior method to the width and height of the input RGB sample image. Because the feature map size obtained by the subsequent ROI pooling is 7×7, for the small feature map, if μ<0.35, it means that the bbox is not enough to crop a 7×7 area on the small feature map, and the subsequent pooling will add 0 as padding. To avoid zero padding, DEN will only use the bbox with μ ≥ 0.35 as the cropping of small feature maps. Similarly, for the bbox with 0.175 ≤ μ < 0.35, DEN will crop the medium feature map, and for the bbox with μ < 0.175, DEN will crop the large feature map. Figure 7As shown in the figure, the data cropped by the bbox of the same μ on feature maps of different sizes are different. On a small-sized feature map, only one data can be cropped, while on a relatively large feature map, multiple data can be cropped, thereby avoiding excessive 0 filling during ROI pooling.

[0088] The detection head receives the output feature map of ROI pooling, and uses the first CSP_0 module and the second CSP_0 module to splice twice for output. The first CSP_0 module and the second CSP_0 module are modified as follows: the convolution kernel size of the first CSP_0 module is 7×7, and the padding is 0, so that the width and height of the output feature map are 5×5. The convolution kernel size of the second CSP_0 module is 5×5, and the padding is 0, so that the width and height of the output feature map are 3×3. Finally, it is output through an ordinary convolution layer Conv. The convolution kernel size of the convolution layer is 3×3, the padding is 0, and the number of output channels is 5, which are the bbox regression values ​​(dx, dy, dh, dw) and confidence confidence of each ROI area. In this embodiment, the output of the deep learning neural network model adopts NMS, where the confidence threshold is set to 0.6 and the IOU threshold is set to 0.7.

[0089] In this embodiment, it is assumed that the RGB image sample in step S1 is as follows: Figure 8 As shown, the grayscale processing process is a conventional operation and will not be described here in detail. The process of sampling the suspected smoke proposal area with grayscale correlation described in step S3 is as follows:

[0090] S31. After grayscale processing of the RGB image sample, a grayscale image is obtained, and the grayscale image is subjected to multiple threshold segmentation to obtain a plurality of threshold segmentation images, and then the threshold segmentation images are subjected to morphological corrosion and expansion processing to eliminate the fragmented areas and obtain images of different grayscale levels;

[0091] Here, the obtained grayscale image is segmented by multiple thresholds according to a certain interval. Here, the interval is 50, that is, pixels with grayscales of 0 to 50, 50 to 100, 100 to 150, 150 to 200, and 200 to 250 are extracted respectively. The extraction expression is:

[0092]

[0093] Through the above process, 5 threshold segmentation images are obtained. In order to remove the fragmented areas in the images, morphological erosion and dilation operations are performed on the 5 images. The erosion filter kernel size is set to 3, and it is iterated 2 times. The dilation filter kernel size is set to 3, and it is iterated 5 times. The fragmented areas are eliminated by removing areas with an area less than 100 pixels.

[0094] S32. Use the watershed algorithm to process each grayscale image to obtain different connected regions as suspected smoke proposal regions.

[0095] In this step, the watershed algorithm is used on the five images in step S31 to obtain a series of connected regions. The details of the watershed algorithm are as follows: the Euclidean distance between each non-zero pixel of the image and the nearest zero-value pixel is calculated to obtain a Euclid image, and then the local maximum value is selected in the Euclid image with a size of 20, and then the pixels of the Euclid image are inverted to obtain a TEuclid image. Finally, the coordinates of these local maximum values ​​are used as the starting point, and the watershed algorithm is applied to the TEuclid image to obtain n 1 A suspected smoke suggestion area.

[0096] In this embodiment, when the grayscale correlation suspected smoke proposal area sampling is performed in step S3, the grayscale similar areas in the image are extracted as proposal areas. The union of these areas covers the entire image. The purpose is to perform grayscale correlation sampling on all positions of the image to prevent sample omission. These proposal areas contain smoke, but the vast majority of the proposal areas do not contain smoke. Negative samples account for the majority, and when the smoke area is large, a complete smoke may be divided into multiple small proposal areas. Therefore, the HSV channel of the image is used to expand the smoke positive samples. The process of expanding the smoke positive samples and then sampling the suspected smoke proposal area in step S4 is as follows. The visualization effect of the overall input RGB image in the HSV conversion process is shown in the figure below. Fig. 9 As shown, specifically:

[0097] S41. After the RGB image samples are converted to HSV channel image samples, the H channel and S channel are separated to obtain the H channel image and the S channel image, and the H channel image and the S channel image are pixel-inverted to obtain the HR image and the SR image; one of the H, S, HR, and SR images must have a high grayscale value in the smoke area, such as Fig. 9 The first 4 pictures shown are shown.

[0098] S42. Perform pixel-by-pixel bit AND operation on the H channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area and the area less than 1000 pixels to obtain the HS image and the HSR image; perform pixel-by-pixel bit AND operation on the HR channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area and the area less than 1000 pixels to obtain the HRS image and the HRSR image, that is, Fig. 9 The last 4 images shown;

[0099] S43. Perform adaptive threshold segmentation on the HS image, HSR image, HRS image, and HRSR image, calculate the four-connected components of all non-zero pixels in the image after adaptive threshold segmentation, and extract the connected domain as the suspected smoke proposal area, a total of n 2 A proposed area containing smoke.

[0100] In this embodiment, the threshold value τ of the number of smoke proposal regions is set to 200, and it is determined whether n1+n2>200 is established. If so, step S6 is executed; otherwise, the number of positive samples containing smoke is expanded, and then the suspected smoke proposal region is sampled to obtain the expanded suspected smoke proposal region, and then step S6 is executed;

[0101] In this embodiment, the method for expanding the number of positive samples containing smoke is the dark channel prior method. The dark channel prior method is a traditional defogging method commonly used in smoke detection. After obtaining the dark channel data, the transmittance is calculated, and then the fog is subtracted from the original image according to the transmittance. In the embodiment of the present invention, the obtained fog information is used to generate a proposed area. If the number is less than 150, the dark channel prior method is used to obtain a dark channel image. The expression is:

[0102]

[0103] Among them, J dark represents the dark channel image, J channel represents one of the RGB channels, Ω(x) represents the domain of pixel x; J dark Adaptive threshold segmentation is performed, and then the four-connected components are extracted to sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area, a total of n3. All suspected smoke proposal areas are essentially binary images. Fig.10 express Figure 8 The schematic diagram of a single suspected smoke region in the input RGB image shown is a visualization result of a suspected smoke proposal region.

[0104] In this embodiment, the process of screening all the suspected smoke proposed areas in step S6 is as follows:

[0105] S61. Perform local binary pattern processing on the grayscale image obtained after grayscale processing of the RGB image sample using the LBP algorithm to obtain a texture image;

[0106] The process of local binary pattern processing of the grayscale image obtained after grayscale processing of RGB image samples using the LBP algorithm is as follows:

[0107] In an area with a radius of R, the central pixel value of the area is used as the threshold, and the grayscale values ​​of the adjacent N pixels are compared with the central pixel value of the area. If it is greater than the central pixel value, the pixel value is set to 1, otherwise, the pixel value is set to 0, and then the N pixel values ​​are integrated into binary form. The number represented by the binary code is the LBP value of the central pixel, and the expression is:

[0108]

[0109] Among them, (x c ,y c ) represents the center pixel of the neighborhood, i c Represents the gray value of the center pixel, i S Represents the grayscale value of the neighborhood pixel. In this embodiment, R is 2 and N is 16.

[0110] S62. Perform a one-dimensional discrete wavelet transform on each row of the texture image to obtain a low-frequency component L and a high-frequency component H of the texture image in the horizontal direction; perform a one-dimensional discrete wavelet transform on L in the vertical direction to obtain a low-frequency component LL and a high-frequency component LH of L in the vertical direction; perform a one-dimensional discrete wavelet transform on H in the vertical direction to obtain a low-frequency component HL and a high-frequency component HH of H in the vertical direction;

[0111] S63. Take the square sum of LH, HL and HH to obtain the local wavelet energy map E, which is expressed as:

[0112]

[0113] Among them, i and j represent the pixel value of the i-th row and the pixel value of the i-th row of the texture image respectively;

[0114] S64. Generate a mask using all the suspected smoke proposal areas, that is, convert the binary image into the corresponding binary mask image, set all non-0 pixels to 1, and use the mask to cut out the corresponding area on the wavelet energy map E. The cutout is based on the expression:

[0115]

[0116] Among them, E a (i, j) represents the intercepted area;

[0117] S65. Calculate the regional average energy P, sort the suspected smoke proposed regions from small to large according to the value of the regional average energy P, and select the required number of suspected smoke proposed regions; the regional average energy P formula is:

[0118]

[0119] Among them, Area Ψ(x)represents the area of ​​the suspected smoke proposal region; Ψ(x) refers to the suspected smoke proposal region in the wavelet energy map E.

[0120] The sorting is just to sort the suspected smoke areas based on wavelet energy. The result of the sorting is still a binary image. After sorting, the first 150 suspected smoke proposal areas are selected, and the bounding box bbox is generated based on the screened suspected smoke proposal areas and the value of the bounding box bbox is calculated as the input of the dynamic embedding layer DEN.

[0121] The process of generating the bounding box bbox based on the screened suspected smoke proposal area and calculating the value of the bounding box bbox is:

[0122] By traversing all non-zero pixels in the suspected smoke proposal area, extract the top pixel (x 1 ,y 1 ), the leftmost pixel (x 2 ,y 2 ), the rightmost pixel (x 3 ,y 3 ) and the bottom pixel (x 4 ,y 4 ), the upper left corner coordinate of the bounding box bbox (x top-left ,y top-left )=(x 1 ,y 2 )、lower right corner coordinate (x bottom-right ,y bottom-right )=(x 4 ,y 3 );

[0123] The coordinate values ​​of the upper left corner and the lower right corner of each bounding box are not less than 0, and the upper left corner horizontal coordinate value and the lower right corner horizontal coordinate value are both less than h, and the upper left corner vertical coordinate value and the lower right corner vertical coordinate value are both less than w, where h represents the image height and w represents the image width; let the number threshold of bounding box bbox be N bbox In this embodiment, 150 is taken. If the number of bounding boxes is less than N bbox , then copy the current bounding box bbox in the image, do random scaling, and fill it into the original bounding box bbox until the number of bounding boxes bbox reaches N bbox Otherwise, no processing is done. Figure 8 The schematic diagram of the bbox generated by the suspected smoke area is as follows Fig.11 As shown, the final effect is Fig.12 As shown in the figure, it can be found that all bounding boxes are no longer randomly distributed. Most of the proposal boxes are concentrated in the suspected smoke area. The number of positive samples is very large, which is conducive to the subsequent training of the deep learning neural network model.

[0124] Before training the deep learning neural network model, the deep learning neural network model is explained again. The deep learning neural network model receives 640×640×3 input image data. The Foucus module is essentially a slicing operation, which samples the image of each channel every other pixel, and then cascades it in the channel dimension to obtain data of 320×320×12 size, and then proceeds to the subsequent stages, including CBL, CSP_N, and SPP modules. In the forward propagation of data, there is no pooling operation, and downsampling only occurs in the CBL module. Although CSP_N contains the CBL module, it does not perform downsampling. The SPP module is a cascade output process using the maximum pooling operations of different sizes implemented in SPPnet. CBL is a basic convolution module, which consists of a convolution operation, a batch normalization operation, and a SiLU activation function. The convolution operation is implemented by a depth-separable convolution DW, which consists of a separation convolution stage and a point-by-point convolution stage. The difference between it and ordinary convolution is that in the separation convolution stage, the convolution kernel channel dimension is always 1, and each channel data of the feature map occupies a convolution kernel alone. The output data of this stage will keep the channel unchanged, followed by a point-by-point convolution stage, which is an ordinary convolution operation with a convolution kernel size of 1. Assume that the input data size is C 1 ×H×W, the convolution kernel size is 3, and the number of output data channels is C 2 , then the total number of parameters of ordinary convolution is C 1 ×3×3×C 2 , and the number of separable convolution parameters is C 1 ×3×3+1×1×C 1 ×C 2 , the number of parameters is significantly reduced.

[0125] The CSP_N module consists of a basic convolution module CBL, N HourglassRes, SE, ordinary convolution Conv, batch normalization BN, and activation function SiLU. HourglassRes is an improvement on the residual block Res based on the hourglass structure in Mobilenet to reduce the amount of data calculation. Each data forward propagation stage is a CBL process, except that the number of channels in the feature map is compressed and then expanded, and the final output channel is consistent with the number of input data channels. The SE module draws on the spatial attention mechanism module in SEnet and integrates it into the CSP_N module. The SE module is a typical residual structure. Its residual branch first compresses the data by global pooling, compressing C×H×W data into C×1×1 data. This is followed by the activation stage, in which the C×1×1 data is output through a fully connected neural network, activated by the ReLU function, and then output through a fully connected neural network and activated by the sigmoid function. Finally, the residual branch and the main branch are merged through the Scale operation. The Scale operation is actually a multiplication operation of the C×H×W input data and the C×1×1 residual output data.

[0126] Because the third CBL module to the fifth CBL module perform downsampling operations, feature maps of different scales will be output. After the dynamic embedding layer DEN calculates the proposed area through the spatial vector, it will determine the size of each proposed area and dynamically embed the data vector of each proposed area into a specific feature map to achieve multi-scale feature sampling. The downsampling operation will increase the receptive field of each pixel block but reduce local details. In the feature extraction layer backbone of the present invention, the more backward the CBL is, the more global features its output feature map contains but the fewer local details. When the bbox of the smoke proposal area in the input image is particularly small, then after cropping on the last CBL, very small data of even C×1×1 will be obtained. In addition to the shortcomings of insufficient local detail features described above, this kind of data also causes the subsequent ROIPooling stage to fill the border with a large number of 0s to output a C×7×7 ROI feature map. To avoid the above situation, the dynamic embedding layer DEN assigns proposals with too small sizes to the first 1 or 2 downsampling modules. The specific implementation method is as follows: the normalized spatial vector data is combined with the feature map size output by the last CBL to calculate the bbox coordinates and width and height of the proposed area. If the bbox boundary height and width are both greater than 7, they are directly output; if the bbox boundary height or width is less than 7 but greater than 3.5, the normalized spatial vector data is combined with the feature map size output by the second to last CBL to recalculate the bbox coordinates and width and height of the proposed area, and the feature map is directly output after cropping; if the bbox boundary height or width is less than 3.5, the normalized spatial vector data is combined with the feature map size output by the fifth CBL module to recalculate the bbox coordinates and width and height of the proposed area, and the feature map is directly output after cropping. The effect diagram is as follows Figure 7 Before receiving the third to fifth CBL modules, the dynamic embedding layer DEN can upsample the data of the latter module to enlarge the resolution and concatenate it with the data of the previous module to achieve the purpose of multi-scale feature fusion. The structure is as follows Figure 2 shown.

[0127] The ROLPooling layer pools the 150 ROI feature vectors into a size of C×7×7.

[0128] The first CSP_0 module or the first CSP_0 module of the detection head part head refers to the CSP module that enables 0 HourclassRes modules. The CBL modules in the two CSP_0 do not change the feature map size, but the Conv module convolution kernel size is set to 3×3, and padding is set to 0. Then the first CSP_0 will output 150×C×5×5, and the second CSP_0 will output 150×C×3×3. The last Conv convolution kernel size is 3×3, the output channel is 5, and the output data size is 150×5, which refers to the (dx, dy, dh, dw, confidence) prediction values ​​corresponding to 150 ROIs.

[0129] When training the deep learning neural network model in step S7, the Adam gradient descent method is used for end-to-end training, and the loss function L used by the deep learning neural network model is:

[0130] L=L conf +L SmoothL1 +L CIOU

[0131] Among them, L conf represents the confidence loss; L SmoothL1 Represents the loss of the bounding box bbox regression value, using the SmoothL1 paradigm, L CIOU Represents the CIOU loss, which is the CIOU loss between the bounding box bbox position obtained by combining the bounding box bbox regression value predicted by the deep learning neural network model and the spatial embedding vector and the true bounding box bbox; the SmoothL1 paradigm expression is:

[0132]

[0133] The predicted value output by the model will be used to calculate the loss with the GT value. The CIOU loss will combine the predicted (dx, dy, dh, dw) with the spatial embedding vector to calculate the value of the predicted box bbox (x 1 ,y 1 ,x 2 ,y 2 ), (x 1 ,y 1 ) represents the coordinate of the upper left corner of the prediction box bbox, (x 2 ,y 2 ) represents the coordinates of the lower right corner of the prediction box bbox. 1 ,y 1 ,x 2 ,y 2 )and Calculate the CIOU value, and then use 1-CIOU as the CIOU loss. confRepresents the cross entropy loss function of confidence. The model training adopts Adam gradient descent method, and the learning rate changes with epoch. The formula is:

[0134]

[0135] Among them, epoch represents the current training round, epochs represents the total rounds, Determine the growth rate, initialized to 0.2, the model output (dx, dy, dh, dw, confidence), and ROI value to calculate (x 1 ,y 1 ,x 2 ,y 2 ,confidence), and then use NMS to output the marked box. The confidence threshold of NMS is set to 0.6, and the IOU threshold is set to 0.7. NMS will first exclude the bbox with a confidence threshold less than 0.6, and then perform IOU calculations between the remaining bboxes. When the IOU of two bboxes is greater than 0.7, the confidence of the two bboxes is compared, and the bbox with a smaller confidence is excluded.

[0136] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent.

[0137] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A smoke detection method based on suspected smoke proposal area and deep learning, It is characterized in that The method comprises the following steps: S1. Collect RGB image samples to be detected; S2. Construct a deep learning neural network model, wherein the deep learning neural network model includes a feature extraction layer backbone, a dynamic embedding layer DEN, a ROIPooling layer and a detection head part head connected in sequence; S3. grayscale the RGB image sample of step S1, and then perform grayscale correlation sampling on the suspected smoke proposal area to obtain n1 suspected smoke proposal areas; S4. Convert the RGB image samples in step S1 into HSV channel image samples, perform smoke positive sample expansion, and then perform sampling of suspected smoke proposal regions to obtain n2 suspected smoke proposal regions; S5. Set a threshold value τ for the number of smoke proposal areas, and determine whether n1+n2>τ holds. If so, execute step S6; otherwise, expand the number of positive samples containing smoke, and then sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area, and execute step S6; S6. Screen all suspected smoke proposal areas, generate bounding boxes bbox based on the screened suspected smoke proposal areas and calculate the value of the bounding box bbox, and use the value of the bounding box bbox as the spatial embedding vector input of the dynamic embedding layer DEN; S7. Perform data enhancement operation on the RGB image samples in step S1, input the data-enhanced RGB image samples into the feature extraction layer backbone, and combine with the spatial embedding vector input into the dynamic embedding layer DEN in S6 to train the deep learning neural network model; S8. Use the trained deep learning neural network model for smoke detection.

2. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 1, It is characterized in that The feature extraction layer backbone described in step S2 adopts the backbone part of yolov5, and finally obtains ROI feature vectors of three scales, specifically including the Focus module, the first CBL module, the CSP_1 module, the second CBL module, the first CSP_3 module, the third CBL module, the second CSP_3 module, the fourth CBL module, the SPP module, the third CSP_3 module and the fifth CBL module connected in sequence; the tail end of the fifth CBL module is connected to the dynamic embedding layer DEN, and the third CBL module, the fourth CBL module and the fifth CBL module are upsampled to enlarge the resolution The third CBL module, the fourth CBL module and the fifth CBL module are cascaded with the previous module to perform Concat feature fusion to obtain a ROI feature vector, and the ROI feature vector is input into the dynamic embedding layer DEN; the value of the bounding box bbox is input into the dynamic embedding layer DEN as a spatial embedding vector, and the tail end of the dynamic embedding layer DEN is connected to the ROIPooling layer, and the ROIPooling layer pools the ROI feature vector; the detection head part head includes a first CSP_0 module, a second CSP_0 module and a convolution module connected in sequence, and outputs a prediction value corresponding to the pooled ROI feature vector.

3. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 1, It is characterized in that The process of sampling the grayscale correlation suspected smoke proposal area in step S3 is as follows: S31. After grayscale processing of the RGB image sample, a grayscale image is obtained, and the grayscale image is subjected to multiple threshold segmentation to obtain a plurality of threshold segmentation images, and then the threshold segmentation images are subjected to morphological corrosion and expansion processing to eliminate the fragmented areas and obtain images of different grayscale levels; S32. Use the watershed algorithm to process each grayscale image to obtain different connected regions as suspected smoke proposal regions.

4. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 3, It is characterized in that The process of expanding the smoke positive samples and then sampling the suspected smoke proposal area in step S4 is as follows: S41. After the RGB image samples are converted into HSV channel image samples, the H channel and the S channel are separated to obtain an H channel image and an S channel image, and the H channel image and the S channel image are pixel-inverted to obtain an HR image and an SR image; S42. Perform pixel-by-pixel bit AND operation on the H channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area to obtain the HS image and the HSR image; perform pixel-by-pixel bit AND operation on the HR channel image, the S channel image, and the SR channel image, and then perform morphological corrosion and expansion processing to eliminate the fragmented area to obtain the HRS image and the HRSR image; S43. Perform adaptive threshold segmentation on the HS image, HSR image, HRS image, and HRSR image, calculate the four-connected components of all non-zero pixels in the image after adaptive threshold segmentation, and thus extract the connected domain as the suspected smoke proposal area.

5. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 1, It is characterized in that The method for expanding the number of positive samples containing smoke in step S5 is the dark channel prior method. The dark channel image is obtained by using the dark channel prior method. The expression is: Among them, J dark represents the dark channel image, J channel represents one of the RGB channels, Ω(x) represents the neighborhood of pixel x; J dark Adaptive threshold segmentation is performed, and then the four-connected components are extracted to sample the suspected smoke proposal area to obtain the expanded suspected smoke proposal area.

6. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 1, It is characterized in that The process of screening all suspected smoke proposed areas described in step S6 is as follows: S61. Perform local binary pattern processing on the grayscale image obtained after grayscale processing of the RGB image sample using the LBP algorithm to obtain a texture image; S62. Perform a one-dimensional discrete wavelet transform on each row of the texture image to obtain a low-frequency component L and a high-frequency component H of the texture image in the horizontal direction; perform a one-dimensional discrete wavelet transform on L in the vertical direction to obtain a low-frequency component LL and a high-frequency component LH of L in the vertical direction; perform a one-dimensional discrete wavelet transform on H in the vertical direction to obtain a low-frequency component HL and a high-frequency component HH of H in the vertical direction; S63. Take the square sum of LH, HL and HH to obtain the local wavelet energy map E, which is expressed as: Among them, i and j represent the pixel value of the i-th row and the pixel value of the i-th row of the texture image respectively; S64. Generate a mask using all the suspected smoke proposal areas, and use the mask to cut out the corresponding area on the wavelet energy map E. The cutout is based on the expression: Among them, E a (i, j) represents the intercepted area; S65. Calculate the regional average energy P, sort the suspected smoke proposed regions from small to large according to the value of the regional average energy P, and select the required number of suspected smoke proposed regions; the regional average energy P formula is: Among them, Area Ψ(x) represents the area of ​​the suspected smoke proposal region; Ψ(x) refers to the suspected smoke proposal region in the wavelet energy map E.

7. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 6, It is characterized in that The process of performing local binary pattern processing on the grayscale image obtained after grayscale processing of the RGB picture sample using the LBP algorithm in step S61 is as follows: In an area with a radius of R, the central pixel value of the area is used as the threshold, and the grayscale values ​​of the adjacent N pixels are compared with the central pixel value of the area. If it is greater than the central pixel value, the pixel value is set to 1, otherwise, the pixel value is set to 0, and then the N pixel values ​​are integrated into binary form. The number represented by the binary code is the LBP value of the central pixel, and the expression is: Among them, (x c ,y c ) represents the center pixel of the neighborhood, i c Represents the gray value of the center pixel, i S Represents the gray value of the neighborhood pixels.

8. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 7, It is characterized in that In step S6, the process of generating a bounding box bbox based on the screened suspected smoke proposal area and calculating the value of the bounding box bbox is as follows: By traversing all non-zero pixels in the suspected smoke proposal area, extract the top pixel (x 1 ,y 1 ), the leftmost pixel (x 2 ,y 2 ), the rightmost pixel (x 3 ,y 3 ) and the bottom pixel (x 4 ,y 4 ), the upper left corner coordinate of the bounding box bbox (x top-left ,y top-left )=(x 1 ,y 2 ), the lower right corner coordinate (x bottom-right ,y bottom-right )=(x 4 ,y 3 ); The coordinate values ​​of the upper left corner and the lower right corner of each bounding box are not less than 0, and the horizontal coordinate values ​​of the upper left corner and the lower right corner are both less than h, and the vertical coordinate values ​​of the upper left corner and the lower right corner are both less than w, where h represents the image height and w represents the image width; Set the number threshold of bounding box bbox to N bbox , if the number of bounding boxes is less than N bbox , then copy the current bounding box bbox in the image, do random scaling, and fill it into the original bounding box bbox until the number of bounding boxes bbox reaches N bbox Otherwise, no processing is done.

9. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 8, It is characterized in that The data augmentation operations performed on RGB image samples include: random cropping, random horizontal flipping, random scaling, random brightness adjustment, and random contrast adjustment.

10. The smoke detection method based on suspected smoke proposal area and deep learning according to claim 9, It is characterized in that When training the deep learning neural network model in step S7, the Adam gradient descent method is used for end-to-end training, and the loss function L used by the deep learning neural network model is: L=L conf +L SmoothL1 +L CIOU Among them, L conf represents the confidence loss; L SmoothL1 Represents the loss of the bounding box bbox regression value, using the SmoothL1 paradigm, L CIOU Represents the CIOU loss, which is the CIOU loss between the bounding box bbox position obtained by combining the bounding box bbox regression value predicted by the deep learning neural network model and the spatial embedding vector and the true bounding box bbox; the SmoothL1 paradigm expression is:

Citation Information

Patent Citations

  • Smoke image segmentation and recognition method based on dictionary and BP neural network

    CN110415260A

  • A forest pyrotechnics detection method based on 3D convolution neural network

    CN109409256A

  • Smoke detection method combining deep convolutional neural network and visual change graph

    CN109961042A