A Smoking Behavior Detection Method Based on Deep Learning

Through semantic segmentation network and timing analysis based on deep learning, the problems of high cost, high false alarm rate and low accuracy of smoking behavior detection in the prior art are solved, and high-precision smoking behavior recognition in complex scenarios are realized.

CN119942653BActive Publication Date: 2025-07-11SICHUAN JISU POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510423869.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing smoking behavior detection methods have problems in public places with high cost, high false alarm rate, low detection accuracy, and poor adaptability to complex backgrounds. Especially when light changes, occlusion or smoke is not obvious, it is difficult to accurately identify smoking behavior.

Method used

Using a deep learning-based method, combined with the Segformer encoder, multi-channel attention module (MCA) and dynamic feature fusion module (DFF), the monitoring images are segmented by human body and smoke through a semantic segmentation network, and timing analysis is performed using continuous frames of the video stream to dynamically determine smoking behavior.

Benefits of technology

It improves the ability to extract key features of smoke and human areas, suppresses background noise interference, enhances the accuracy of segmentation and the accuracy of smoking behavior detection in complex scenarios, and reduces the rate of false detection and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942653B_ABST
    Figure CN119942653B_ABST
Patent Text Reader

Abstract

The present invention discloses a smoking behavior detection method based on deep learning, belonging to the technical field of image data processing, and comprising the steps of: obtaining a dataset for semantic segmentation; constructing a semantic segmentation network, including a Segformer encoder, an MCA, and a DFF; training the semantic segmentation network with the dataset to obtain a human body and smoke segmentation model; and detecting and judging smoking behavior. The present invention improves the semantic segmentation model, enhances the key feature extraction ability for smoke and human body regions, and simultaneously suppresses the interference of background noise, can enhance the ability to capture smoke edge details and global semantic information, and improves the segmentation accuracy. By performing temporal analysis in combination with consecutive frames of a video stream, the occurrence and duration of smoking behavior are dynamically determined through the change of the smoke region and the overlapping situation of the human body region. This method improves the accuracy and robustness of detecting smoking behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular, to a smoking behavior detection method based on deep learning. Background Art

[0002] With the increasing social concern about health and environmental issues, the detection of smoking behavior in public places plays a crucial role in maintaining public health, improving environmental quality, and protecting the rights and interests of non-smokers. Especially in no-smoking areas such as hospitals, schools, public transportation, and office buildings, smoking behavior not only violates relevant regulations but also may pose a hazard to the health of others and even cause potential safety hazards. Therefore, timely and accurate detection of smoking behavior is of great significance for creating a smoke-free environment, ensuring public health, and maintaining social order. Currently, the demand for smoking behavior detection is increasing, especially in crowded public places, and there is an urgent need for an efficient and accurate detection solution to assist management.

[0003] Current detection solutions for smoking behavior mainly include: (1) Manual inspection: Arranging management personnel in no-smoking areas for regular inspections to dissuade and punish detected smoking behavior. This method relies on manpower, has a high cost, and is limited by the number and energy of management personnel, making it difficult to achieve round-the-clock and full-coverage monitoring. In addition, manual inspections are prone to missing hidden smoking behavior, especially in places with large crowds or complex environments. (2) Sensor detection: Using smoke sensors to detect the concentration of smoke or harmful gases in the air to determine whether there is smoking behavior. This method requires the smoke to reach a certain concentration to trigger the sensor alarm and is easily interfered by other smoke or gases in the environment, resulting in a high false alarm rate and being unable to accurately distinguish smoking behavior from other smoke sources, lacking accurate determination of smoking behavior. (3) Traditional vision detection based on images: Capturing images of no-smoking areas through ordinary cameras and using traditional image processing techniques (such as edge detection, color segmentation, etc.) to identify smoking behavior. This method has a low cost and can detect the actions of people and the visual features of smoke, but its limitation is that traditional algorithms have poor adaptability to complex backgrounds, especially in cases of light changes, occlusion, or unclear smoke, resulting in low detection accuracy. (4) Object detection method based on deep learning: This method identifies cigarettes or smoke in the image as the target. The area of cigarettes is small, and with blind spots in monitoring or the dynamic behavior of the human body, they may not be effectively identified. And through smoke recognition, there may be interference from other smoke or gases, and due to the dynamic behavior of the human body and the diffusion characteristics of smoke, the target area usually shows irregular changes, making it inaccurate to identify smoking behavior. Summary of the Invention

[0004] The object of the present invention is to provide a deep learning-based smoking behavior detection method that solves the above problems, combines object detection based on deep learning with human behavior and smoke diffusion characteristics to accurately identify smoking behavior.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: A deep learning-based smoking behavior detection method, comprising the following steps;

[0006] S1, Obtain a dataset for semantic segmentation, where the samples in the dataset are images with object category annotations, and the object categories are human bodies and smoke;

[0007] S2, Construct a semantic segmentation network, including a Segformer encoder, MCA, and DFF;

[0008] The Segformer encoder is used to input samples and sequentially output a first feature map C1, a second feature map C2, a third feature map C3, and a fourth feature map C4;

[0009] The MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer, and a first multiplication layer;

[0010] Among them, the first convolutional splicing layer is used to perform 3×3 convolution on C1, C2, and C3 respectively and adjust them to the scale of C2, and then splice them to obtain a fifth feature map C5; C5 is sequentially subjected to average pooling, one-dimensional serialization, and 1×1 convolution operations through the average pooling layer, one-dimensional serialization layer, and first 1×1 convolutional layer to obtain a sixth feature map C6; the normalization and dimension adjustment layer is used to perform normalization processing on C6 and then adjust the dimension to the scale of C5; the 3×3 convolutional layer is used to perform 3×3 convolution operation on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of the MCA;

[0011] The DFF is used to perform feature fusion on F1 and C4 to obtain an output feature F;

[0012] S3, Train the semantic segmentation network with the dataset to obtain a human body and smoke segmentation model for identifying and segmenting human bodies and smoke in images;

[0013] S4, Smoking behavior detection and judgment, including steps S41~S44;

[0014] S41 Split the video stream of the monitoring area into single-frame monitoring images, and sequentially perform human body and smoke recognition through the human body and smoke segmentation model;

[0015] S42, When a human body is recognized in the monitoring image, mark the monitoring image as It , I t The front and rear monitoring images are respectively marked as I t-1 , I t+1 ;

[0016] S43, obtain I t-1 , I t , I t+1 The corresponding human body area, smoke area, and smoke area area in, if in the three monitoring images, the human body area and the smoke area overlap, then execute S44, otherwise do not judge;

[0017] S44, smoking behavior judgment;

[0018] Mark the smoke area areas in I t-1 , I t , I t+1 as S t-1 , S t , S t+1 ;

[0019] If S t-1 = 0, S t+1 > S t > 0, it is determined that a smoking behavior has just occurred;

[0020] If S t+1 > S t > S t-1 , it is determined that the smoking behavior is ongoing;

[0021] In other cases, do not judge.

[0022] Preferably: The data set is the Smoke - Segmentation data set.

[0023] Preferably: The method for the DFF to obtain the output feature F based on F1 and C4 is;

[0024] Perform dilated convolution with a dilation rate of 3 on F1 to obtain the first fused feature R1;

[0025] Perform dilated convolution with a dilation rate of 5 on C4 to obtain the second fused feature R2;

[0026] Perform 3×3 convolution on F1 and C4 respectively and then splice them to obtain the third fused feature R3;

[0027] Perform global average pooling, one - dimensional convolution, and activation function operations on R3 in sequence to obtain the fourth fused feature R4;

[0028] Multiply R3 and R4 to generate the fifth fused feature R5;

[0029] After splicing R1 and R2, add the result to R5, and then adjust the sample size back through a 1×1 convolution to obtain the output feature F.

[0030] Preferably, the DFF includes a first dilated convolutional layer, a second dilated convolutional layer, a second convolutional splicing layer, a global average pooling layer, a one-dimensional convolutional layer, an activation function layer, a second multiplication layer, a splicing layer, an addition layer, and a second 1×1 convolutional layer;

[0031] The first dilated convolutional layer is used to perform a dilated convolution with a dilation rate of 3 on F1 to obtain R1;

[0032] The second dilated convolutional layer is used to perform a dilated convolution with a dilation rate of 5 on C4 to obtain R2;

[0033] The second convolutional splicing layer is used to perform 3×3 convolutions on F1 and C4 respectively and then splice them to obtain R3;

[0034] The global average pooling layer is used to perform a global average pooling operation on R3;

[0035] The one-dimensional convolutional layer is used to perform a one-dimensional convolution operation on the output of the global average pooling;

[0036] The activation function layer is used to perform an activation function operation on the output of the one-dimensional convolutional layer to obtain R4;

[0037] The second multiplication layer is used to multiply R3 and R4 to generate R5;

[0038] The splicing layer is used to splice R1 and R2;

[0039] The addition layer is used to add the output of the splicing layer to R5;

[0040] The second 1×1 convolutional layer is used to adjust the output of the addition layer back to the sample size through a 1×1 convolution to obtain F.

[0041] Preferably, the Segformer encoder includes four Transformer blocks connected in sequence. When a sample is input into the Segformer encoder, the four Transformer blocks sequentially output a first feature map C1, a second feature map C2, a third feature map C3, and a fourth feature map C4.

[0042] In the present invention: Regarding the MCA (Multi-channel Attention Module), the feature maps of the first three lower-level semantics of the Segformer encoder are processed in this module. Through operations such as convolutional splicing, global average pooling, serialization, and 1×1 convolution for inter-channel operations, and then through operations such as normalization, dimension adjustment, and subsequent convolution, the output is obtained, enabling the model to analyze the importance of channels, capture the dependencies between channels, and enhance the channel features useful for segmenting the human body and smoke. For subsequent model processing, the representation ability, dynamic feature fusion ability, and background noise suppression ability are enhanced.

[0043] Regarding the DFF (Dynamic feature fusion module), it fuses the output of the MCA and the feature map of the fourth high-level semantics of the Segformer encoder. By using dilated convolution to obtain a larger receptive field, and better combining the detailed information and semantic information through weighted calculation of the fused feature map, thereby improving the ability to understand complex scenes. Finally, the trained human body and smoke segmentation model can more accurately and completely segment the smoke area for smoking behavior in complex scenes, and better handle the edge details of the smoke.

[0044] Compared with the prior art, the advantages of the present invention are as follows:

[0045] (1) Aiming at the problems of complex multi-object segmentation of the human body and smoke in surveillance images, large interference of smoke by background noise, and difficult accurate segmentation of edges, a multi-channel attention module MCA and a dynamic feature fusion module DFF are designed by using four different-scale features encoded by the Segformer. The four different-scale features are processed, and by adaptively weighting the feature channels, the key feature extraction ability for the smoke and human body regions is enhanced, while the interference of background noise is suppressed. At the same time, on the basis of the human body and smoke segmentation model segmenting the human body region, the ability to capture the edge details and global semantic information of the smoke is improved, and the segmentation accuracy of the model is increased.

[0046] (2) On the basis of accurately segmenting the human body region and the smoke region by the above human body and smoke segmentation model, temporal analysis is performed using consecutive frames of the video stream, and through the change of the smoke region and the overlapping situation of the human body region, the occurrence and continuation of smoking behavior are dynamically determined. This method abandons the direct detection of the small target of the cigarette, avoiding the problems of false detection and missed detection caused by the too small size or occlusion of the cigarette target, and improving the accuracy and robustness of detecting smoking behavior. Brief Description of the Drawings

[0047] Figure 1 It is a structural diagram of the human body and smoke segmentation model;

[0048] Figure 2 It is the structural diagram of MCA;

[0049] Figure 3 It is the structural diagram of DFF. Specific implementation manners

[0050] The present invention will be further described below in conjunction with embodiments and the accompanying drawings.

[0051] Embodiment 1: Refer to Figures 1 to 3 , a smoking behavior detection method based on deep learning, including the following steps;

[0052] A smoking behavior detection method based on deep learning, including the following steps;

[0053] S1. Obtain a dataset for semantic segmentation, where the samples in the dataset are images with target category annotations, and the target categories are human bodies and smoke;

[0054] S2. Construct a semantic segmentation network, including a Segformer encoder, MCA, and DFF;

[0055] The Segformer encoder is used to input samples and sequentially output a first feature map C1, a second feature map C2, a third feature map C3, and a fourth feature map C4;

[0056] The MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer, and a first multiplication layer;

[0057] Among them, the first convolutional splicing layer is used to perform 3×3 convolution on C1, C2, and C3 respectively and adjust them to the C2 scale, and then splice them to obtain a fifth feature map C5; C5 is sequentially subjected to average pooling, one-dimensional serialization, and 1×1 convolution operations through the average pooling layer, the one-dimensional serialization layer, and the first 1×1 convolutional layer to obtain a sixth feature map C6; the normalization and dimension adjustment layer is used to perform normalization processing on C6 and then adjust the dimension to the C5 scale; the 3×3 convolutional layer is used to perform 3×3 convolution operation on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of MCA;

[0058] The DFF is used to perform feature fusion on F1 and C4 to obtain an output feature F;

[0059] S3. Train the semantic segmentation network with the dataset to obtain a human body and smoke segmentation model for identifying and segmenting human bodies and smoke in images;

[0060] S4. Smoking behavior detection and judgment, including steps S41 to S44;

[0061] S41 splits the video stream of the monitored area into single-frame monitoring images, and successively performs human body and smoke recognition through the human body and smoke segmentation model;

[0062] S42, when a human body is recognized in the monitoring image, mark the monitoring image as I t , I t Mark the monitoring images of one frame before and one frame after I as I t-1 , I t+1 ;

[0063] S43, obtain the corresponding human body area, smoke area, and smoke area area in I t-1 , I t , I t+1 If the human body area and the smoke area overlap in all three frames of monitoring images, execute S44, otherwise do not make a determination;

[0064] S44, smoking behavior judgment;

[0065] Mark the smoke area areas in I t-1 , I t , I t+1 as S t-1 , S t , S t+1 ;

[0066] If S t-1 = 0, S t+1 > S t > 0, it is determined that a smoking behavior has just occurred;

[0067] If S t+1 > S t > S t-1 , it is determined that the smoking behavior is ongoing;

[0068] In other cases, no determination is made.

[0069] In this embodiment, the data set is the Smoke-Segmentation data set.

[0070] DFF is also called a multi-channel attention module. The method for obtaining the output feature F based on F1 and C4 is as follows;

[0071] Perform dilated convolution with a dilation rate of 3 on F1 to obtain the first fusion feature R1;

[0072] Perform dilated convolution with a dilation rate of 5 on C4 to obtain the second fusion feature R2;

[0073] Perform 3×3 convolution on F1 and C4 respectively and then splice them to obtain the third fusion feature R3;

[0074] Perform global average pooling, one-dimensional convolution, and activation function operations on R3 in sequence to obtain the fourth fused feature R4;

[0075] Multiply R3 and R4 to generate the fifth fused feature R5;

[0076] Concatenate R1 and R2, then add the result to R5, and then adjust the size back to the sample size through 1×1 convolution to obtain the output feature F.

[0077] The Segformer encoder includes four consecutive Transformer blocks. When a sample is input into the Segformer encoder, the four Transformer blocks output the first feature map C1, the second feature map C2, the third feature map C3, and the fourth feature map C4 in sequence.

[0078] Example 2: Refer to Figures 1 to 3 , based on Example 1, a specific structure of DFF is given. DFF is also called a dynamic feature fusion module. The DFF includes a first dilated convolutional layer, a second dilated convolutional layer, a second convolutional concatenation layer, a global average pooling layer, a one-dimensional convolutional layer, an activation function layer, a second multiplication layer, a concatenation layer, an addition layer, and a second 1×1 convolutional layer. The first dilated convolutional layer is used to perform dilated convolution with a dilation rate of 3 on F1 to obtain R1. The second dilated convolutional layer is used to perform dilated convolution with a dilation rate of 5 on C4 to obtain R2. The second convolutional concatenation layer is used to perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain R3. The global average pooling layer is used to perform global average pooling operation on R3. The one-dimensional convolutional layer is used to perform one-dimensional convolution operation on the output of the global average pooling. The activation function layer is used to perform activation function operation on the output of the one-dimensional convolutional layer to obtain R4. The second multiplication layer is used to multiply R3 and R4 to generate R5. The concatenation layer is used to concatenate R1 and R2. The addition layer is used to add the output of the concatenation layer to R5. The second 1×1 convolutional layer is used to adjust the output of the addition layer back to the sample size through 1×1 convolution to obtain F.

[0079] Example 3: Based on Example 1, we select the Smoke-Segmentation dataset and divide the dataset into a training set and a validation set at a ratio of 8:2. Four existing semantic segmentation models and the model of the present invention are used for comparative experiments, each trained for 100 rounds, and the objective comparison indicators accuracy Accuracy and mean intersection over union mIoU are selected for the experiments. The experimental results of each model's objective comparison indicators are shown in Table 1.

[0080] Table 1. Comparison table of multi-model experimental results

[0081] ,

[0082] As can be seen from Table 1, the human body and smoke segmentation model of the present invention is superior to four existing methods in terms of accuracy and mean intersection over union. By combining this human body and smoke segmentation model with step S4 of the present invention for smoking behavior detection, smoking behavior can ultimately be accurately identified.

[0083] Example 4: In the first-floor area of a municipal first-class office building, a video lasting 10 minutes was collected, and it was manually marked that a total of 15 people had smoking behaviors. Then, the present method was used for smoking behavior detection, and a total of 14 people were detected to have smoking behaviors, and another 2 people were misdetected as having smoking behaviors. When using the marked smoking images to conduct detection through the method of image classification, 10 people were correctly detected to have smoking behaviors, and another 7 people were misdetected as having smoking behaviors. Obviously, the output result of the present method is more continuous and has stronger accuracy compared with the current image classification method.

[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A smoking behavior detection method based on deep learning, characterized in that: Including the following steps; S1. Obtain a dataset for semantic segmentation. Samples in the dataset are images with target category annotations, and the target categories are human bodies and smoke; S2. Construct a semantic segmentation network, including a Segformer encoder, a multi-channel attention module MCA, and a dynamic feature fusion module DFF; The Segformer encoder is used to input samples and sequentially output a first feature map C1, a second feature map C2, a third feature map C3, and a fourth feature map C4; The multi-channel attention module MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer, and a first multiplication layer; Among them, the first convolutional splicing layer is used to perform 3×3 convolutions on C1, C2, and C3 respectively and adjust them to the scale of C2, and then splice them to obtain a fifth feature map C5; C5 is sequentially subjected to average pooling through the average pooling layer, one-dimensional serialization through the one-dimensional serialization layer, and 1×1 convolution through the first 1×1 convolutional layer to obtain a sixth feature map C6; the normalization and dimension adjustment layer is used to perform normalization processing on C6 and then adjust the dimension to the scale of C5; the 3×3 convolutional layer is used to perform 3×3 convolution operations on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of the multi-channel attention module MCA; The dynamic feature fusion module DFF is used to fuse the features of F1 and C4 to obtain an output feature F; S3. Train the semantic segmentation network with the dataset to obtain a human body and smoke segmentation model for identifying and segmenting human bodies and smoke in images; S4. Smoking behavior detection and judgment, including steps S41 to S44; S41. Split the video stream of the monitoring area into single-frame monitoring images, and sequentially perform human body and smoke recognition through the human body and smoke segmentation model; S42, when a human body is recognized in the monitoring image, the monitoring image is marked as I t , I t The first and last surveillance images are marked as I t-1 ,I t+1 ; S43, Obtain I t-1 、I t 、I t+1 The corresponding human body area, smoke area, and smoke area area in it. If the human body area and the smoke area overlap in all three frames of surveillance images, then execute S44, otherwise do not determine; S44. Smoking behavior judgment; Label I t-1 、I t 、I t+1 The smoke area areas in are respectively labeled as S t-1 、S t 、S t+1 ; If S t-1 = 0, S t+1 > S t > 0, it is determined that a smoking behavior has just occurred; If S t+1 > S t > S t-1 , it is determined that the smoking behavior is ongoing; In other cases, no judgment is made.

2. The method for detecting smoking behavior based on deep learning according to claim 1, wherein: The dataset is the Smoke-Segmentation dataset.

3. The method for detecting smoking behavior based on deep learning according to claim 1, characterized in that: The method by which the dynamic feature fusion module DFF obtains the output feature F based on F1 and C4 is as follows; Perform dilated convolution with a dilation rate of 3 on F1 to obtain a first fusion feature R1; Perform dilated convolution with a dilation rate of 5 on C4 to obtain a second fusion feature R2; Perform 3×3 convolutions on F1 and C4 respectively and then splice them to obtain a third fusion feature R3; Perform global average pooling, one-dimensional convolution, and activation function operations on R3 in sequence to obtain a fourth fusion feature R4; Multiply R3 and R4 to generate a fifth fusion feature R5; Splice R1 and R2, then add them to R5, and then adjust them back to the sample size through 1×1 convolution to obtain the output feature F.

4. The method for detecting smoking behavior based on deep learning according to claim 3, wherein: The dynamic feature fusion module DFF includes a first dilated convolutional layer, a second dilated convolutional layer, a second convolutional splicing layer, a global average pooling layer, a one-dimensional convolutional layer, an activation function layer, a second multiplication layer, a splicing layer, an addition layer, and a second 1×1 convolutional layer; The first dilated convolutional layer is used to perform dilated convolution with a dilation rate of 3 on F1 to obtain R1; The second dilated convolutional layer is used to perform dilated convolution with a dilation rate of 5 on C4 to obtain R2; The second convolutional splicing layer is used to perform 3×3 convolution on F1 and C4 respectively and then splice them to obtain R3; The global average pooling layer is used to perform global average pooling operation on R3; The one-dimensional convolutional layer is used to perform one-dimensional convolution operation on the output of the global average pooling; The activation function layer is used to perform activation function operation on the output of the one-dimensional convolutional layer to obtain R4; The second multiplication layer is used to multiply R3 and R4 to generate R5; The splicing layer is used to splice R1 and R2; The addition layer is used to add the output of the splicing layer and R5; The second 1×1 convolutional layer is used to adjust the output of the addition layer back to the sample size through 1×1 convolution to obtain F.

5. The method for detecting smoking behavior based on deep learning according to claim 1, characterized in that: The Segformer encoder includes four Transformer blocks connected in sequence. When a sample is input into the Segformer encoder, the four Transformer blocks output the first feature map C1, the second feature map C2, the third feature map C3, and the fourth feature map C4 in sequence.

Citation Information

Patent Citations

  • Smoking smoke determination method and device, storage medium and electronic device

    CN110276310A

  • Smoking condition detection method and device, electronic equipment and storage medium

    CN118097718A