Smoking behavior detection method based on deep learning
Through a deep learning-based semantic segmentation network combining human body and smoke characteristics, dynamically analyzes video streaming frames, the existing smoking behavior detection methods are solved in the high cost, low accuracy and susceptible to environmental interference, and achieves high accuracy and robust smoking behavior detection.
Patent Information
- Application Number
- CN202510423869.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing smoking behavior detection methods have problems such as high cost, low accuracy and susceptibility to environmental interference, especially in the context of dense crowds and complex situations.
The object detection method based on deep learning is used, combining human behavior and smoke diffusion characteristics, and the human body and smoke in the image are identified and segmented through semantic segmentation networks (including Segformer encoder, MCA and DFF), and timing analysis is performed through continuous frames of the video stream to dynamically determine the occurrence and persistence of smoking behavior.
It improves the accuracy and robustness of smoking behavior detection, reduces the problems of missed detection and missed detection, can accurately segment the smoke area in complex scenarios, handles smoke edge details, and improves the model segmentation accuracy and ability to identify smoking behavior.
Smart Images

Figure CN119942653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a smoking behavior detection method based on deep learning. Background Art
[0002] As society pays more and more attention to health and environmental issues, the detection of smoking behavior in public places plays a vital role in maintaining public health, improving environmental quality, and protecting the rights and interests of non-smokers. Especially in non-smoking areas such as hospitals, schools, public transportation, and office places, smoking behavior not only violates relevant regulations, but may also harm the health of others and even cause safety hazards. Therefore, timely and accurate detection of smoking behavior is of great significance for creating a smoke-free environment, protecting public health, and maintaining social order. At present, the demand for detection of smoking behavior is increasing, especially in public places with dense traffic, and an efficient and accurate detection solution is urgently needed to assist management.
[0003] The current detection schemes for smoking behavior mainly include: (1) Manual inspection: By arranging management personnel to conduct regular inspections in no-smoking areas, any smoking behavior found will be dissuaded and punished. This method relies on manpower and is costly. It is also limited by the number and energy of management personnel, making it difficult to achieve all-weather, full-coverage monitoring. In addition, manual inspections are prone to miss hidden smoking behaviors, especially in places with large traffic or complex environments. (2) Sensor detection: Using smoke sensors to detect the concentration of smoke or harmful gases in the air to determine whether there is smoking behavior. This method requires smoke to reach a certain concentration to trigger the sensor alarm and is easily interfered by other smoke or gases in the environment, resulting in a high false alarm rate. It is unable to accurately distinguish smoking behavior from other smoke sources and lacks accurate judgment of smoking behavior. (3) Traditional image-based visual detection: Using ordinary cameras to capture images of no-smoking areas, traditional image processing techniques (such as edge detection, color segmentation, etc.) are used to identify smoking behavior. This method is low-cost and can detect the visual characteristics of human movements and smoke, but its limitation is that traditional algorithms have poor adaptability to complex backgrounds, especially in cases of light changes, occlusion or unclear smoke, resulting in low detection accuracy. (4) Target detection method based on deep learning: This method identifies cigarettes or smoke in the image as targets. Cigarettes are small in size and may not be effectively identified due to blind spots or dynamic behavior of the human body. Smoke recognition may be interfered by other smoke or gases. In addition, due to the dynamic behavior of the human body and the diffusion characteristics of smoke, the target area usually shows irregular changes, and the identification of smoking behavior is not accurate enough. Summary of the invention
[0004] The purpose of the present invention is to provide a method for smoking behavior detection based on deep learning that solves the above-mentioned problem, combining target detection based on deep learning with human behavior and smoke diffusion characteristics, so as to accurately identify smoking behavior.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows: a smoking behavior detection method based on deep learning, comprising the following steps; S1, obtaining a data set for semantic segmentation, wherein samples in the data set are images with target category annotations, and the target categories are human body and smoke; S2, construct a semantic segmentation network, including Segformer encoder, MCA and DFF; The Segformer encoder is used to input a sample and sequentially output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4; The MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer and a first multiplication layer; Among them, the first convolution splicing layer is used to perform 3×3 convolution on C1, C2, and C3 respectively and adjust them to the scale of C2, and then splice to obtain the fifth feature map C5; C5 is average pooled in turn through the average pooling layer, the one-dimensional serialization layer, and the first 1×1 convolution layer, and expanded into one-dimensional serialization and 1×1 convolution operations to obtain the sixth feature map C6; the normalization and dimension adjustment layer is used to normalize C6 and then adjust the dimension to the scale of C5; the 3×3 convolution layer is used to perform a 3×3 convolution operation on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of MCA; The DFF is used to fuse the features of F1 and C4 to obtain the output feature F; S3, uses the dataset to train the semantic segmentation network to obtain a human body and smoke segmentation model, which is used to identify and segment the human body and smoke in the image; S4, smoking behavior detection and judgment, including steps S41 to S44; S41 splits the video stream of the monitoring area into single-frame monitoring images, and performs human body and smoke recognition through human body and smoke segmentation models in turn; S42, when a human body is recognized in the monitoring image, the monitoring image is marked as I t , I t The first and last surveillance images are marked as I t-1 ,I t+1 ; S43, Get I t-1 ,I t ,I t+1The human body region, smoke region, and smoke region area corresponding to the three monitoring images are as follows: if the human body region and the smoke region overlap in the three monitoring images, then S44 is executed, otherwise no determination is made; S44, judgment of smoking behavior; Will I t-1 ,I t ,I t+1 The area of the smoke region is marked as S t-1 , S t , S t+1 ; If S t-1 =0, S t+1 >S t >0, it is judged as the smoking behavior just occurred; If S t+1 >S t >S t-1 , it is judged that the smoking behavior is continuing; Other situations are not judged.
[0006] Preferably, the data set is a Smoke-Segmentation data set.
[0007] As a preferred embodiment: the method by which the DFF obtains the output feature F based on F1 and C4 is: Perform dilated convolution on F1 with a dilation rate of 3 to obtain the first fusion feature R1; Perform a dilated convolution on C4 with a dilation rate of 5 to obtain the second fusion feature R2; Perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain the third fusion feature R3; Perform global average pooling, one-dimensional convolution and activation function operations on R3 in sequence to obtain the fourth fusion feature R4; Multiply R3 and R4 to generate the fifth fusion feature R5; R1 and R2 are concatenated and then added to R5, and then adjusted back to the sample size through 1×1 convolution to obtain the output feature F.
[0008] Preferably, the DFF includes a first dilated convolution layer, a second dilated convolution layer, a second convolution splicing layer, a global average pooling layer, a one-dimensional convolution layer, an activation function layer, a second multiplication layer, a splicing layer, an addition layer, and a second 1×1 convolution layer; The first dilated convolution layer is used to perform a dilated convolution on F1 with a dilation rate of 3 to obtain R1; The second dilated convolution layer is used to perform a dilated convolution on C4 with a dilation rate of 5 to obtain R2; The second convolutional concatenation layer is used to perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain R3; The global average pooling layer is used to perform a global average pooling operation on R3; The one-dimensional convolution layer is used to perform a one-dimensional convolution operation on the output of the global average pooling; The activation function layer is used to perform an activation function operation on the output of the one-dimensional convolution layer to obtain R4; The second multiplication layer is used to multiply R3 and R4 to generate R5; The splicing layer is used to splice R1 and R2; The addition layer is used to add the output of the concatenation layer to R5; The second 1×1 convolutional layer is used to adjust the output of the addition layer back to the sample size through a 1×1 convolution to obtain F.
[0009] Preferably, the Segformer encoder comprises four layers of Transformer blocks connected in sequence, and when a sample is input into the Segformer encoder, the four layers of Transformer blocks output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4 in sequence.
[0010] In the present invention: Regarding MCA (Multi-channel Attention Module), the module processes the feature maps of the first three lower-level semantics of the Segformer encoder, performs operations between channels through convolution splicing, global average pooling, serialization, 1×1 convolution, and other operations, and then obtains the output through normalization, dimension adjustment, and subsequent convolution operations, so that the model analyzes the importance of channels, captures the dependencies between channels, and enhances the channel features useful for segmenting human bodies and smoke. For subsequent model processing, the representation ability, dynamic feature fusion ability, and background noise suppression ability are enhanced.
[0011] Regarding DFF (Dynamic feature fusion module), the output of MCA is fused with the fourth high-level semantic feature map of the Segformer encoder. The dilated convolution is used to obtain a larger receptive field, and the fusion feature map is calculated by weight to better combine the detail information and semantic information, thereby improving the ability to understand complex scenes. The human body and smoke segmentation model finally trained can more accurately and completely segment the smoke area in complex scenes for smoking behavior, and better handle the edge details of the smoke.
[0012] Compared with the prior art, the advantages of the present invention are: (1) In order to solve the problem that the multi-target segmentation of human body and smoke in surveillance images is complex, smoke is greatly disturbed by background noise, and the edges are difficult to accurately segment, the multi-channel attention module MCA and the dynamic feature fusion module DFF are designed by using the four different scale features encoded by Segformer. The four different scale features are processed and the feature channels are adaptively weighted to enhance the key feature extraction ability of smoke and human body areas, while suppressing the interference of background noise. In addition, the human body and smoke segmentation model can improve the ability to capture smoke edge details and global semantic information on the basis of segmenting the human body area, thereby improving the segmentation accuracy of the model.
[0013] (2) Based on the accurate segmentation of the human body and smoke area by the human body and smoke segmentation model, the continuous frames of the video stream are used for time series analysis, and the occurrence and continuation of smoking behavior are dynamically determined by the changes in the smoke area and the overlap of the human body area. This method abandons the direct detection of the small target of cigarettes, avoids the problem of false detection and missed detection caused by the cigarette target being too small or blocked, and improves the accuracy and robustness of detecting smoking behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is the structural diagram of the human body and smoke segmentation model; Figure 2 It is the structural diagram of MCA; Figure 3 This is the structural diagram of DFF. DETAILED DESCRIPTION
[0015] The present invention will be further described below in conjunction with embodiments and drawings.
[0016] Example 1: See Figures 1 to 3 ,A smoking behavior detection method based on deep learning comprises the following steps; A smoking behavior detection method based on deep learning, comprising the following steps; S1, obtaining a data set for semantic segmentation, wherein samples in the data set are images with target category annotations, and the target categories are human body and smoke; S2, construct a semantic segmentation network, including Segformer encoder, MCA and DFF; The Segformer encoder is used to input a sample and sequentially output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4; The MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer and a first multiplication layer; Among them, the first convolution splicing layer is used to perform 3×3 convolution on C1, C2, and C3 respectively and adjust them to the scale of C2, and then splice to obtain the fifth feature map C5; C5 is average pooled in turn through the average pooling layer, the one-dimensional serialization layer, and the first 1×1 convolution layer, and expanded into one-dimensional serialization and 1×1 convolution operations to obtain the sixth feature map C6; the normalization and dimension adjustment layer is used to normalize C6 and then adjust the dimension to the scale of C5; the 3×3 convolution layer is used to perform a 3×3 convolution operation on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of MCA; The DFF is used to fuse the features of F1 and C4 to obtain the output feature F; S3, uses the dataset to train the semantic segmentation network to obtain a human body and smoke segmentation model, which is used to identify and segment the human body and smoke in the image; S4, smoking behavior detection and judgment, including steps S41 to S44; S41 splits the video stream of the monitoring area into single-frame monitoring images, and performs human body and smoke recognition through human body and smoke segmentation models in turn; S42, when a human body is recognized in the monitoring image, the monitoring image is marked as I t , I t The first and last surveillance images are marked as I t-1 ,I t+1 ; S43, Get I t-1 ,I t ,I t+1 The human body region, smoke region, and smoke region area corresponding to the three monitoring images are as follows: if the human body region and the smoke region overlap in the three monitoring images, then S44 is executed, otherwise no determination is made; S44, judgment of smoking behavior; Will I t-1 ,I t ,I t+1 The area of the smoke region is marked as S t-1 , S t , S t+1 ; If S t-1 =0, S t+1 >S t >0, it is judged as the smoking behavior just occurred; If S t+1 >S t >S t-1 , it is judged that the smoking behavior is continuing; Other situations are not judged.
[0017] In this embodiment, the data set is a Smoke-Segmentation data set.
[0018] DFF is also called the multi-channel attention module. The method by which this module obtains the output feature F based on F1 and C4 is as follows; Perform dilated convolution on F1 with a dilation rate of 3 to obtain the first fusion feature R1; Perform a dilated convolution on C4 with a dilation rate of 5 to obtain the second fusion feature R2; Perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain the third fusion feature R3; Perform global average pooling, one-dimensional convolution and activation function operations on R3 in sequence to obtain the fourth fusion feature R4; Multiply R3 and R4 to generate the fifth fusion feature R5; R1 and R2 are concatenated and then added to R5, and then adjusted back to the sample size through 1×1 convolution to obtain the output feature F.
[0019] The Segformer encoder includes four layers of Transformer blocks connected in sequence. When a sample is input into the Segformer encoder, the four layers of Transformer blocks output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4 in sequence.
[0020] Example 2: See Figures 1 to 3 On the basis of Example 1, a specific structure of DFF is given. DFF is also called a dynamic feature fusion module. The DFF includes a first dilated convolution layer, a second dilated convolution layer, a second convolution splicing layer, a global average pooling layer, a one-dimensional convolution layer, an activation function layer, a second multiplication layer, a splicing layer, an addition layer, and a second 1×1 convolution layer; the first dilated convolution layer is used to perform a dilated convolution on F1 with a dilation rate of 3 to obtain R1; the second dilated convolution layer is used to perform a dilated convolution on C4 with a dilation rate of 5 to obtain R2; the second convolution splicing layer is used to respectively perform a dilated convolution on F1 and C4. After 3×3 convolution, they are spliced to obtain R3; the global average pooling layer is used to perform a global average pooling operation on R3; the one-dimensional convolution layer is used to perform a one-dimensional convolution operation on the output of the global average pooling; the activation function layer is used to perform an activation function operation on the output of the one-dimensional convolution layer to obtain R4; the second multiplication layer is used to multiply R3 and R4 to generate R5; the splicing layer is used to splice R1 and R2; the addition layer is used to add the output of the splicing layer to R5; the second 1×1 convolution layer is used to adjust the output of the addition layer back to the sample size through a 1×1 convolution to obtain F.
[0021] Example 3: Based on Example 1, we selected the Smoke-Segmentation dataset and divided the dataset into a training set and a validation set at a ratio of 8:2. We used four existing semantic segmentation models and the model of the present invention, each trained for 100 rounds for comparative experiments, and selected objective comparison indicators such as accuracy and mean intersection over Union (mIoU) for experiments. The objective comparison indicators of the experimental results of each model are shown in Table 1.
[0022] Table 1. Comparison of multi-model experimental results , As can be seen from Table 1, the human body and smoke segmentation model of the present invention is superior to the four existing methods in terms of accuracy and average intersection-over-union ratio. Combining the human body and smoke segmentation model with step S4 of the present invention to perform smoking behavior detection can ultimately accurately identify smoking behavior.
[0023] Example 4: A 10-minute video was collected on the first floor of a first-level office building in a certain city, and a total of 15 people were manually marked as smoking. Then, the smoking behavior detection was performed using this method, and a total of 14 people were detected to have smoked, and 2 people were falsely detected to have smoked. By using the labeled smoking images and performing detection through image classification, 10 people were correctly detected to have smoked, and 7 people were falsely detected to have smoked. Obviously, compared with the current image classification method, the output results of this method are more continuous and accurate.
[0024] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A smoking behavior detection method based on deep learning, characterized in that: The steps include: S1, obtaining a data set for semantic segmentation, wherein samples in the data set are images with target category annotated, and the target categories are human body and smoke; S2, construct a semantic segmentation network, including Segformer encoder, MCA and DFF; The Segformer encoder is used to input a sample and sequentially output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4; The MCA includes a first convolutional splicing layer, an average pooling layer, a one-dimensional serialization layer, a first 1×1 convolutional layer, a normalization and dimension adjustment layer, a 3×3 convolutional layer and a first multiplication layer; Among them, the first convolution splicing layer is used to perform 3×3 convolution on C1, C2, and C3 respectively and adjust them to the scale of C2, and then splice to obtain the fifth feature map C5; C5 is average pooled in turn through the average pooling layer, the one-dimensional serialization layer, and the first 1×1 convolution layer, and expanded into one-dimensional serialization and 1×1 convolution operations to obtain the sixth feature map C6; the normalization and dimension adjustment layer is used to normalize C6 and then adjust the dimension to the scale of C5; the 3×3 convolution layer is used to perform a 3×3 convolution operation on C6 to generate a channel weight W; the first multiplication layer is used to multiply the channel weight W and C5 to generate the output feature F1 of MCA; The DFF is used to fuse the features of F1 and C4 to obtain the output feature F; S3, uses the dataset to train the semantic segmentation network to obtain a human body and smoke segmentation model, which is used to identify and segment the human body and smoke in the image; S4, smoking behavior detection and judgment, including steps S41 to S44; S41 splits the video stream of the monitoring area into single-frame monitoring images, and performs human body and smoke recognition through human body and smoke segmentation models in turn; S42, when a human body is recognized in the monitoring image, the monitoring image is marked as I t , I t The first and last surveillance images are marked as I t-1 ,I t+1 ; S43, Get I t-1 ,I t ,I t+1 The human body region, smoke region, and smoke region area corresponding to the three frames of monitoring images are overlapped with the human body region, then S44 is executed, otherwise no determination is made; S44, judgment of smoking behavior; Will I t-1 ,I t ,I t+1 The area of the smoke region is marked as S t-1 , S t , S t+1 ; If S t-1 =0, S t+1 >S t >0, it is judged as the smoking behavior just occurred; If S t+1 >S t >S t-1 , it is judged that the smoking behavior is continuing; Other situations are not judged.
2. The method for detecting smoking behavior based on deep learning according to claim 1, characterized in that: The data set is a Smoke-Segmentation data set.
3. The method for detecting smoking behavior based on deep learning according to claim 1, characterized in that: The method of DFF obtaining the output feature F based on F1 and C4 is: Perform dilated convolution on F1 with a dilation rate of 3 to obtain the first fusion feature R1; Perform a dilated convolution on C4 with a dilation rate of 5 to obtain the second fusion feature R2; Perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain the third fusion feature R3; Perform global average pooling, one-dimensional convolution and activation function operations on R3 in sequence to obtain the fourth fusion feature R4; Multiply R3 and R4 to generate the fifth fusion feature R5; R1 and R2 are concatenated and then added to R5, and then adjusted back to the sample size through 1×1 convolution to obtain the output feature F.
4. The method for detecting smoking behavior based on deep learning according to claim 3, characterized in that: The DFF includes a first dilated convolution layer, a second dilated convolution layer, a second convolution splicing layer, a global average pooling layer, a one-dimensional convolution layer, an activation function layer, a second multiplication layer, a splicing layer, an addition layer, and a second 1×1 convolution layer; The first dilated convolution layer is used to perform a dilated convolution on F1 with a dilation rate of 3 to obtain R1; The second dilated convolution layer is used to perform a dilated convolution on C4 with a dilation rate of 5 to obtain R2; The second convolutional concatenation layer is used to perform 3×3 convolution on F1 and C4 respectively and then concatenate them to obtain R3; The global average pooling layer is used to perform a global average pooling operation on R3; The one-dimensional convolution layer is used to perform a one-dimensional convolution operation on the output of the global average pooling; The activation function layer is used to perform an activation function operation on the output of the one-dimensional convolution layer to obtain R4; The second multiplication layer is used to multiply R3 and R4 to generate R5; The splicing layer is used to splice R1 and R2; The addition layer is used to add the output of the concatenation layer to R5; The second 1×1 convolutional layer is used to adjust the output of the addition layer back to the sample size through a 1×1 convolution to obtain F.
5. The method for detecting smoking behavior based on deep learning according to claim 1, characterized in that: The Segformer encoder includes four layers of Transformer blocks connected in sequence. When a sample is input into the Segformer encoder, the four layers of Transformer blocks output a first feature map C1, a second feature map C2, a third feature map C3 and a fourth feature map C4 in sequence.
Citation Information
Patent Citations
Smoking smoke determination method and device, storage medium and electronic device
CN110276310A
Smoking behavior detection method under video monitoring
CN116630878A
Smoking condition detection method and device, electronic equipment and storage medium
CN118097718A
Smoke detection method based on deep learning, device and storage medium
US20240037950A1
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1