A Smoking Behavior Detection Method Based on the SDVGNet Network

Through the smoking behavior detection method based on the SDVGNet network, the feature extraction and multi-scale feature fusion are used to use convolution operation and attention mechanism to perform feature extraction and multi-scale feature fusion, which solves the problems of high false alarm rate, high computational complexity and instability in the prior art, and achieves high accuracy and stable smoking behavior detection.

CN118587762BActive Publication Date: 2025-06-20GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410633363.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-06-20
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

The existing smoking behavior detection methods have problems such as high false alarm rate, high computational complexity, and unstable changes in light and angles, making it difficult to achieve high accuracy detection in various scenarios.

Method used

The smoking behavior detection method based on the SDVGNet network is adopted, and the SDVGNet network is designed by constructing an image data set containing cigarettes, and the convolution operation and attention mechanism are used to perform feature extraction and multi-scale feature fusion, reducing redundant features and improving detection accuracy.

Benefits of technology

It improves the accuracy and stability of smoking behavior detection, reduces the computational complexity, enhances the generalization ability of the network, and is suitable for detection of various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118587762B_ABST
    Figure CN118587762B_ABST
Patent Text Reader

Abstract

The present invention discloses a smoking behavior detection method based on the SDVGNet network, which relates to the technical field of behavior detection. The present invention includes the following steps: Step1: Construct an image dataset containing cigarettes, collect images containing cigarettes in various scenarios, and use the LabelImg annotation software to make the collected cigarette images into a dataset in the PASCAL VOC format; Step2: Data preprocessing and dataset division, annotate the image data containing cigarettes, label them with smoking and non-smoking labels, and then divide the annotated data into a training set, a validation set, and a test set. The present invention can capture local cross-channel information interaction of the feature map containing cigarettes. Compared with traditional networks, the present invention introduces the MRA attention mechanism, which can not only process local features but also capture global context, thereby improving the overall performance and generalization ability of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of behavior detection. Specifically, it relates to a smoking behavior detection method based on the SDVGNet network. Background Art

[0002] The research results of smoking behavior detection can be divided into detection methods based on hardware devices and wireless signals, and those based on computer vision. Among them, the detection methods based on hardware devices and wireless signals have poor adaptability and do not work well in some special scenarios. In view of these limitations, in recent years, the smoking behavior detection method based on computer vision has been widely used. At the same time, the monitoring system has gradually entered the intelligent era from the simulation era, the network era, and the high-definition era. Monitoring resources are no longer used as a local monitoring function, but are combined with computer vision to achieve intelligent monitoring. According to the characteristics of the smoke generated by smoking, using relevant image processing algorithms to extract the suspected smoke area and accurately segment and identify it can greatly improve the high false alarm rate and low detection rate of traditional smoke detectors.

[0003] Comparative Document 1: Chinese Patent Document No. 202210652952.6, publication date July 15, 2022, discloses a "kitchen smoking detection method and system for matching a neural network with an infrared image". The method includes: obtaining an infrared image and a visible light image of a target area; marking a fixed high-temperature area in the visible light image; performing position registration on the infrared image and the visible light image; extracting a high-temperature area that does not coincide with the fixed high-temperature area from the infrared image as a high-temperature candidate area; using a multi-task convolutional neural network to detect the head area in the visible light image, and deforming the head area to obtain a candidate area of cigarette position information; according to the position registration relationship between the infrared image and the visible light image, detecting whether there is an overlap between the high-temperature candidate area and the candidate area of cigarette position information. If there is an overlap, it is determined that there is a smoking behavior; otherwise, it is determined that there is no smoking behavior. The invention can significantly improve the detection speed and detection accuracy of smoking detection in the kitchen scenario.

[0004] In Comparative Document 1, it is necessary to process both the visible light image and the infrared image simultaneously. If the quality of one of the images is poor or the acquisition is incomplete, it may affect the final detection result. Steps such as multi-task convolutional neural network, head area detection, and position registration make the complexity of the entire system relatively high, and a large amount of computing resources and time are required to complete these steps, resulting in a relatively high computational complexity.

[0005] Comparative Document 2: Chinese Patent Document No. 202111402315.5, publication date March 22, 2022, discloses a "Method for Detecting and Warning Smoking Behaviors at Gas Stations Based on Deep Learning". The method includes: collecting real-time images of the target gas station through a preset camera; identifying and locating the positions of smoking personnel in the current target gas station image through a preset smoking detection algorithm; controlling a flash lamp to shine on the smoking personnel for flash illumination to warn the smoking personnel. The present invention uses a smoking detection algorithm to analyze the real-time collected gas station images to detect whether there is a smoking behavior, and when there is, it timely warns of the smoking behavior at the gas station; it also reduces the burden on gas station staff, can work 24 hours a day, ensures the safety of the gas station, and ensures zero smoking in the gas station scenario, avoiding the situation of misseeing and missing in the existing smoking detection methods for gas station scenarios.

[0006] The accuracy and stability of the algorithm in Comparative Document 2 are affected by factors such as light, angle, and occlusion, resulting in unsatisfactory detection effects.

[0007] Comparative Document 3: Chinese Patent Document No. 202110038437.4, publication date April 30, 2021, discloses a "Method and System for Identifying and Handling Smoking Personnel at Oil Production Operation Sites". The method calls a camera to take pictures of the oil production operation site to obtain environmental photos; uses an identification model to determine whether there are cigarette butts in the environmental photos; among them, the identification model is obtained through machine learning training using multiple sets of data; the multiple sets of data include the first type of data and the second type of data; each set of data in the first type of data includes: a photo containing a cigarette butt and a label indicating that the photo contains a cigarette butt; each set of data in the second type of data includes: a photo not containing a cigarette butt and a label indicating that the photo does not contain a cigarette butt; in the case where there are cigarette butts in the environmental photos, it is determined that there is a smoking behavior among the personnel at the oil production operation site, and a warning signal is issued. The present invention extracts multi-scale features of the image, obtains multi-layer feature maps, and obtains a feature pyramid, solving the problem of scale change and enabling rapid and accurate identification of cigarette butts in on-site photos.

[0008] In Comparative Document 3, there may be misjudgments when judging whether there is a smoking behavior only based on whether there are cigarette butts in the photo. For example, a cigarette butt does not necessarily mean that someone is smoking, and it may also be caused by other reasons.

[0009] Finally, the traditional method for cigarette butt recognition uses the method based on the Faster RCNN model or the YOLOv5 model for recognition. The Faster RCNN model has a high complexity and cannot perform real-time detection, and the model does not consider the different scale problems of cross-domain object detection. The YOLOv5 model is limited by data samples, and the effect of the model in actual applications cannot reach the final test results; moreover, the generalization ability of the model is not strong, and the detection effect depends on the self-made data set.

[0010] In view of this, the present invention is specifically proposed. Summary of the Invention

[0011] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a smoking behavior detection method based on the SDVGNet network, which solves the problems raised in the above background technology.

[0012] To solve the above technical problems, the basic concept of the technical solution adopted by the present invention is:

[0013] A smoking behavior detection method based on the SDVGNet network includes the following steps:

[0014] Step1: Construct an image data set containing cigarettes, collect images containing cigarettes in various scenarios, and use the LabelImg annotation software to make the collected cigarette images into a data set in the PASCAL VOC format;

[0015] Step2: Data preprocessing and data set division, annotate the image data containing cigarettes, label them with smoking and non-smoking labels, and then divide the annotated data into a training set, a validation set, and a test set;

[0016] Step3: Design the SDVGNet network. In the SDVGNet network, for the input image containing cigarettes, first perform convolution operations and feature extraction operations based on the attention mechanism to extract features to obtain a feature map containing cigarettes, and then perform multi-scale feature fusion and context enhancement operations using upsampling and spatial channel refinement fusion operations to perform multi-scale feature fusion on the feature map containing cigarettes, and transfer these features to the output layer for final regression prediction.

[0017] Step 4: Training and validation of the network. Based on the SDVGNet network designed in Step 3, train and update the parameters of each layer. First, initialize all neural network parameters and use the cross-entropy loss function to measure the difference between the model prediction results and the true labels. Then, input the test set obtained in Step 2 into the SDVGNet network, and evaluate the performance of the SDVGNet network in the smoking behavior detection task by the accuracy of the SDVGNet network in identifying the smoking behavior of the pictures in the test set. Finally, according to the verification results, adjust the hyperparameters and structure of the network.

[0018] Step 5: Apply to the smoking behavior detection task. Apply the trained SDVGNet network to the smoking behavior detection task.

[0019] Optionally, the SDVGNet network consists of a C3MRA module, a PCAK module, and an SMConv module. The PCAK module includes a channel attention module PCAM and a spatial attention module PSAM.

[0020] Optionally, the following steps are performed when designing the SDVGNet network:

[0021] Step 3.1: Capture the local cross-channel information interaction of the feature map containing cigarettes. Input the feature map containing cigarettes into the C3MRA module, and then divide it into two paths. The lower path goes through a Conv for 1×1 convolution operation, and the upper path goes through a Conv for 1×1 convolution operation and a Bottleneck to increase the receptive field and reduce the computational amount. Then, pass the results of the two paths through a Concat, and finally go through a Conv for 1×1 convolution operation and an MRA to capture the local cross-channel information interaction.

[0022] Step 3.2: Feature enhancement of the feature map containing cigarettes. The PCAK module uses the attention mechanism to perform feature enhancement on the feature map containing cigarettes. The channel attention module PCAM and the spatial attention module PSAM are in parallel. Perform the attention mechanism in the space and channels, respectively infer the channel attention weight and the spatial attention weight along the two dimensions of the channel and space, and then perform feature fusion (Concat) on these two attention weights to obtain the spatial attention weight representing the fused channel. Then, obtain the spatial attention weights representing the two fused channels through a multi-layer perceptron and activate them using the sigmoid function. Finally, multiply the spatial attention weights representing the two fused channels by the input feature map containing cigarettes respectively and add the obtained results to get the feature-enhanced feature map containing cigarettes.

[0023] Step3.3: Reduce the redundant features of the feature map containing cigarettes. Use the SMConv module to reduce the redundant features of the feature map containing cigarettes. In the SMConv module, for the input feature map F∈R containing cigarettes c×h×w It is divided into two paths. The upper path is input into the Spatial Reconstruction Unit (SRU) to obtain the spatially refined feature map F containing cigarettes 1 , and the lower path is input into the Channel Reconstruction Unit (CRU) to obtain the channel-refined feature map F containing cigarettes 2 . Then, the spatially refined feature map F containing cigarettes 1 and the channel-refined feature map F containing cigarettes 2 are concatenated using the feature fusion (Concat) operation, and then passed through a Multi-Layer Perceptron layer (MLP) to obtain the spatially and channel-refined fused feature maps P and Q of cigarettes. Finally, the spatially and channel-refined fused feature maps P and Q of cigarettes are added together to obtain the spatially and channel-refined fused feature map K of cigarettes;

[0024] Step3.4: For the input image F containing cigarettes, first perform two Conv operations to downsample it to obtain the compressed feature map F 1 ; then apply the C3ECA module for feature extraction to obtain the feature map F 2 ; then perform another Conv operation on the obtained feature map and apply the C3ECA module for feature extraction to obtain the feature map F3; then input the feature map F3 into the PCAK module to enhance its features to obtain the feature map F4; then input the feature map F4 into the SMConv module to reduce the redundant features in the feature map to obtain the feature map F5; apply the Upsample operation to the feature map F5 for upsampling to obtain the feature map F6; perform a feature fusion operation (Concat) on the feature map F 2 and the feature map F6 to obtain the feature map F7; then input F7 into the SMConv module to reduce the redundant features in the feature map to obtain the feature map F8; finally, input F8 into Conv2d for a two-dimensional convolution operation to obtain the final smoking behavior detection result.

[0025] Optionally, the steps for Step3.1 to capture the local cross-channel information interaction of the feature map containing cigarettes are as follows:

[0026] Step3.1.1: Perform a global max pooling operation on the input feature map containing cigarettes, then perform a one-dimensional convolution operation with a convolution kernel size of M, and pass through the ReLU activation function to obtain the weights ω of each channel, as shown in the following formula:

[0027] ω = R(C1D M (x)) (1)

[0028] where ω is the weight of the channel; R is the ReLU activation function; C1D represents one-dimensional convolution, M is the size of the convolution kernel, x is the input feature map containing cigarettes, and finally, the weight is multiplied by the corresponding elements of the input feature map containing cigarettes to obtain the final output feature map containing cigarettes.

[0029] Optionally, the steps for feature enhancement of the feature map containing cigarettes in Step 3.2 are as follows:

[0030] Step 3.2.1: The input of PCAM is the feature map F containing cigarettes, and the dimension is set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map respectively. Global L2 norm pooling (L2Pool) and global max pooling (MaxPool) are performed on the input feature map layer containing cigarettes to obtain two feature maps containing cigarettes of 1×1×C. Then, the two feature maps containing cigarettes obtained by average pooling and max pooling are processed using a shared fully connected layer (SharedMLP), and then the results obtained by the shared fully connected layer are added and activated using the Sigmoid activation function to obtain the channel attention weight F of the input feature map F containing cigarettes c (between 0 and 1), with a size of 1×1×C. Since each channel of the feature map layer containing cigarettes extracts a certain level of feature information, the attention of PPCAM focuses on what level of feature information in the image is more important. To effectively calculate the channel attention, a method of compressing the spatial dimension of the input feature map is adopted, and its calculation formula is as follows:

[0031] F c =σ(MLP(L2Pool(F)) + MLP(MaxPool(F))) (2)

[0032] In the formula, σ is the Sigmoid activation function, MLP is the fully connected layer, L2Pool is the L2 norm pooling, MaxPool is the max pooling, F is the input feature map containing cigarettes, and Fc is the channel attention weight of the feature map F containing cigarettes.

[0033] Step 3.2.2: The input of PSAM is the feature map F containing cigarettes, and the dimension is set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map respectively. It is passed through max pooling and L2 norm pooling to obtain two feature maps of H×W×1, and then the two feature maps containing cigarettes are concatenated through a feature map fusion (Concat) operation to obtain a feature map of H×W×2, and then it becomes a feature map containing cigarettes of H×W×1 through a 7×7 convolution, and then is activated through a sigmoid function to obtain the spatial attention weight F of the feature map F containing cigarettes s(between 0 and 1), with a size of H×W×1. Its calculation formula is as follows:

[0034] F s =σ(Cov 7×7 ([L2Pool(F)+MaxPool(F)])) (3)

[0035] In the formula, σ is the Sigmoid activation function, Cov 7×7 is a 7×7 convolution operation, L2Pool is L2-norm pooling, MaxPool is max pooling, F is the input feature map containing cigarettes, and F s is the spatial attention weight of the feature map F of cigarettes;

[0036] Step3.2.3: Multiply the channel attention weight F c of the feature map F containing cigarettes obtained by PCAM with the spatial attention weight F s of the feature map F containing cigarettes obtained by PSAM through a Concat for feature fusion to obtain the spatial attention weight F cs of the fused channel representation, with a size of (H×W + C)×1×1. Then, pass through an mlp layer to obtain two spatial attention weights of the fused channel representation and Then apply the sigmoid function to activate the spatial attention weights and of the fused channel representation;

[0037] Step3.2.4: Finally, multiply the spatial attention weights and of the fused channel representation with the input feature map F containing cigarettes respectively to obtain two feature maps with enhanced spatial attention of the fused channel representation, and then add the two feature maps to obtain the final feature-enhanced feature map F PCAK .

[0038] Optionally, the steps to reduce the redundant features of the feature map containing cigarettes in Step3.3 are as follows:

[0039] Step3.3.1: In the spatial reconstruction unit, first subtract the average value from the input feature map F containing cigarettes (and divide by the standard deviation (to standardize it, and then use the trainable parameter γ ∈ R in the normalization (GN) layer c to measure the variance of the spatial pixels of each batch and channel, and normalize the relevant weight P γ ∈ R c obtained by formula (4),

[0040]

[0041] Step3.3.2: Then, through the Sigmoid function, map the weight values of the feature map containing cigarettes reweighted by P γ to the range (0, 1), and perform gating through a threshold. Set the threshold weight to 1 to obtain the information weight P1, and set it to 0 to obtain the non-information weight P2. Finally, multiply the input feature map F containing cigarettes by P1 and P2 respectively to obtain two weighted features: the feature map containing cigarettes with a large amount of information and the feature map containing cigarettes with a small amount of information

[0042] Step3.3.3: Then perform a reconstruction operation on the feature map containing cigarettes and . First, split the feature map containing cigarettes into two feature maps with channels being 1 / 2 of the original channel and Then split the feature map containing cigarettes into two feature maps with channels being 1 / 2 of the original channel and Then perform cross reconstruction, that is, fuse the features of and to obtain F P1 ; fuse the features of and to obtain F P2 . Finally, add the cross-reconstructed feature maps containing cigarettes and to obtain the spatially refined feature map F 1 . The calculation process is shown in formula (5).

[0043]

[0044] Among them, is element-wise multiplication; is element-wise summation; ∪ is feature fusion (Concat). After applying MSU to the intermediate input feature map F containing cigarettes, not only are the features with a large amount of information separated from the features with a small amount of information, but also they are reconstructed to enhance the representative features and suppress the redundant features in the spatial dimension.

[0045] Step3.3.4: In the channel reconstruction unit (MCU), for the input feature map F∈R c×h×w , first divide the channels of F into two equal parts, each with 1 / 2C channels. Subsequently, use 1×1 convolution on the two parts to compress the channels of the feature map containing cigarettes to improve its computational efficiency, and obtain F up and F lowFor F up Efficient convolution operations GWC and PWC are respectively used, and then the obtained results are added together to obtain the feature map M containing cigarettes, so as to extract high-level representative information and reduce the computational cost; for F low The efficient convolution operation PWC is adopted, and then the obtained result is combined with F low Take the union to obtain the feature map N containing cigarettes, and then use the simplified SKNet method to adaptively merge the feature maps M and N containing cigarettes to obtain the channel-refined feature map F containing cigarettes 2 。

[0046] After adopting the above technical solutions, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described below at the same time:

[0047] 1. The present invention can capture the local cross-channel information interaction of the feature map containing cigarettes. Compared with the traditional network, the present invention introduces the MRA attention mechanism, which can not only process local features but also capture global context, thereby improving the overall performance and generalization ability of the network.

[0048] 2. The present invention can respectively infer the channel attention weight and the spatial attention weight along the two dimensions of the channel and the space, and then perform feature fusion on these two attention weights to obtain the spatial attention weight representing the fused channel, so as to perform feature enhancement to make the network focus on more important features and improve the accuracy of the SDVGNet network for detecting smoking behavior. Moreover, the present application can perform feature enhancement on the feature map containing cigarettes. Compared with the traditional network, the present invention constructs a PCAK module to perform feature enhancement using the attention mechanism, which can improve the accuracy of the network for identifying smoking behavior and requires fewer parameters.

[0049] 3. The present invention can reduce the redundant features of the feature map containing cigarettes. Compared with the traditional network, the present invention constructs an SMConv module to reduce redundant features, refines features at the spatial level and the channel level and fuses the two to reduce the redundant features of the feature map containing cigarettes.

[0050] The following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The following drawings in the description are only some embodiments. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts. In the attached

[0052] In the figure:

[0053] Figure 1 Is the flow chart of the smoking behavior detection method;

[0054] Figure 2 This is the network structure diagram of the SDVGNet of the present invention;

[0055] Figure 3 This is the structure diagram of C3MRA of the present invention;

[0056] Figure 4 This is the structure diagram of the MRA module of the present invention;

[0057] Figure 5 This is the structure diagram of the PCAK attention mechanism of the present invention;

[0058] Figure 6 This is one of the structure diagrams of PCAM of the present invention;

[0059] Figure 7 This is the second structure diagram of PSAM of the present invention;

[0060] Figure 8 This is the structure diagram of the SMConv module of the present invention;

[0061] Figure 9 This is one of the structure diagrams of MSU of the present invention;

[0062] Figure 10 This is the second structure diagram of MCU of the present invention.

[0063] It should be noted that these drawings and textual descriptions are not intended to limit the scope of the concept of the present invention in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0064] Now, the present invention will be further described in detail with reference to the accompanying drawings.

[0065] Please refer to Figures 1-10 As shown, in this embodiment, a smoking behavior detection method based on the SDVGNet network is provided, including the following steps:

[0066] Step1: Construct an image dataset containing cigarettes, collect images of cigarettes in various scenarios, and use the LabelImg annotation software to make the collected cigarette images into a dataset in the PASCAL VOC format; the acquisition method of the image data is to construct a dataset by collecting images of cigarettes taken in the foreground and background in scenarios such as kitchens, scenic spots, exhibition rooms, buildings, squares, offices, cabins, and urban streets through monitoring devices on the Internet.

[0067] Step 2: Data preprocessing and dataset partitioning. Label the image data containing cigarettes with tags for smoking and non-smoking, and then divide the labeled data into a training set, a validation set, and a test set. The ratio of the training set, validation set, and test set is 6:2:2. The network structure is as Figure 2 shown.

[0068] Step 3: Design the SDVGNet network. In the SDVGNet network, for the input image containing cigarettes, first perform convolution operations and feature extraction operations based on the attention mechanism to obtain a feature map containing cigarettes. Then, perform multi-scale feature fusion and context enhancement operations using upsampling and spatial-channel refinement fusion operations on the feature map containing cigarettes, and pass these features to the output layer for the final regression prediction. The SDVGNet network consists of a C3MRA module, a PCAK module, and a SMConv module. The PCAK module includes a channel attention module PCAM and a spatial attention module PSAM.

[0069] Step 4: Training and validation of the network. Based on the SDVGNet network designed in Step 3, train and update the parameters of each layer. First, initialize all neural network parameters and use the cross-entropy loss function to measure the difference between the model prediction results and the true labels. Then, input the test set obtained in Step 2 into the SDVGNet network, and evaluate the performance of the SDVGNet network in the smoking behavior detection task by the accuracy of the SDVGNet network in recognizing the smoking behavior of the pictures in the test set. Finally, according to the validation results, adjust the hyperparameters and structure such as the learning rate and regularization parameter of the network. Input the training set into the SDVGNet network for training. During the training process, set Batch_size to 20 and the number of iterations to 200 rounds. Use the stochastic gradient descent algorithm to complete the backpropagation during the training process, and use the Adam optimizer to adaptively adjust the learning rate to help the SDVGNet network converge quickly. Then, input the preprocessed validation set into the trained network, obtain the accuracy of the network in detecting the smoking behavior of the images containing cigarettes, and use the accuracy to evaluate the performance of the network on the validation set.

[0070] Input the test set obtained in Step 2 into the SDVGNet network, evaluate the performance of the SDVGNet network in the smoking behavior detection task by the accuracy of the SDVGNet network in recognizing the smoking behavior of the pictures in the test set, and then adjust the relevant parameters of the SDVGNet network according to the performance to improve the performance of the SDVGNet network.

[0071] Step 5: Apply to the smoking behavior detection task. Apply the trained SDVGNet network to the smoking behavior detection task.

[0072] The following steps are performed when designing the SDVGNet network:

[0073] Step3.1: Capture the local cross-channel information interaction of the feature map containing cigarettes, input the feature map containing cigarettes into the C3MRA module, and then divide it into two paths. The lower path performs a 1×1 convolution operation through a Conv, and the upper path performs a 1×1 convolution operation through a Conv and a Bottleneck to increase the receptive field and reduce the computational complexity. Then, the results of the two paths are passed through a Concat, and finally, a 1×1 convolution operation is performed through a Conv and an MRA to capture the local cross-channel information interaction; the structure of the C3MRA module is as Figure 3 shown.

[0074] Step3.1.1: Perform a global max pooling operation on the input feature map containing cigarettes, then perform a one-dimensional convolution operation with a convolution kernel size of M, and obtain the weights ω of each channel through the ReLU activation function, as shown in the following formula:

[0075] ω = R(C1D M (x)) (1)

[0076] where ω is the weight of the channel; R is the ReLU activation function; C1D represents one-dimensional convolution, M is the size of the convolution kernel, x is the input feature map containing cigarettes, and finally, the weights are multiplied by the corresponding elements of the original input feature map containing cigarettes to obtain the final output feature map containing cigarettes. The structure of the MRA module is as Figure 4 shown.

[0077] It should be noted that the role of the Bottleneck module is to reduce the number of parameters and the computational complexity. In this module, the input feature map first performs a 1×1 convolution operation to reduce the number of channels to half of the original, and then performs a 3×3 convolution operation to double the number of channels. First, dimensionality reduction is beneficial for the convolution kernel to better understand the feature information, and then dimensionality increase is beneficial for extracting more detailed features. Finally, the residual structure is used to add the obtained result to the input feature map to avoid the problem of gradient disappearance.

[0078] Example: Assume that the dimension of the feature map A containing cigarettes is 160×160×128 (height×width×channels). After inputting it into the C3MRA module, the C3MRA module first divides it into two paths. The lower path passes through a 1×1 convolution to halve the number of channels, obtaining the feature map B with a dimension of 160×160×64. The upper path first passes through a 1×1 convolution to halve the number of channels, obtaining the feature map C with a dimension of 160×160×64. Then, it passes through a Bottleneck to first reduce the dimension of the feature map and then increase the dimension, obtaining the feature map D with a dimension of 160×160×64. Then, it is added to the input feature map C to obtain the feature map E with a dimension of 160×160×64. Then, through the Concat operation, the feature map B and the feature map E are fused to obtain the feature map F with a dimension of 160×160×128. Then, the feature map F passes through a 1×1 convolution to double the number of channels, obtaining the feature map G with a dimension of 160×160×256.

[0079] Input the feature map G into the MRA module. First, perform a global max pooling operation on it to obtain the feature map H with a dimension of 1×1×256. Then, perform a one-dimensional convolution operation with a convolution kernel size of k to learn the importance between different channels, obtaining the feature map I with a dimension of 1×1×256, and then activate it through the sigmoid function. Finally, multiply the feature map I and the input feature map G channel by channel to obtain the feature map J with a dimension of 160×160×256.

[0080] Step3.2: Feature enhancement of the feature map containing cigarettes. The PCAK module uses the attention mechanism to perform feature enhancement on the feature map containing cigarettes. The channel attention module PCAM and the spatial attention module PSAM are in parallel. The attention mechanism is carried out in the space and channels, and the channel attention weight and the spatial attention weight are respectively inferred along the two dimensions of the channel and the space. Then, these two attention weights are feature fused (Concat) to obtain the spatial attention weight representing the fused channel. Then, through the multi-layer perceptron, the spatial attention weights representing the two fused channels are obtained and activated by applying the sigmoid function. Finally, the spatial attention weights representing the two fused channels are respectively multiplied by the input feature map containing cigarettes, and the obtained results are added to obtain the feature-enhanced feature map containing cigarettes; The PCAK structure is as Figure 5 shown.

[0081] Step3.2.1: The input of PCAM is the feature map F containing cigarettes, with the dimension set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map respectively. Global L2-norm pooling (L2Pool) and global max pooling (MaxPool) are performed on the input feature map layer containing cigarettes to obtain two 1×1×C feature maps containing cigarettes. Then, the two feature maps containing cigarettes obtained by average pooling and max pooling are processed using a shared fully connected layer (SharedMLP). Then, the results obtained by the shared fully connected layer are added together and activated using the Sigmoid activation function to obtain the channel attention weight F of the input feature map F containing cigarettes c (between 0 and 1), with a size of 1×1×C. Since each channel of the feature map layer containing cigarettes extracts a certain level of feature information, the attention of PPCAM focuses on what level of feature information in the image is more important. To effectively calculate channel attention, a method of compressing the spatial dimension of the input feature map is adopted, and its calculation formula is as follows:

[0082] F c =σ(MLP(L2Pool(F)) + MLP(MaxPool(F))) (2)

[0083] In the formula, σ is the Sigmoid activation function, MLP is the fully connected layer, L2Pool is the L2-norm pooling, MaxPool is the max pooling, F is the input feature map containing cigarettes, and Fc is the channel attention weight of the feature map F containing cigarettes. The PPCAM structure is as Figure 6 shown.

[0084] Step3.2.2: The input of PSAM is the feature map F containing cigarettes, with the dimension set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map respectively. It is passed through max pooling and L2-norm pooling to obtain two H×W×1 feature maps, and then the two feature maps containing cigarettes are concatenated through a feature map fusion (Concat) operation to obtain an H×W×2 feature map, and then it becomes an H×W×1 feature map containing cigarettes through a 7×7 convolution, and then is activated through a sigmoid function to obtain the spatial attention weight F of the feature map F containing cigarettes s (between 0 and 1), with a size of H×W×1. Its calculation formula is as follows:

[0085] F s =σ(Cov 7×7 ([L2Pool(F) + MaxPool(F)])) (3)

[0086] In the formula, σ is the Sigmoid activation function, Cov7×7 is a 7×7 convolution operation, L2Pool is L2 norm pooling, MaxPool is max pooling, F is the input feature map containing cigarettes, and F s is the spatial attention weight of the feature map F of cigarettes; The PSAM structure is as Figure 7 shown.

[0087] Step3.2.3: Multiply the channel attention weight Fc of the feature map F containing cigarettes obtained by PCAM and the spatial attention weight F of the feature map F containing cigarettes obtained by PSAM s through a Concat for feature fusion to obtain the spatial attention weight F of the fused channel representation cs , with a size of (H×W + C)×1×1. Then, through an mlp layer, two spatial attention weights of the fused channel representation are obtained and Then apply the sigmoid function to activate the spatial attention weights of the fused channel representation and ;

[0088] Step3.2.4: Finally, multiply the spatial attention weights of the fused channel representation and by the input feature map F containing cigarettes respectively to obtain two feature maps with enhanced spatial attention of the fused channel representation, and then add the two feature maps to obtain the finally feature-enhanced feature map F PCAK

[0089] It should be noted that: L2 norm pooling selects the L2 norm within the pooling window as the output.

[0090] Example: Assume that the dimension of the input feature map F of cigarettes to PCAK is 20×20×1024 (height×width×channels). Input it into the PCAM module, perform global L2 norm pooling (L2Pool) and global max pooling (MaxPool) on the feature map A of cigarettes respectively to obtain two 1×1×1024 feature maps of cigarettes, then process the two feature maps of cigarettes obtained by average pooling and max pooling using a shared fully connected layer (SharedMLP), then add the results obtained by the shared fully connected layer and use the Sigmoid activation function for activation to obtain the channel attention weight F of the input feature map F of cigarettes c(Between 0 and 1), with a size of 1×1×1024. Then, the feature map F containing cigarettes is input into the PSAM module. Through max pooling and L2 norm pooling, two feature maps C and D of 20×20×1 are obtained. Then, through a feature map fusion (Concat) operation, the two feature maps C and D containing cigarettes are concatenated to obtain a feature map E of H×W×2. Then, a 7×7 convolution operation is performed on the feature map E to obtain a feature map G with halved channels. Then, it is activated through a sigmoid function to obtain the spatial attention weight F of the feature map F containing cigarettes s (Between 0 and 1), with a size of 20×20×1. Then, the channel attention weight F of the feature map F containing cigarettes obtained by PCAM c and the spatial attention weight F of the feature map F containing cigarettes obtained by PSAM s undergo a feature fusion through a Concat to obtain the spatial attention weight F representing the fused channel cs (with a size of 1064×1×1). Then, through an mlp layer, two spatial attention weights representing the fused channels are obtained (with a size of 20×20×1) and (with a size of 1×1×1024). Then, the sigmoid function is applied to activate the spatial attention weights representing the fused channels and Finally, the spatial attention weights representing the fused channels and are respectively multiplied by the input feature map F containing cigarettes to obtain two feature maps with enhanced spatial attention representing the fused channels. Then, the two feature maps are added to obtain the finally feature-enhanced feature map F PCAK (with a size of 20×20×1024)

[0091] Step3.3: Reduce the redundant features of the feature map containing cigarettes. Use the SMConv module to reduce the redundant features of the feature map containing cigarettes. In the SMConv module, for the input feature map F containing cigarettes ∈ R c×h×w it is divided into two paths. The upper path is input into the spatial reconstruction unit (SRU) to obtain the spatially refined feature map F containing cigarettes 1 , and the lower path is input into the channel reconstruction unit (CRU) to obtain the channel-refined feature map F containing cigarettes 2 . Then, the spatially refined feature map F containing cigarettes 1 and the channel-refined feature map F containing cigarettes 2Perform splicing using the feature fusion (Concat) operation, and then obtain the feature maps P and Q of cigarettes with refined spatial channels through a multi-layer perceptron layer (MLP). Finally, add the feature maps P and Q of cigarettes with refined spatial channels to obtain the feature map K of cigarettes with refined spatial channels; the SMConv module structure is as follows Figure 8 shown.

[0092] Step3.3.1: In the spatial reconstruction unit, first normalize the input feature map F of cigarettes by subtracting the mean μ and dividing by the standard deviation σ, and then use the trainable parameter γ ∈ R in the group normalization (GN) layer c to measure the variance of the spatial pixels for each batch and channel, and normalize the relevant weight P γ ∈ R c obtained by formula (4),

[0093]

[0094] Step3.3.2: Then map the weight values of the feature map of cigarettes re-weighted by P γ to the range (0, 1) through the Sigmoid function, and perform gating through a threshold. Set the threshold weight to 1 to obtain the information weight P1, and set it to 0 to obtain the non-information weight P2; finally, multiply the input feature map F of cigarettes by P1 and P2 respectively to obtain two weighted features: the feature map of cigarettes with a large amount of information and the feature map of cigarettes with a small amount of information

[0095] Step3.3.3: Then perform a reconstruction operation on the feature map of cigarettes and . First, split the feature map of cigarettes into two feature maps with channels being 1 / 2 of the original channels and Then split the feature map of cigarettes into two feature maps with channels being 1 / 2 of the original channels and Then perform cross reconstruction, that is, perform feature fusion on and to obtain F P1 ; perform feature fusion on and to obtain F P2 . Finally, add the cross-reconstructed feature maps of cigarettes and to obtain the spatially refined feature map F of cigarettes 1The calculation process is shown in formula (5).

[0096]

[0097] Wherein, is element-wise multiplication; is element-wise summation; ∪ is feature fusion (Concat). After applying MSU to the feature map F containing cigarettes in the intermediate input, not only are the features with large amounts of information separated from the features with small amounts of information, but they are also reconstructed to enhance the representative features and suppress the redundant features in the spatial dimension. The MSU structure is as Figure 9 shown.

[0098] Step3.3.4: In the channel reconstruction unit (MCU), for the input feature map F ∈ R c×h×w containing cigarettes, first divide the channels of F into two equal parts, each with 1 / 2C channels. Subsequently, use 1×1 convolutions on the two parts to compress the channels of the feature map containing cigarettes to improve its computational efficiency, and obtain F up and F low respectively. Apply the efficient convolution operations GWC and PWC to F up respectively, and then add the obtained results to get the feature map M containing cigarettes to extract high-level representative information and reduce the computational cost; apply the efficient convolution operation PWC to F low , and then take the union of the obtained result and F low to get the feature map N containing cigarettes. Then use the simplified SKNet method to adaptively merge the feature maps M and N containing cigarettes to obtain the channel-refined feature map F 2 . The MCU structure is as Figure 10 shown.

[0099] It should be noted that: The reconstruction operation is to add the features with more information and the features with less information to generate features with more information and save space.

[0100] Further explanation: SKNet (Selective Kernel Network) is a deep neural network architecture for image classification and object detection tasks. Its core innovation is the introduction of selective multi-scale convolutional kernels (Selective Kernel) and a novel attention mechanism, which improves the feature extraction ability without increasing the network complexity. The design of SKNet aims to solve the problem of multi-scale information fusion, enabling the network to adapt to features of different scales. Specifically, for the input feature maps M and N, global average pooling operations are first used to combine global spatial information and channel statistical information, obtaining the pooled M1 and N1. Then, Softmax is applied to M1 and N1 to obtain the feature weight vectors β1 and β2. Finally, the output F is obtained using the feature weight vectors. 2 = β1 × M1 + β2 × N1, F 2 That is, the features refined by channel.

[0101] Step3.4: For the input image F containing cigarettes, first perform two Conv operations to downsample it to obtain the compressed feature map F 1 ; then apply the C3ECA module for feature extraction to obtain the feature map F 2 ; then perform another Conv operation on the obtained feature map and apply the C3ECA module for feature extraction to obtain the feature map F3; then input the feature map F3 into the PCAK module to enhance its features to obtain the feature map F4; then input the feature map F4 into the SMConv module to reduce the redundant features in the feature map to obtain the feature map F5; apply the Upsample operation to the feature map F5 for upsampling to obtain the feature map F6; perform a feature fusion operation (Concat) on the feature map F 2 and the feature map F6 to obtain the feature map F7; then input F7 into the SMConv module to reduce the redundant features in the feature map to obtain the feature map F8; finally, input F8 into Conv2d to perform a two-dimensional convolution operation to obtain the final smoking behavior detection result.

[0102] The present invention is not limited to the above embodiments. Anyone should know that structural changes made under the inspiration of the present invention, as long as they have the same or similar technical solutions as the present invention, fall within the protection scope of the present invention. The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.

Claims

1. A smoking behavior detection method based on SDVGNet network, characterized in that: The following steps are involved: S1: Construct an image dataset containing cigarettes, collect images containing cigarettes in various scenes, and use LabelImg annotation software to make the collected cigarette images into a dataset in PASCAL VOC format; S2: Data preprocessing and data set division: annotate the image data containing cigarettes, label them as smoking and non-smoking, and then divide the annotated data into training set, validation set and test set; S3: Design the SDVGNet network. In the SDVGNet network, for the input image containing cigarettes, convolution operation and feature extraction operation based on attention mechanism are first used to extract features to obtain feature maps containing cigarettes. Then, upsampling operation, spatial channel refinement fusion operation and context enhancement operation are used to perform multi-scale feature fusion on the feature maps containing cigarettes, and these features are passed to the output layer for the final regression prediction. The SDVGNet network is used to perform the following steps: For the input image F containing cigarettes, it is first downsampled through two convolution operations to obtain the compressed feature map F1; then the C3ECA module is applied to extract features to obtain the feature map F2; then the obtained feature map is convolved once and the C3ECA module is applied to extract features to obtain the feature map F3; then the feature map F3 is input to the PCAK module to enhance its features to obtain the feature map F4; then the feature map F4 is input to the SMConv module to reduce the redundant features in the feature map to obtain the feature map F5; the feature map F5 is upsampled to obtain the feature map F6; the feature map F2 and the feature map F6 are subjected to a feature fusion operation to obtain the feature map F7; then F7 is input to the SMConv module to reduce the redundant features in the feature map to obtain the feature map F8; finally, F8 is input to the Conv2d module to perform a two-dimensional convolution operation to obtain the final smoking behavior detection result; S4: Network training and verification: Based on the SDVGNet network designed in S3, train and update the parameters of each layer. First, initialize all neural network parameters and use the cross entropy loss function to measure the difference between the model prediction results and the true labels. Then input the test set obtained in S2 into the SDVGNet network. The performance of the SDVGNet network in the smoking behavior detection task is evaluated by the accuracy of the SDVGNet network in recognizing smoking behaviors in the test set images. Finally, adjust the network's hyperparameters and structure based on the verification results. S5: Apply the trained SDVGNet network to the smoking behavior detection task.

2. The smoking behavior detection method based on SDVGNet network according to claim 1, characterized in that: The steps performed by the SDVGNet network include: S3.1: Capture local cross-channel information interaction of feature maps containing cigarettes: Input the feature map containing cigarettes into the C3MRA module, and then divide it into two paths. The lower path passes through a Conv module for a 1×1 convolution operation, and the upper path passes through a Conv module for a 1×1 convolution operation and a Bottleneck module to increase the receptive field and reduce the amount of calculation. Then the two results pass through a Concat module for feature fusion, and finally pass through a Conv module for a 1×1 convolution operation and an MRA module to capture local cross-channel information interaction; S3.2: Feature enhancement of feature maps containing cigarettes. The PCAK module uses the attention mechanism to enhance the feature maps containing cigarettes. The channel attention module PCAM and the spatial attention module PSAM are connected in parallel to perform the attention mechanism in space and channels. The channel attention weight and spatial attention weight are inferred along the two dimensions of channel and space respectively. Then, the two attention weights are feature fused to obtain the spatial attention weight represented by the fused channel. The spatial attention weights represented by the two fused channels are obtained through a multi-layer perceptron and activated by a sigmoid function. Finally, the spatial attention weights represented by the two fused channels are multiplied by the input feature map containing cigarettes respectively, and the results are added to obtain the feature-enhanced feature map containing cigarettes. S3.3: Use the SMConv module to reduce the redundant features of the feature map containing cigarettes. In the SMConv module, for the input feature map containing cigarettes F∈R c×h×w It is divided into two paths, where c represents the number of channels of the feature map F, h represents the height of the feature map F, and w represents the width of the feature map F. The upper path is input to the spatial reconstruction unit to obtain the spatially refined feature map F containing the cigarette. 1 , the lower path is input to the channel reconstruction unit to obtain the channel refinement containing the cigarette feature map F 2 , and then refine the space to include the cigarette feature map F 1 and channel refinement contains the feature map F of cigarettes 2 The feature fusion operation is used for splicing, and then the feature maps P and Q containing cigarettes are obtained through a multi-layer perceptron layer. Finally, the feature maps P and Q are added together to obtain the feature map K containing cigarettes.

3. The smoking behavior detection method based on SDVGNet network according to claim 2 is characterized in that: S3.1 The steps for capturing the local cross-channel information interaction of the feature graph containing cigarettes are: S3.1.1: Perform a global maximum pooling operation on the input feature map containing cigarettes, and then perform a one-dimensional convolution operation with a convolution kernel size of M. After the ReLU activation function, the weight ω of each channel is obtained, as shown in the formula below: ω=R(C1D M (x)) (1) Where ω is the weight of the channel; R is the ReLU activation function; C1D represents one-dimensional convolution, M is the size of the convolution kernel, x is the input feature map containing cigarettes, and finally the weight is multiplied by the corresponding element of the original input feature map containing cigarettes to obtain the final output feature map containing cigarettes.

4. The smoking behavior detection method based on SDVGNet network according to claim 3 is characterized in that: S3.2 contains the steps of feature enhancement of the feature map of cigarettes: S3.2.1: The input of PCAM is a feature map F containing cigarettes, and the dimension is set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map, respectively. The input feature map containing cigarettes is subjected to global L2 norm pooling and global maximum pooling to obtain two 1×1×C feature maps containing cigarettes. The two feature maps containing cigarettes obtained by average pooling and maximum pooling are then processed using a shared fully connected layer. The results obtained by the shared fully connected layer are then added and activated using the Sigmoid activation function to obtain the channel attention weight F of the input feature map F containing cigarettes. c , F c The value range is between 0 and 1, and the size is 1×1×C. In order to effectively calculate the channel attention, the spatial dimension method of compressing the input feature map is adopted. The calculation formula is as follows: F c =σ(MLP(L2Pool(F))+MLP(MaxPool(F))) (2) In the formula, σ is the Sigmoid activation function, MLP is the fully connected layer, L2Pool is the L2 norm pooling, MaxPool is the maximum pooling, F is the input feature map containing cigarettes, and F c is the channel attention weight of the feature map F containing cigarettes; S3.2.2: The input of PSAM is a feature map F containing cigarettes, and the dimension is set to H×W×C; where H, W, and C are the height, width, and number of channels of the feature map, respectively. The feature map is pooled by maximum pooling and L2 norm pooling to obtain two H×W×1 feature maps, and then the two feature maps containing cigarettes are concatenated by feature map fusion to obtain a H×W×2 feature map, which is then converted into a H×W×1 feature map containing cigarettes by 7×7 convolution, and then activated by a sigmoid function to obtain the spatial attention weight F of the feature map F containing cigarettes. s , F s The value range is between 0 and 1, and the size is H×W×1; the calculation formula is as follows: F s =σ(Cov 7×7 ([L2Pool(F)+MaxPool(F)])) (3) Where σ is the Sigmoid activation function, Cov 7×7 is a 7×7 convolution operation, L2Pool is L2 norm pooling, MaxPool is maximum pooling, F is the input feature map containing cigarettes, F s is the spatial attention weight of the feature map F of the cigarette; S3.2.3: The channel attention weight F of the feature map F containing cigarettes obtained by PCAM c The spatial attention weight F of the feature map F containing cigarettes obtained by PSAM s After a Concat, feature fusion is performed to obtain the spatial attention weight F represented by the fusion channel. cs , the size is (H×W+C)×1×1; then after an MLP layer, the spatial attention weights represented by the two fusion channels are obtained and Then apply the sigmoid function to fuse the spatial attention weights represented by the channel and activation; S3.2.4: Finally, the spatial attention weights represented by the fusion channel and Multiply them with the input feature map F containing cigarettes to obtain two spatial attention-enhanced feature maps represented by the fusion channel, and then add the two feature maps to obtain the final feature-enhanced feature map F PCAK .

5. The smoking behavior detection method based on SDVGNet network according to claim 3 is characterized in that: The steps in S3.3 to reduce redundant features of feature graphs containing cigarettes are: S3.3.1: In the spatial reconstruction unit, the input feature map F containing cigarettes is first normalized by subtracting the mean and dividing by the standard deviation, and then the trainable parameter γ∈R in the normalization (GN) layer is used. c To measure the variance of spatial pixels for each batch and channel, normalize the correlation weight P γ ∈R c From formula (4), S3.3.2: Then the Sigmoid function is used to convert P γ The weight value of the re-weighted feature map containing cigarettes is mapped to the range (0, 1) and gated by the threshold. The threshold weight is set to 1 to obtain the informative weight P1, and it is set to 0 to obtain the non-informative weight P2. Finally, the input feature map F containing cigarettes is multiplied by P1 and P2 respectively to obtain two weighted features: the feature map containing cigarettes with large information content and the feature map containing cigarettes with less information S3.3.3: Then, for the feature map containing cigarettes and To perform the reconstruction operation, first convert the feature map containing cigarettes Split into two channels with a feature map of 1 / 2 of the original channel and Then the feature map containing cigarettes Split into two channels with a feature map of 1 / 2 of the original channel and Then perform cross reconstruction, that is, and Perform feature fusion to obtain F P1 ; Bundle and Perform feature fusion to obtain F P2 Finally, the cross-reconstructed feature map F containing cigarettes P1 and F P2 Add together to obtain the spatially refined feature map F containing cigarettes 1 ; The calculation process is shown in formula (5); in, is element-wise multiplication; is the sum of elements; ∪ is feature fusion; S3.3.4: In the channel reconstruction unit, for the input feature map F∈R containing cigarettes c×h×w First, the channel of F is evenly divided into two parts, each with 1 / 2C. Then, 1×1 convolution is used to compress the channel containing the feature map of cigarettes to improve its computational efficiency. up and F low ; for F up The efficient convolution operations GWC and PWC are used respectively, and the results are added to obtain the feature map M containing cigarettes to extract high-level representative information and reduce the computational cost; low Use efficient convolution operation PWC, and then combine the result with F low Take the union to get the feature map N containing cigarettes, and then use the simplified SKNet method to adaptively merge the feature maps M and N containing cigarettes to get the channel-refined feature map F containing cigarettes 2 .

Citation Information

Patent Citations

  • Oil production operation site smoker recognition processing method and system

    CN112733730A

  • Gas station smoking behavior detection alarm method based on deep learning

    CN114220000A

  • A method and system for detecting kitchen smoke by matching neural networks with infrared images.

    CN114757979B

  • Light smoking detection method and system

    CN115240118A