A smoking behavior recognition method and system
By using the deep learning model efficientDet and the hand node model to identify smoking behavior, the high cost and low accuracy of smoking monitoring in public places have been solved. This has enabled low-cost, flexible deployment of multi-target smoking identification, which is adaptable to complex environments and improves the accuracy and robustness of identification.
Patent Information
- Application Number
- CN202210799484.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-07-06
AI Technical Summary
Existing technologies for monitoring smoking in public places suffer from high human resource costs, poor equipment adaptability, and low accuracy, especially in complex environments where it is difficult to effectively identify smoking behavior.
Real-time video stream analysis is performed using the deep learning model efficientDet. By acquiring the regions of interest for the hand and mouth, the distribution of joint positions is calculated by combining the hand node model. The Pixel2Mesh algorithm is used to map the data to a 3D space, and the SpatialRelationshipNet network is used to rearrange the feature maps to determine whether the hand movements are consistent with smoking behavior.
It achieves low-cost, flexible deployment of multi-target smoking behavior recognition, adapts to complex scenarios, reduces unnecessary detection, improves recognition accuracy and robustness, and can monitor the smoking behavior of multiple individuals simultaneously.
Smart Images

Figure CN115359381B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular, the present application relates to a smoking behavior recognition method and system. BACKGROUND
[0002] Thanks to the rapid development of high-performance computing and artificial intelligence in recent years, computer vision technology has been well developed and utilized. In the field of safety production, the behavior attribute and standard production of personnel have become an increasingly important problem. With the improvement of people's health consciousness and safety awareness, more and more public places and production environments have clear regulations prohibiting smoking, and the demand for smoking detection is increasing. Schools, stations, office buildings, shopping malls, gas stations and factory workshops are key smoking areas. Using manual supervision in these relatively wide areas requires a large amount of human resource cost, and using existing surveillance cameras and artificial intelligence computer vision technology to monitor smoking can save a lot of human resources.
[0003] Therefore, in view of the problems existing in the current smoking monitoring in public scenes, a smoking behavior recognition scheme using deep learning modeling is designed. SUMMARY
[0004] In order to overcome the shortcomings of the prior art, the present application provides a smoking behavior recognition method and system to solve the above technical problems.
[0005] The technical method adopted by the present application to solve its technical problems is: a smoking behavior recognition method, the improvement lies in: comprising the following steps: S1, acquiring real-time video stream, calculating the video stream through the deep learning target detection model efficientDet, acquiring the region of interest of the personnel, the region of interest includes the position of the hand and the position of the mouth; S2, acquiring the position of the joint node in the region of interest through the hand node model, and calculating the joint node position distribution information; S3, using the joint node position distribution information, and comparing with the existing standard hand-held cigarette palm node distribution, judging whether the current palm action conforms to the palm node position distribution when smoking, if yes, identifying as smoking behavior.
[0006] In the above method, the step further comprises:
[0007] S4, comparing the hand-held cigarette hand in the target detection result with the hand holding other objects or the hand not holding objects, judging whether the current action behavior is smoking behavior.
[0008] In the above method, in the step S1, the deep learning target detection model efficientDet calculates the video stream, comprising the following steps:
[0009] S101, performing feature extraction on the video stream by a dynamic scalable convolutional neural network;
[0010] S102, fusing the extracted features in a form of recurrent skip combination;
[0011] S103, calculating the fused features to obtain target information.
[0012] In the above method, the step S2 comprises the following steps:
[0013] S21, collecting a large amount of palm data, labeling each node of the palm of the palm data using an automatic node labeling method, and mapping the obtained nodes to a three-dimensional space by using a Pixel2Mesh algorithm;
[0014] S22, calculating the spatial node data by using a SpatialRelationshipNet network, and rearranging the obtained feature map;
[0015] S23, the rearranged feature map is calculated by a full connection layer to output the final result, i.e. the joint position distribution information.
[0016] In the above method, in the step S21, the Pixel2Mesh algorithm is used to map the obtained nodes to a three-dimensional space, comprising the following steps:
[0017] S201, randomly initializing an outer ellipse of an arbitrary input picture as an initial three-dimensional feature shape, the network model has two branches, one branch is responsible for feature extraction of the input picture by using a full convolutional neural network, and the other branch is responsible for three-dimensional grid feature extraction;
[0018] S202, deforming the three-dimensional grid into a required three-dimensional model by continuously fusing and calculating the two branches.
[0019] In the above method, the step S22 comprises the following steps:
[0020] S221, converting the input picture calculation into specific dense spatial point position information;
[0021] S222, the convolutional neural network CBL structure outputs two types of feature maps, which are obtained by maximum pooling and average pooling respectively, and the two types of pooling results are fused;
[0022] S223, the fused features are sent to the next network structure, and the obtained feature map is rearranged.
[0023] The application further provides a smoking behavior recognition system, comprising a deep learning target detection model efficientDet, a hand node model and a judgment module,
[0024] The deep learning target detection model efficientDet is used for calculating the acquired real-time video stream, and acquiring a region of interest of a person, wherein the region of interest comprises a position of a hand and a position of a mouth.
[0025] The hand node model is used for acquiring a position of a joint node in the region of interest, and calculating joint node position distribution information.
[0026] The judgment module is used for comparing the joint node position distribution information with a standard hand-held cigarette palm node distribution, judging whether a palm action currently occurring conforms to a palm node position distribution when smoking, and if yes, recognizing as a smoking behavior.
[0027] The application has the beneficial effects that: the smoking behavior is recognized by using time sequence deep learning modeling, the candidate region of interest of a person is acquired through target detection, thereby reducing unnecessary detection, and the calculation resources can be better utilized; and the deep learning target detection method is beneficial to removing interference of other non-specific targets, and has stronger robustness and is suitable for various complex scenes; meanwhile, the target detection can acquire multiple persons of interest at the same time, and whether different individuals have smoking behaviors can be recognized at the same time. BRIEF DESCRIPTION OF DRAWINGS
[0028] FIG. 1 is a flowchart of a smoking behavior recognition method according to the application. Figure 1 FIG. 1 is a flowchart of a smoking behavior recognition method according to the application.
[0029] FIG. 2 is a structure diagram of a SpatialRelationshipNet network in the smoking behavior recognition method according to the application. Figure 2 FIG. 2 is a structure diagram of a SpatialRelationshipNet network in the smoking behavior recognition method according to the application. DETAILED DESCRIPTION
[0030] The application will be further described below in combination with the drawings and examples.
[0031] The concept, specific structure and generated technical effects of the present application will be described clearly and completely in combination with the embodiments and drawings, so as to fully understand the purposes, features and effects of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments, and other embodiments obtained by those skilled in the art on the basis of the embodiments of the present application without creative labor are within the scope of protection of the present application. In addition, all the coupling / connection relationships involved in the patent do not mean that the components are directly connected, but that a better coupling structure can be composed by adding or reducing coupling accessories according to the specific implementation. The technical features in the present application can be combined interactively without conflict.
[0032] Currently, the smoking monitoring in specific scenarios mainly includes the following methods: infrared thermal induction-based method, motion trajectory analysis-based method, depth image-based method, and wearable device monitoring-based method.
[0033] 1. The infrared thermal induction-based method uses a special infrared camera for smoking detection. The infrared thermal imaging camera can present different colors for objects of different temperatures in the imaging picture compared with the human body and the surrounding environment. The temperature at the cigarette burning end is higher, so the color of the burning cigarette in the infrared thermal imaging camera is darker, thereby judging it as a smoking action. The current method has low accuracy when the surrounding temperature is high, and when smoking far away from the infrared device, the temperature sensing effect of the device will also decrease, so it cannot accurately identify whether a smoking action occurs.
[0034] 2. The motion trajectory analysis-based method extracts a video frame, realizes smoking recognition in the dynamic trajectory of the features in the video frame, and the dynamic features mainly refer to the motion trajectory of the burning generated smoke area and the center point of the smoke area. Then, through the means of deep learning, the potential smoke area and the trajectory of the center point are learned to identify whether a smoking action occurs. This method processes video frame sequences, which requires a large amount of computing resources.
[0035] 3. The depth learning-based method uses a portable mobile device to obtain depth images for analysis. The sensing distance of the portable device is limited, and only the smoking behavior in front of the device can be accurately identified. This method is not suitable for large-scale smoking behavior recognition in public scenarios.
[0036] 4. The terminal device-based cigarette smoking identification method uses a wearable device to detect cigarette smoking. The wearable device collects real-time momentum information in the X, Y, and Z directions of the acceleration sensor, judges whether there is a reciprocating hand movement, and combines the micro smoke sensing device on the device to sense whether there is smoke in the surrounding space, thereby identifying the smoking behavior. The current method requires each member to wear a terminal device, has low accuracy, high cost, and does not meet the actual scene usage.
[0037] The cigarette smoking behavior identification method of the present application uses existing monitoring cameras to obtain real-time video frame sequences and uses deep learning methods to perform overall target detection on the images. When the smoking behavior occurs on a person's body, the target detection can obtain the person and the hand and mouth regions of interest; a specific model is established to calculate the positions of the hand joints and the mouth, and the smoking behavior is identified and judged by analyzing the shape of the hand joints, the position information, and the recognition effect of the regions of interest. The method has multiple advantages: (1) it uses time-series deep learning modeling to identify smoking behavior, has a simple device, low cost, and wide coverage; (2) the cigarette smoking monitoring camera of the present application is installed at a certain distance from the ground, only a single monitoring device is needed, and no modification of the existing environment is required during installation and deployment, and no additional devices with specific markers need to be installed on the person to be identified, which has the characteristics of flexible deployment, adaptability to various environments; (3) the method uses artificial intelligence machine learning to detect targets in the monitoring area, which can effectively remove interference from different temperature objects, machines, and other devices in infrared imaging, has strong robustness, and can adapt to complex environments; (4) the method uses target detection and hand joint, mouth position distance and picture to monitor smoking, can simultaneously calculate multiple regions of interest in real time, and has the functions of multiple target smoking monitoring and strong practicality. The following is a detailed description of the present application.
[0038] Referring to Figure 1 The cigarette smoking behavior identification method includes the following S1-S3 steps:
[0039] S1, acquire real-time video stream through a monitoring device, calculate the video stream through a deep learning target detection model efficientDet, and obtain the region of interest of each person, including the position of the hand and the position of the mouth.
[0040] Specifically, the deep learning target detection model efficientDet calculates the video stream, including the following steps:
[0041] S101, perform feature extraction on the video stream through a dynamically scalable convolutional neural network;
[0042] S102, fuse the extracted features in the form of a recurrent skip combination;
[0043] S103, calculate the target information according to the fused features.
[0044] S2, obtain the positions of the key points in the region of interest through the hand key point model, and calculate the key point position distribution information. As a deep learning neural network model, the hand key point model needs to be trained through a large amount of hand data.
[0045] Specifically, the step S2 includes the following steps:
[0046] S21, a large amount of palm data, for example, 65000, is collected by using an existing monitoring device, the nodes of the palm are labeled by using an automatic node labeling method, and the obtained nodes are mapped to a three-dimensional space by using a Pixel2Mesh algorithm. Pixel2Mesh is a 3D mesh model generated from a single RGB picture pixel;
[0047] Specifically, the Pixel2Mesh algorithm is used to map the obtained nodes to a three-dimensional space, including the following steps:
[0048] S201, an arbitrary input picture is randomly initialized as an initial three-dimensional feature shape, and the network model has two branches. One branch is responsible for feature extraction of the input picture by using a full convolutional neural network, and the other branch is responsible for three-dimensional mesh feature extraction;
[0049] S202, the three-dimensional mesh is deformed into a required three-dimensional model by continuously fusing and calculating the two branches.
[0050] S22, the spatial node data is sent to a spatial relationship network SpatialRelationshipNet for calculation, and the obtained feature map is rearranged;
[0051] Specifically, the step S22 includes the following steps:
[0052] S221, the input picture is converted into specific dense spatial point position information by calculation;
[0053] S222, a convolutional neural network CBL structure outputs two types of feature maps, which are obtained by maximum pooling and average pooling respectively, and the two types of pooling results are fused;
[0054] S223, the fused features are sent to the next network structure, and the obtained feature map is rearranged.
[0055] The structure of SpatialRelationshipNet is shown in the following figure:Figure 2 As shown, A is an input picture, an SPP structure is added in the model, the length and width of the input picture can be any 8 times multiple size, Block1 is a Pixel2Mesh algorithm, the Pixel2Mesh algorithm converts the input picture into specific dense spatial point position information through calculation;Block2-1 is a convolutional neural network CBL structure, Block2-1 outputs two types of feature maps, which are obtained by maximum pooling and average pooling respectively, B performs feature fusion on the two pooling results, and then sends the fused features to the next network structure Block-2, Block-2 re-arranges the obtained feature map, increases the judgment ability of the model and improves the robustness of the model through the combination of different feature positions. The re-arranged feature map is calculated by the full connection layer to output the final result.
[0056] S23, the re-arranged feature map is calculated by the full connection layer to output the final result, i.e. the joint position distribution information.
[0057] S3, when holding a cigarette with hands, the positions between the nodes of the palms will be in several special distribution forms to complete specific actions, in this step, the joint position distribution information is compared with the existing standard palm node distribution when holding a cigarette to judge whether the current palm action conforms to the palm node position distribution when smoking, if yes, it is recognized as a smoking behavior. Further, the method further comprises step S4, comparing the hand holding the cigarette in the target detection result with the hand holding other objects or the hand without holding objects to comprehensively judge whether the current action behavior is a smoking behavior, further improving the accuracy of smoking recognition.
[0058] The application also provides a smoking behavior recognition system, comprising a deep learning target detection model efficientDet, a hand node model and a judgment module,
[0059] The deep learning target detection model efficientDet is used for calculating the obtained real-time video stream to obtain a region of interest of a person, and the region of interest includes the position of the hand and the position of the mouth;
[0060] The hand node model is used for obtaining the position of the joint in the region of interest and calculating the joint position distribution information;
[0061] The judgment module is used for comparing the joint position distribution information with the existing standard palm node distribution when holding a cigarette to judge whether the current palm action conforms to the palm node position distribution when smoking, if yes, it is recognized as a smoking behavior.
[0062] The smoking behavior recognition method and system of the present application utilizes time series deep learning modeling to recognize smoking behavior, first obtains the candidate region of interest of personnel through target detection, thereby reducing unnecessary detection, so that the computing resources can be better utilized; and the target detection method of deep learning can remove the interference of other non-specific targets, is more robust, and is suitable for various complex scenes; at the same time, target detection can obtain multiple personnel of interest at the same time, and can simultaneously recognize whether different individuals have smoking behavior. Moreover, compared with the traditional smoking recognition method, the recognition method implemented by the present application using time series deep learning modeling is convenient and flexible to deploy, does not need to be additionally equipped with special equipment, only needs the existing monitoring camera, and is low in cost; and can simultaneously monitor the behavior of multiple personnel, and has strong practicability.
[0063] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the described embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application. These equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A smoking behavior recognition method, characterized by: The steps include the following: S1, acquiring a real-time video stream, calculating the video stream through a deep learning target detection model efficientDet to obtain a region of interest of a person, the region of interest including a position of a hand and a position of a mouth; The deep learning target detection model efficientDet calculates the video stream, including the following steps: S101, performing feature extraction on the video stream through a dynamic scalable convolutional neural network; S102, fusing the extracted features in a form of recurrent skip combination; S103, calculating the fused features to obtain target information; S2, acquiring positions of joints in the region of interest through a hand joint model, and calculating joint position distribution information; The step S2 includes the following steps: S21, collecting a large amount of palm data, labeling each node of the palm of the palm data using an automatic node labeling method, and mapping the obtained nodes to a three-dimensional space using a Pixel2Mesh algorithm; S22, calculating spatial node data through a SpatialRelationshipNet network, and rearranging the obtained feature maps; S23, the rearranged feature maps are calculated through a full connection layer to output the final result, i.e., the joint position distribution information; S3, comparing the joint position distribution information with an existing standard hand-held cigarette palm node distribution to determine whether the current palm action conforms to the palm node position distribution when smoking, and if so, identifying the action as a smoking behavior.
2. The smoking behavior recognition method of claim 1, wherein: Further including the step: S4, comparing the hand holding the cigarette in the target detection result with the hand holding other objects or the hand not holding objects to determine whether the current action is a smoking behavior.
3. The smoking behavior recognition method of claim 1, wherein: In the step S21, the Pixel2Mesh algorithm is used to map the obtained nodes to a three-dimensional space, including the following steps: S201, initializing an arbitrary input picture as an initial three-dimensional feature shape of a target outer ellipse, the network model having two branches, one branch being responsible for feature extraction of the input picture using a full convolutional neural network, and the other branch being responsible for three-dimensional mesh feature extraction; S202, deforming the three-dimensional mesh into a required three-dimensional model through continuous fusion calculation of the two branches.
4. The smoking behavior recognition method of claim 3, wherein: The step S22 includes the following steps: S221, converting the input picture calculation into specific dense spatial point position information; S222, a convolutional neural network CBL structure outputs two types of feature maps, respectively obtained by maximum pooling and average pooling, and fuses the two types of pooling results; S223, the fused features are sent to the next network structure, and the obtained feature maps are rearranged.
5. A smoking behavior recognition system, characterized by: The system is used to implement the smoking behavior recognition method according to any one of claims 1-4, including a deep learning target detection model efficientDet, a hand joint model, and a judgment module. The deep learning target detection model efficientDet is used for calculation on the acquired real-time video stream to acquire a region of interest of a person, the region of interest including a position of a hand and a position of a mouth; The hand joint node model is used for acquiring positions of joint nodes in the region of interest and calculating position distribution information of the joint nodes; The judging module is used for comparing the position distribution information of the joint nodes with a standard hand-held cigarette palm joint node distribution to judge whether a palm action currently occurring conforms to a palm joint node position distribution when smoking, and if yes, the smoking behavior is recognized.
Citation Information
Patent Citations
Smoking behavior detection method in monitoring scene based on computer vision
CN112115775A