A smoking detection method and device based on a low-power GPU device

A smoking detection method combining an RGBD camera and a low-power GPU device utilizes changes in the chest and abdomen region and smoke diffusion sequences to construct a smoking behavior model, solving the problem of high false alarm rates in traditional cigarette and e-cigarette detection and achieving more efficient smoking detection.

CN117253288BActive Publication Date: 2026-01-20FOSHAN HONGSHI INTELLIGENT INFORMATION TECH CO LTD
View PDF 23 Cites 0 Cited by

Patent Information

Application Number
CN202311273074.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-01-20
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

Existing smoking detection methods have a high false alarm rate when detecting traditional cigarettes and e-cigarettes, especially e-cigarettes, and fail to effectively reflect smoking behavior.

Method used

An RGBD camera is used to collect depth stream data and video stream data. Human body localization and smoke detection are performed using a low-power GPU device. By combining chest and abdominal region change sequences and smoke diffusion image sequences, a smoking behavior detection model is constructed to reduce the false alarm rate.

Benefits of technology

It improves detection effectiveness, significantly reduces false alarm rate, and can more accurately identify smoking behavior. It is applicable to both traditional cigarettes and e-cigarettes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253288B_ABST
    Figure CN117253288B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based on low-power GPU equipment's smoking detection method and device, it is related to intelligent smoke detection technical field.The application discloses a kind of based
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent smoke detection, and particularly relates to a smoking detection method and device based on a low-power GPU device. BACKGROUND

[0002] Most of the existing smoking detection is the detection of traditional cigarette smoking. With the popularity of electronic cigarettes, the smoking method and holding method have also changed greatly. In order to have good detection effect for both traditional cigarette smoking and electronic cigarette smoking, the present application designs a smoking detection method and device based on a low-power GPU device. The methods close to the present application are still mainly the detection methods of traditional cigarette smoking, including detection based on a fire sensor, cigarette image detection, joint detection of cigarette and human body, continuous action detection, video combined with sound detection, detection based on an infrared device, smoke-based detection, etc. The advantages and disadvantages of several commonly used smoking detection methods are introduced below.

[0003] Detection based on a fire sensor: this is the most common fire alarm method in daily life, which is convenient to install and low in cost. However, the fire sensor often needs to reach a certain degree of flame or smoke to produce an alarm, and many times the situation of smoking cannot trigger the fire sensor.

[0004] Cigarette image-based detection: this method mainly detects the image features of the cigarette itself, and mostly uses a deep learning detection scheme. A plurality of angle and scene cigarette images are pre-collected for training, and a video frame based on monitoring is used for cigarette image detection. For example, CN202310042572.5, a smoking detection data labeling method based on active learning, CN202210508435.1, a smoking recognition method, system, device and storage medium, CN202111511659.X, a safety hat wearing and smoking detection safety monitoring method, CN202211710731.6, a smoking behavior detection model construction method based on a double-main network, etc. The advantages of this method are low cost and the use of existing monitoring networks to realize cigarette detection; the disadvantage is that there are many false positives. Since the structural features of the cigarette itself are very simple, there may be a large number of targets similar to the cigarette image in daily life, which is easy to produce false positives. In addition, for electronic cigarettes, due to the holding posture, the electronic cigarette is often blocked by the hand, and it is not easy to detect the target image of the electronic cigarette.

[0005] Based on the joint detection of cigarettes and human body: This method not only detects cigarettes, but also makes a related judgment on the human body or human joint points, such as judging the relationship between cigarettes and human faces, heads, hands, and mouths, or judging the human body smoking posture based on bones. For example, CN202211482519.9 A smoking detection method, device, equipment and storage medium, CN202211261654.0 A smoking detection method and device, electronic equipment, computer readable storage medium, CN202110885747.X A real-time multi-human angle smoking behavior detection method, CN202110795174.1 A driver smoking detection method based on multiple models, CN202011017414.7 A method for identifying illegal smoking based on AlphaPose, CN201911014203.5 A driver smoking detection method and system based on computer vision technology, CN202211621833.0 A smoking behavior detection method and system in industrial scene monitoring video, CN202211148966.0 Smoking detection method, device, equipment and storage medium, CN202210576591.1 A method for detecting smoking of personnel in substation monitoring, CN202010650354.6 A smoking detection method and device, etc. The advantage of such methods is that on the basis of pure cigarette detection, the interaction between the human body and the cigarette is increased, which can filter more false positives; the disadvantage is that pure image judgment still has some false positives similar to smoking.

[0006] Based on continuous action smoking detection: This method is based on the joint detection of cigarettes and human body, and further increases the detection of continuous action, such as CN202211125461.2 A smoking behavior detection method, system and device, which identifies suspected smoking behavior through image recognition, combines continuous frame action judgment, uses spatiotemporal convolution to judge the action of the hand close to the nose when smoking, and determines whether it is a true smoking behavior. The advantage of this method is that it more accurately reflects the characteristics of smoking behavior and reduces false positives; the disadvantage is that it does not consider features other than action such as smoke.

[0007] Smoking detection based on video combined with sound: This method mainly detects suspected smoking images first, and then judges whether there is a sound of a lighter or a deep breathing sound of smoking. The sound of the lighter is collected by monitoring or wearable devices, while the deep breathing sound of smoking is usually collected by the microphone of the wearable device. For example, CN202211274850.1 A multi-angle smoking detection alarm system, CN201610356958.3 A daily smoking behavior detection method based on wearable devices, etc. The advantage of this method is that it refers to audio features other than video, which are closely related to traditional smoking and can detect smoking behavior from another perspective. The disadvantage is that if a camera sound collector is used, the sound collection is greatly affected by the environment, and if a wearable device is used, it has a more stringent restriction on the smoking person being detected, making it difficult to be universal.

[0008] Smoking detection based on infrared camera: This method uses infrared equipment to detect the high temperature characteristics of cigarette burning, and combines visible light images to judge smoking. For example, CN202210652952.6 A neural network and infrared image matching kitchen smoking detection method and system, CN202210343702.4 A smoking early warning method based on image recognition, CN202111438613.X An infrared image smoking detection method and system based on deep learning, etc. The advantage of this method is that it has fewer false positives and more accurate detection; the disadvantage is that it needs to add expensive infrared temperature detection equipment separately, and in addition, the heating of electronic cigarettes is much lower than that of traditional cigarettes, and due to the holding posture, electronic cigarettes are often blocked by hands, making it difficult to detect temperature.

[0009] Smoking detection based on smoke: This method directly detects smoke and judges whether someone is smoking by detecting smoke in the scene. Some methods of smoke detection in fire scene smoke detection are also similar. For example, CN202210284951.0 A smoking detection method, device, electronic equipment and storage medium. The advantage of this method is that it is very intuitive and has universality for electronic cigarette smoke; the disadvantage is that it does not associate smoke with some details of human smoking.

[0010] Traditional detection methods mostly focus on cigarette detection or smoking action posture detection, but the cigarette itself is small and difficult to detect, and some new electronic cigarettes are more difficult to detect because they are often held. Therefore, the characteristics of smoke itself are the key characteristics of smoking. However, previous smoke detection-based schemes do not consider the relationship between smoke and human respiration, and simply judging smoke can easily produce false positives due to environmental smoke or some environmental light problems. In addition, smoke is thick or thin, and in some monitoring scenes, it is difficult to see the smoke of smoking.

[0011] Therefore, how to more objectively reflect the action of smoking behavior and reduce the possibility of false positives is a problem that those skilled in the art need to solve. SUMMARY

[0012] Therefore, the present application provides a low-power GPU device-based smoking detection method and device, which can enhance the detection effect and greatly reduce the false positive rate.

[0013] To achieve the above purpose, the present application adopts the following technical solutions:

[0014] A low-power GPU device-based smoking detection method, comprising the following steps:

[0015] S1, the RGBD camera inputs the collected depth stream data and video stream data to the low-power GPU device;

[0016] S2, the low-power GPU device performs human body positioning operation on the head position in the video stream data and the chest and abdomen position in the depth stream data;

[0017] S3, smoke detection is performed in the head position preset area, the detected smoke area is recorded, and a smoke diffusion image sequence is calculated;

[0018] S4, the depth data change of the chest and abdomen preset area is calculated, and a chest and abdomen area change sequence is extracted;

[0019] S5, a smoking behavior detection model is constructed, the smoke diffusion image sequence in S3 and the chest and abdomen area change sequence in S4 are input to the smoking behavior detection model, and when the model judges that it is a smoking behavior, the low-power GPU device generates an alarm signal.

[0020] The above method, optionally, the human body positioning operation in S2 comprises the following steps:

[0021] S201, data preprocessing: pre-processing the input image, including but not limited to cropping, scaling, and normalization;

[0022] S202, feature extraction: using two parallel CNN models, set as the first CNN model and the second CNN model, to extract the features of different parts of the body, wherein the first CNN model is used to extract the features of key parts including but not limited to the head, shoulders, hips, and knee joints, and the second CNN model is used to extract the features of the hands and feet;

[0023] S203, PAF calculation: calculating the vector relationship between each pixel and adjacent pixels in the image to determine the joint position of the human body;

[0024] S204, pose estimation: full-body pose estimation is performed, a third CNN model is used to calculate the probability of each pixel corresponding to a joint, the position of each joint is calculated according to the probability, and the positions of the joints are connected using geometric rules and thresholds to form a full-body pose estimation result;

[0025] S205, result output: output the full-body pose estimation result as key point coordinates and joint angle information. The above method, optionally in S3, calculates the smoke diffusion image sequence, including the following steps:

[0026] S301, let the nose feature point be P0, the right eye point be P 15 , the left eye point be P 16 , the right ear point be P 17 , and the left ear point be P 18 When the human feature point detects any 3 or more points in the 5 points, take the minimum circumscribed rectangle of the point set as the human head region, and take 3 times the head region size as the smoke diffusion image region;

[0027] S302, in the RGB video stream, let the nose point P0 coordinate be (x0, y0);

[0028] Let the right eye point P 15 coordinate be (x 15 , y 15 );

[0029] Let the left eye point P 16 coordinate be (x 16 , y 16 );

[0030] Let the right ear point P 17 coordinate be (x 17 , y 17 );

[0031] Let the left ear point P 18 coordinate be (x 18 , y 18 );

[0032] If the feature point does not exist, the x and y coordinates of the corresponding feature point are both assigned a value of -1;

[0033] S303, in the human skeleton feature points P0, P 15 , P 16 , P 17 , P 18

[0034] Let the smallest x coordinate be

[0035] Min_x = min(x0, x 15 , x 16 , x​17 ,x 18 ),

[0036] The smallest y-coordinate is

[0037] Min_y = min(y0, y 15 ,y 16 ,y 17 ,y 18 ),

[0038] The largest x-coordinate is

[0039] Max_x = max(x0, x 15 ,x 16 ,x 17 ,x 18 ),

[0040] The largest y-coordinate is

[0041] Max_y = max(y0, y 15 ,y 16 ,y 17 ,y 18 ),

[0042] Here, min() and max() are the functions for calculating the minimum and maximum values, respectively. When a feature point does not exist, that is, when the x and y coordinates of a feature point are both -1, the feature point does not participate in the calculation of the minimum and maximum values. The coordinates of the upper left corner of the human head region rectangle are (Min_x, Min_y), and the coordinates of the lower right corner are (Max_x, Max_y).

[0043] S304. Take three times the size of the head region as the smoke diffusion image area, and set the coordinates of the upper left corner of the smoke diffusion area rectangle as (x... L ,y L The coordinates of the lower right corner are (x R ,y R ),in,

[0044] x L =Min_x - (Max_x - Min_x),

[0045] y L =Min_y - (Max_y - Min_y),

[0046] x R =Max_x + (Max_x - Min_x),

[0047] y R =Max_y + (Max_y - Min_y);

[0048] S305, real-time image region is intercepted from the RGB video stream as a smoke diffusion image sequence L with the smoke diffusion region B .

[0049] The method described above, optionally, the chest and abdomen region change sequence is extracted in S4, comprising the following steps:

[0050] S401, the neck point is P1 and the hip center point is P8 in the RGB video stream, the neck point P1 coordinate is (x1, y1); the hip center point P8 coordinate is (x8, y8); the midpoint of P1 and P8 is P A , the coordinate is (x A ,y A ), wherein,

[0051] x A =(x1+x8) / 2,

[0052] y A =(y1+y8) / 2;

[0053] S402, the depth stream is registered with the video stream so that the coordinate points of x and y are one-to-one, then,

[0054] The coordinate point corresponding to P1 in the depth stream is (x1, y1, z1);

[0055] The coordinate point corresponding to P8 in the depth stream is (x8, y8, z8);

[0056] The coordinate point corresponding to P A in the depth stream is (x A ,y A ,z A );

[0057] S403, the three-dimensional space distance D A between P A1 and P1 is calculated:

[0058]

[0059] S404, the three-dimensional space distance D A between P A8 and P8 is calculated:

[0060]

[0061] S405, the average distance D A is calculated:

[0062] D A =(D A1 +D A8 ) / 2;

[0063] S406, record the mean distance D A the sequence of changes, and take the sequence characteristic value as the sequence of changes L of the chest and abdominal region A .

[0064] The method described above, optionally, the method for judging the smoking behavior in S5 is that the smoking detection behavior model outputs the classification result confidence of the current data collection in real time, the confidence is a floating point value between 0 and 1, and when the confidence is greater than a set threshold, the system judges that the classification result is smoking.

[0065] A smoking detection device based on a low-power GPU device, which executes the smoking detection method based on the low-power GPU device, and comprises an RGBD camera, a low-power GPU device and a smoking detection alarm connected in sequence.

[0066] According to the technical solution described above, compared with the prior art, the present application provides a smoking detection method and device based on a low-power GPU device, which has the following beneficial effects: 1) the present application uses an RGBD camera to collect synchronous sequence of changes of the chest and abdominal region and sequence of smoke diffusion images, which can better enhance the movement of breathing and smoke diffusion, and the synchronous data more objectively reflects the action of smoking, greatly reducing the possibility of false positives; 2) the present application designs two synchronous sequence input models of the sequence of changes of the chest and abdominal region and the sequence of smoke diffusion images, which can not only describe two independent actions of exhalation and smoking, but also extract features that can reflect the relationship between exhalation and smoke diffusion over time, greatly improving the overall detection effect and false positive filtering. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0068] Figure 1 A flow chart of the smoking detection method based on a low-power GPU device disclosed by the present application is shown in the figure.

[0069] Figure 2 An OPENPOSE network structure diagram disclosed by the present embodiment is shown in the figure.

[0070] Figure 3 A human feature point sequence schematic diagram disclosed by the present embodiment is shown in the figure.

[0071] Figure 4 A smoking detection behavior model schematic diagram disclosed by the present embodiment is shown in the figure. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0073] In the present application, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or equipment including the element.

[0074] Referring to Figure 1 The present application discloses a smoking detection method based on a low-power GPU device, comprising the following steps:

[0075] S1, the RGBD camera inputs the collected depth stream data and video stream data to the low-power GPU device;

[0076] S2, the low-power GPU device performs human body positioning operation on the head position in the video stream data and the chest and abdomen position in the depth stream data;

[0077] S3, smoke detection is performed in the head position preset area, the detected smoke area is recorded, and a smoke diffusion image sequence is calculated;

[0078] S4, the depth data change of the chest and abdomen preset area is calculated, and a chest and abdomen area change sequence is extracted;

[0079] S5, a smoking behavior detection model is constructed, the smoke diffusion image sequence in S3 and the chest and abdomen area change sequence in S4 are input to the smoking behavior detection model, and when the model judges that it is a smoking behavior, the low-power GPU device generates an alarm signal.

[0080] Further, referring to Figure 2 The human body positioning operation in S2 comprises the following steps:

[0081] S201, data preprocessing: pre-processing the input image, operations include but are not limited to cropping, scaling, normalization;

[0082] S202, feature extraction: using two parallel CNN models, set as the first CNN model and the second CNN model, respectively extracting the features of different parts of the body, wherein the first CNN model is used to extract the features of key parts including but not limited to head, shoulder, hip and knee joint, and the second CNN model is used to extract the features of hand and foot;

[0083] S203, PAF calculation: calculating the vector relationship between each pixel and adjacent pixel in the image, determining the joint position of the human body;

[0084] S204, pose estimation: full body pose estimation, using a third CNN model to calculate the probability of each pixel corresponding to a joint, calculating the position of each joint according to the probability, using geometric rules and threshold to connect the positions of the joints, forming the full body pose estimation result;

[0085] S205, result output: outputting the full body pose estimation result as key point coordinates and joint angle information.

[0086] Specifically, in S202, the input of the two models is the pre-processed image, and they learn a set of feature representations related to their respective tasks. After the operation of multiple convolutional layers, pooling layers and fully connected layers, the feature extraction model will output a set of feature vectors. For each pixel input, there will be a corresponding feature vector. This feature vector contains posture information related to the pixel, such as which joint the pixel belongs to, the position and posture of the joint, etc. These feature vectors will be used as input for subsequent PAF calculation and pose estimation.

[0087] Specifically, PAF is the abbreviation of Pairwise Aggregation Field, which represents the vector relationship between each pixel and adjacent pixel in the image. Specifically, PAF is a two-dimensional vector field, where each vector represents the position and direction relationship of a pixel relative to its adjacent pixel. PAF can represent the joint position of the human body, and by calculating PAF, the joint position of the human body can be determined, so as to perform full body pose estimation.

[0088] Further, referring to Figure 3 , in S3, the smoke diffusion image sequence is calculated, including the following steps:

[0089] S301, let the nose feature point be P0, the right eye point be P 15 , the left eye point be P 16 , the right ear point be P 17 , and the left ear point be P 18When the human feature points detect more than 3 points in the 5 points, the minimum circumscribed rectangle of the point set is taken as the human head region, and 3 times the head region size is taken as the smoke diffusion image region;

[0090] S302, in the RGB video stream, let the nose point P0 coordinate as (x0, y0);

[0091] Let the right eye point P 15 coordinate as (x 15 ,y 15 );

[0092] Let the left eye point P 16 coordinate as (x 16 ,y 16 );

[0093] Let the right ear point P 17 coordinate as (x 17 ,y 17 );

[0094] Let the left ear point P 18 coordinate as (x 18 ,y 18 );

[0095] If the feature point does not exist, the x and y coordinates of the corresponding feature point are both assigned as -1;

[0096] S303, in the human body skeleton feature points P0, P 15 , P 16 , P 17 , P 18 ,

[0097] Let the minimum x coordinate as

[0098] Min_x=min(x0,x 15 ,x 16 ,x 17 ,x 18 ),

[0099] The minimum y coordinate as

[0100] Min_y=min(y0,y 15 ,y 16 ,y 17 ,y 18 ),

[0101] The maximum x coordinate as

[0102] Max_x=max(x0,x 15 ,x 16 ,x 17 ,x 18 ),

[0103] Max_y = max(y0, y

[0104] Max_y = max(y0, y 15 ,y 16 ,y 17 ,y 18 ),

[0105] wherein min() and max() are the minimum and maximum calculation functions respectively, when a certain feature point does not exist, i.e. the x and y coordinates of a certain feature point are both -1, the feature point does not participate in the minimum and maximum calculation; the head region rectangular frame left upper corner coordinate is (Min_x, Min_y), and the right lower corner coordinate is (Max_x, Max_y);

[0106] S304, taking 3 times the head region size as the smoke diffusion image region, and letting the smoke diffusion region rectangular frame left upper corner coordinate be (x L ,y L ), and the right lower corner coordinate be (x R ,y R ), wherein,

[0107] x L = Min_x - (Max_x - Min_x),

[0108] y L = Min_y - (Max_y - Min_y),

[0109] x R = Max_x + (Max_x - Min_x),

[0110] y R = Max_y + (Max_y - Min_y);

[0111] S305, real-time image region is intercepted from the RGB video stream in the smoke diffusion region, as the smoke diffusion image sequence L B .

[0112] Further, the chest and abdomen region change sequence is extracted in S4, including the following steps:

[0113] S401, letting the neck point be P1 and the hip center point be P8 in the RGB video stream, letting the neck point P1 coordinate be (x1, y1); letting the hip center point P8 coordinate be (x8, y8); the midpoint of the line connecting P1 and P8 is P A , the coordinate of which is (x A ,y A ), wherein,

[0114] x A= (x1 + x8) / 2,

[0115] y A = (y1+y8) / 2;

[0116] S402. Registering the depth stream and the video stream ensures a one-to-one correspondence between the x and y coordinates.

[0117] The coordinates of P1 in the depth flow are (x1, y1, z1);

[0118] The coordinates of P8 in the depth flow are (x8, y8, z8);

[0119] P A The corresponding coordinate point in the depth flow is (x A ,y A ,z A );

[0120] S403, Calculate P A The three-dimensional spatial distance D from P1 A1 :

[0121]

[0122] S404, Calculate P A The three-dimensional spatial distance D from P8 A8 :

[0123]

[0124] S405, Calculate the mean distance D A :

[0125] D A =(D A1 +D A8 ) / 2;

[0126] S406, Record the mean distance D A The change sequence is used as the characteristic value of this sequence as the change sequence L of the thoracic and abdominal region. A .

[0127] Specifically, when smoking, whether it is a traditional cigarette or an electronic cigarette, the action of inhaling and exhaling will be generated, and the chest and abdomen area change sequence is used to describe the action of inhaling and exhaling. Let the neck point be P1 and the hip center point be P8, considering that the human body may be in a state of motion and different angles, here, only in the case that the neck point P1 and the hip center point P8 can be detected, the chest and abdomen area change sequence feature extraction is carried out. Since the positions of the neck point P1 and the hip center point P8 are the upper end and the lower end of the chest and abdomen, when breathing, the most obvious change is the midpoint A of the line connecting the neck point P1 and the hip center point P8. In order to avoid the interference of people walking and body movement, the spatial relative relationship between point A and the neck point P1 and the hip center point P8 is calculated to describe the chest and abdomen area change.

[0128] Further, the method for judging the input video sequence as a smoking behavior in S5 is that the smoking detection behavior model outputs the classification result confidence of the current data collection in real time, the confidence is a floating point value between 0 and 1, and when the confidence is greater than a set threshold, the system judges that the classification result is smoking.

[0129] Further, as shown in Figure 4 , the chest and abdomen area change sequence L A and the smoke diffusion image sequence L B are combined, and the features of the chest and abdomen area change sequence L A are one-dimensional data sequences, and one-dimensional convolution can learn the features of the chest and abdomen fluctuation change process; the smoke diffusion image sequence L B is a video sequence, and three-dimensional convolution can learn the video features of the smoke image diffusion with exhalation. L A and L B have synchronous correlation in time, so the application designs a smoking detection behavior model combining one-dimensional convolution and three-dimensional convolution.

[0130] Specifically, the chest and abdomen area change sequence L A is input into a one-dimensional convolution subnetwork, and the subnetwork includes convolution layers Conv1D_1, convolution layers Conv1D_2, maximum pooling layers MaxPooling1D, convolution layers Conv1D_3, convolution layers Conv1D_4, global average pooling layers GlobalAveragePooling1D, random inactivation layers Dropout and full connection layers fc_1D connected in sequence. The smoke diffusion image sequence L BThe input three-dimensional convolution subnetwork includes convolution layer Conv1a, pooling layer Pool1, convolution layer Conv2a, pooling layer Pool2, convolution layers Conv3a and Conv3b, pooling layer Pool3, convolution layers Conv4a and Conv4b, pooling layer Pool4, convolution layers Conv5a and Conv5b, pooling layer Pool5, two fully connected layers fc6 and fc7 connected in sequence. The one-dimensional convolution subnetwork fully connected layer fc_1D is added with the three-dimensional convolution subnetwork fully connected layer fc7 features, and finally fused through the loss function softmax layer to output the current classification confidence. The confidence is a floating point value between 0 and 1.

[0131] Specifically, in actual smoking detection, the model is deployed in a low-power GPU device, and the chest and abdominal region change sequence L A and the smoke diffusion image sequence L B are collected, input into the smoking detection behavior model for classification, and the classification result confidence of the current data collection is output in real time. When the confidence is greater than a set threshold, the system judges that the classification result is smoking, and the system generates an alarm.

[0132] A smoking detection device based on a low-power GPU device, which executes the smoking detection method based on a low-power GPU device, and includes an RGBD camera, a low-power GPU device and a smoking detection alarm connected in sequence.

[0133] Each of the embodiments in the specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments. Especially, the system or system embodiment is described relatively simply because it is basically similar to the method embodiment, and the relevant part can be referred to the part of the method embodiment. The above-described system and system embodiment are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to the actual needs. Those skilled in the art can understand and implement without creative labor.

[0134] In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been described in the above description in general terms. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0135] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Numerous modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without the use of the inventive faculty. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting smoke using a low-power GPU device, characterized in that, Includes the following steps: The S1 and RGBD cameras input the acquired depth stream data and video stream data to the low-power GPU device; S2. Low-power GPU devices perform human body localization operations by taking the head position in the video stream data and the chest and abdomen position in the depth stream data. S3. Perform smoke detection in a preset area at the head position, record the detected smoke area and calculate the smoke diffusion image sequence; S4. Calculate the depth data changes of the preset chest and abdomen region and extract the chest and abdomen region change sequence; S5. Construct a smoking behavior detection model. Input the smoke diffusion image sequence in S3 and the chest and abdominal region change sequence in S4 into the smoking behavior detection model. When the model determines that it is a smoking behavior, the low-power GPU device generates an alarm signal. In S3, the calculation of the smoke diffusion image sequence includes the following steps: S301, Let the nose feature point be P0 and the right eye point be P. 15 The left eye point is P. 16 The right ear point is P. 17 The left ear point is P. 18 When human feature points detect any 3 or more of these 5 points, the smallest bounding rectangle of the point set is taken as the human head region, and 3 times the size of the head region is taken as the smoke diffusion image region. S302. In the RGB video stream, let the coordinates of the nose point P0 be (x0, y0); Make right eye point P 15 Coordinates are (x 15 ,y 15 ); Make the left eye point P 16 Coordinates are (x 16 ,y 16 ); Point P to the right ear 17 Coordinates are (x 17 ,y 17 ); Point P on the left ear 18 Coordinates are (x 18 ,y 18 ); If the feature point does not exist, assign -1 to both the x and y coordinates of the corresponding feature point; S303, at human skeletal feature points P0, P... 15 P 16 P 17 P 18 middle Let the smallest x-coordinate be , The smallest y-coordinate is , The largest x-coordinate is , The largest y-coordinate is , Here, min() and max() are the functions for calculating the minimum and maximum values, respectively. When a feature point does not exist, that is, when the x and y coordinates of a feature point are both -1, the feature point does not participate in the calculation of the minimum and maximum values. The coordinates of the upper left corner of the human head region rectangle are (Min_x, Min_y), and the coordinates of the lower right corner are (Max_x, Max_y). S304. Take three times the size of the head region as the smoke diffusion image area, and set the coordinates of the upper left corner of the smoke diffusion area rectangle as (x... L ,y L The coordinates of the lower right corner are (x R ,y R ),in, , , , ; S305. Extract an image region from the RGB video stream in real time, based on the smoke diffusion area, as the smoke diffusion image sequence L. B ; Extracting the thoracic and abdominal region variation sequence from S4 includes the following steps: S401. Let the neck point be P1 and the center point of the buttocks be P8. In the RGB video stream, let the coordinates of the neck point P1 be (x1, y1); let the coordinates of the center point of the buttocks P8 be (x8, y8); and let the midpoint of the line connecting P1 and P8 be P. A Its coordinates are (x A ,y A ),in, , ; S402. Registering the depth stream and the video stream ensures a one-to-one correspondence between the x and y coordinates. The coordinates of P1 in the depth flow are (x1, y1, z1); The coordinates of P8 in the depth flow are (x8, y8, z8); P A The corresponding coordinate point in the depth flow is (x A ,y A ,z A ); S403, Calculate P A Three-dimensional spatial distance from P1 D A1 : ; S404, Calculate P A Three-dimensional spatial distance from P8 D A8 : ; S405, Calculate the mean distance D A : ; S406, Record the mean distance D A The change sequence is used as the characteristic value of this sequence as the change sequence L of the thoracic and abdominal region. A .

2. The method for detecting smoke based on a low-power GPU device according to claim 1, characterized in that, The human body positioning operation in S2 includes the following steps: S201. Data preprocessing: Preprocess the input image, including but not limited to cropping, scaling, and normalization. S202, Feature Extraction: Two parallel CNN models, designated as the first CNN model and the second CNN model, are used to extract features from different parts of the body. The first CNN model is used to extract features from key parts, including but not limited to the head, shoulders, hips, and knees, while the second CNN model is used to extract features from the hands and feet. S203, PAF calculation: Calculate the vector relationship between each pixel in the image and its neighboring pixels to determine the joint positions of the human body; S204, Pose Estimation: Perform full-body pose estimation, use the third CNN model to calculate the probability of the joint point corresponding to each pixel, calculate the position of each joint point based on the probability, and use geometric rules and thresholds to connect the positions of the joint points to form the full-body pose estimation result. S205. Output Results: Output the full-body pose estimation results as keypoint coordinates and joint angle information.

3. The method for detecting smoke based on a low-power GPU device according to claim 1, characterized in that, The method for determining smoking behavior in S5 is as follows: the smoking detection behavior model outputs the confidence level of the classification result of the current data collection in real time. The confidence level is a floating-point value between 0 and 1. When the confidence level is greater than the set threshold, the system determines the classification result as smoking.

4. A smoke detection device based on a low-power GPU, characterized in that, The method for detecting smoking based on a low-power GPU device as described in any one of claims 1-3 includes an RGBD camera, a low-power GPU device, and a smoking detection alarm connected in sequence.

Citation Information

Patent Citations

  • Daily smoking behavior detection method based on wearable equipment

    CN106056061A

  • Driver smoking detection method and system based on computer vision technology

    CN110738186A

  • A smoking detection method and device

    CN111914667B

  • Illegal smoking identification method based on AlphaPose

    CN112668387A

  • Driver smoking detection method based on multiple models

    CN113591615A