A Smoking and Phone Call Detection Method Based on Deep Learning and Behavioral Priors

Through detection methods based on deep learning and behavior priors, a deep convolutional neural network for multi-task object detection is trained, and combined with logical reasoning rules, the accuracy and reliability problems of smoking and calling behavior detection in the existing technology are solved, and efficient and real-time supervision is achieved in complex scenarios.

CN112883755BActive Publication Date: 2025-05-23WUHAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911196057.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-29
Publication Date
2025-05-23
Estimated Expiration
2039-11-29

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve all-round, real-time and effective supervision of smoking and calling behaviors. Especially in complex scenarios, deep learning models are prone to behavioral missed and mis-checked due to incomplete training set coverage and inconsistent annotation standards.

Method used

Using a detection method based on deep learning and behavioral priors, a multi-task object detection deep convolutional neural network is trained through offline processes. During the online process, the trained deep network model is used to perform forward inference on the input image or video frame, and logical reasoning rules are established based on the prior knowledge when the behavior occurs, and further determine whether smoking or calling behavior occurs.

Benefits of technology

It realizes the accuracy and reliability of smoking and calling behavior detection in practical applications, reduces missed and missed detection, and improves the credibility of security monitoring. Moreover, this method is easy to deploy and transform quickly to adapt to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112883755B_ABST
    Figure CN112883755B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting smoking and making phone calls based on deep learning and behavioral priors, which belongs to the field of safety supervision and image processing and analysis. The method includes two processes, offline and online: the offline process trains a multi-task target detection deep convolutional neural network through a self-built smoking and phone call behavior image data set, and the online process uses the trained deep network model to perform forward reasoning on the input image or video frame after face detection, first preliminarily predicts the label, confidence and location information of the smoking or phone call behavior, and also predicts the label, confidence and location information of specific targets related to these behaviors, such as human hands, cigarettes or mobile phones, etc., and then establishes logical reasoning rules between these information based on the prior knowledge when the behavior occurs, and further determines whether the smoking or phone call behavior occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of safety supervision and image processing and analysis, and specifically relates to a smoking and phone call detection method based on deep learning and behavioral priors. Background Art

[0002] In places such as gas stations, specific laboratories and factory sites, as well as when drivers are driving, smoking and making phone calls are strictly prohibited behaviors and are also behaviors that are monitored in safety management. Traditional video surveillance systems mainly rely on manual uninterrupted monitoring of human behavior in the surveillance screen or recording videos for post-playback to identify human behavior. Due to manpower limitations and low efficiency, it is difficult to achieve full and effective supervision of these strictly prohibited behaviors at any time. Intelligent analysis and detection based on machine vision technology has become a trend, which is more real-time and efficient than traditional manual video surveillance methods. The commonly used method is to first extract manually designed visual features from the collected video frames or images and then use classifiers to distinguish behaviors. Since the feature extraction algorithm is manually designed, it is not distinguishable enough, and the results of human behavior detection in complex actual scenarios are not reliable. In recent years, with the development of deep learning technology, people have begun to use deep convolutional neural networks to automatically learn visual features from a large amount of image data to characterize behavior, and realize end-to-end behavior detection. For example, Liu Xueqi et al. proposed an abnormal behavior detection method based on the YOLO network model (see "Electronic Design Engineering" Journal, Vol. 26, No. 20, pp. 154-158, 2018). Compared with traditional behavior detection methods, deep learning has shown great advantages in the field of behavior detection, but the effect of deep learning models depends largely on the training set. Since human behaviors such as smoking and making phone calls vary greatly in actual performance, it is difficult for general training sets to cover all situations. There are insufficient samples and uneven distribution. It is also difficult to have a unified standard for labeling behavior training sets, which easily leads to behaviors such as missed detection and false detection in end-to-end prediction methods such as deep learning. Summary of the invention

[0003] In order to overcome the shortcomings of the above-mentioned technology, the present invention provides a smoking and phone call detection method based on deep learning and behavior prior, which is characterized in that the method includes two processes, offline and online. The offline process trains a multi-task target detection deep convolutional neural network by self-built smoking and phone call behavior image data sets, and the online process uses the trained deep network model to perform forward reasoning on the input image or video frame after face detection, first preliminarily predicts the label, confidence and location information of the smoking or phone call behavior, and also predicts the label, confidence and location information of specific targets related to these behaviors, such as human hands, cigarettes or mobile phones, etc., and then establishes logical reasoning rules between these information based on the prior knowledge when the behavior occurs, and further determines whether the smoking or phone call behavior occurs.

[0004] Specifically, the present invention provides a smoking and phone call detection method based on deep learning and behavioral priors, and its offline process includes the following steps:

[0005] Step 1: Collect training videos or images, and use face detection methods to filter out video frames or images containing face information as valid training samples; Step 2: Label the filtered valid training samples, including labels and corresponding bounding box information for smoking, making phone calls or normal behaviors, as well as labels and corresponding bounding box information for targets related to smoking and making phone calls, such as hands, cigarettes or mobile phones; Step 3: Use data enhancement methods on the labeled samples to obtain more samples, which together form a training sample set; Step 4: Use all training samples and annotation information to train a multi-task target detection deep convolutional neural network based on the principles of deep learning.

[0006] In the above technical solution, the data collection method in step one is to record human behavior at different locations and lighting conditions indoors and outdoors, record videos of different people smoking or making phone calls, and also record some videos of not smoking or making phone calls as normal behavior samples; in addition, images downloaded from the Internet or images directly photographed of different behaviors can also be used as training data; in order to establish the correlation between behavior and people and consider the redundancy between consecutive video frames, the data screening method is to collect one frame of the video file every few frames and use the face detection algorithm to process it, and directly use the face detection algorithm to process the image file, and only retain those images in which faces can be detected as valid training samples.

[0007] In the above technical solution, the method for labeling the effective training samples in step 2 is: on the one hand, the behavior information is labeled, and a larger image area containing the face is framed as a behavior boundary box. When smoking and making phone calls occur, the corresponding labels are set to smoking and calling respectively, otherwise it is regarded as a normal behavior and the label is set to normal; on the other hand, the target information related to smoking and making phone calls is also labeled, that is, when targets such as human hands, cigarettes, and mobile phones appear in the image, their boundary boxes are marked, and the labels are set to hand, cigarette, and phone accordingly.

[0008] In the above technical solution, the data enhancement methods used in step three include image scaling, horizontal mirror flipping, random adjustment of brightness and hue, etc., keeping the label information of each behavior or target unchanged while updating the bounding box coordinate information according to the corresponding geometric transformation method.

[0009] In the above technical solution, the multi-task target detection network used in the step 4 can be modified based on the existing network structure in the field, such as Fast / Faster R-CNN, SSD or YOLO series, and share the backbone network structure to achieve simultaneous training of the behavior detection classifier and the corresponding target detection classifier. The behavior detection classifier is used to predict the label, confidence and location information of smoking, making a phone call or normal behavior, while the target detection classifier is used to predict the label, confidence and location information of a hand, a cigarette or a mobile phone. Here, the behavior detection problem is also regarded as a target detection problem, and the loss function form of the two tasks is the same during training.

[0010] The present invention provides a smoking and phone call detection method based on deep learning and behavioral priors, wherein the online process includes the following steps:

[0011] Step 1: For the input surveillance video or single image, use the face detection method to filter out the video frames or images containing face information as valid test samples; Step 2: Send the valid test samples to the multi-task target detection network trained in the offline process for forward reasoning, and predict the behavior, i.e. smoking, making a phone call or normal behavior, as well as the label, confidence and location information of the target related to the behavior, i.e. human hands, cigarettes or mobile phones; Step 3: Based on the prior knowledge when the behavior occurs, establish the logical reasoning rules between these predicted information to further determine whether the smoking or phone calling behavior occurs.

[0012] In the above technical solution, the face detection method used in the offline process is used in step 1, and the video frame or image containing face information is sent to the deep network model as a valid test sample for forward reasoning, and the position information of the face is recorded for logical reasoning in step 3;

[0013] In the above technical solution, in step 2, when using the trained deep network model for forward reasoning, the behavior label L, L∈{smoking, calling, normal}, the confidence level p are predicted for the valid test samples at the same time. 0 , location information (x, y, h, w), i.e., the horizontal coordinate, vertical coordinate, width and height of the center point of the behavior detection box, and the target label L′, L′∈{hand, cigarette, phone}, related to the behavior, and the confidence p 0 ′, position information (x′, y′, w′, h′), namely the horizontal coordinate, vertical coordinate, width and height of the center point of the target detection box.

[0014] In the above technical solution, the prior knowledge related to smoking and making phone calls used in step three includes: (1) the predicted behavior box should include the face area. For the situation where multiple people may appear in the image at the same time, the face included in the behavior box indicates that the behavior corresponds to the person; (2) when smoking or making phone calls occur in real life, the positional relationship between the face, hand, and object, i.e., cigarette or mobile phone, also has certain constraints. When the confidence corresponding to the behavior label predicted by the trained network model is low or the actual behavior is missed or misdetected, this constraint relationship can be used to establish a logical reasoning rule based on behavior priors to further perform behavior judgment.

[0015] Let Dist(face, object), Dist(hand, object) and Dist(face, hand) represent the distance between the face and the object, i.e., the cigarette or mobile phone, the distance between the hand and the object, i.e., the cigarette or mobile phone, and the distance between the face and the hand, respectively. The distance can be obtained by calculating the distance between the center points of the detection frame, and the possibility of smoking or making a phone call in the image is associated with these distance information. Since the absolute distance between pixels will change with the image scale, the side length of the detected square face frame Len(face) is used as the reference distance here, and the following rules are established:

[0016] (1) When Dist(face, object) ≤ a·Len(face), the confidence of smoking or making phone calls increases by p 1 ;

[0017] (2) When Dist(hand, object)≤b·Len(face), the confidence of smoking or making phone calls increases by p 2 ;

[0018] (3) When Dist(face, hand)≤c·Len((face)), the confidence of smoking or making phone calls increases by p3 ;

[0019] When determining the parameters a, b, and c, we can first perform statistical analysis on the labeled information of the training samples and then fine-tune based on human experience to determine the parameter p. 1 , p 2 , p 3 When, according to people's experience, according to the contribution degree p to the occurrence of smoking or phone calls 1 ≥p 2 >>p 3 ≥0, and the above three conditions are met at the same time p 1 +p 2 +p 3 =1;

[0020] When judging whether a certain behavior, such as smoking or making a phone call, occurs in an image, the label L is used to represent the behavior, and the confidence threshold of the occurrence of the behavior is represented by T. The behavior predicted by the target detection network and the labels, confidence, and location information of the related targets are processed according to different situations:

[0021] (1) When the detection result predicts a specific behavior label L and the confidence level p 0 Higher i.e. p 0 >T, directly determine that behavior L occurs;

[0022] (2) When the detection result predicts the behavior label L and the confidence level p 0 Lower that is p 0 ≤T, it is necessary to re-judge whether the behavior L occurs based on the distance relationship with the behavior-related target. The judgment rule is: calculate the distance information based on the relevant position information, determine whether the above three distance conditions are met, and obtain the confidence increase of the occurrence of behavior L as p 1 , p 2 , p 3 , then the confidence of behavior L is corrected to p 0 +p 1 +p 2 +p 3 , if the corrected confidence is higher than the threshold T, then it is determined that behavior L occurs, otherwise behavior L does not occur;

[0023] (3) When the detection result does not predict the behavior label L, p 0 = 0, it is also necessary to re-judge whether the behavior L occurs based on the distance relationship between the behavior-related target. The judgment rule is: calculate the distance information based on the relevant position information, judge whether the above three conditions are met, and obtain the confidence increase of the occurrence of behavior L as p 1 , p 2 , p 3, then the confidence of behavior L is calculated as p 1 +p 2 +p 3 , if the confidence is higher than the threshold T, it is determined that behavior L occurs, otherwise behavior L does not occur.

[0024] The present invention provides a method for detecting smoking and making phone calls based on deep learning and behavioral priors, which has the following beneficial effects: (1) The offline process is highly operable. For specific application scenarios, it can realize on-site video or image acquisition and timely model training, achieve rapid deployment, and be easy to promote and apply in actual systems; (2) The multi-task target detection model training is carried out using a deep learning method, which overcomes the limitation of the traditional method of manually extracted features with weak discrimination. At the same time, logical reasoning rules are established based on behavioral priors, and the results of the initial prediction of the deep network are further analyzed and reasoned, which is conducive to improving the behavior omissions and false detections that are easily caused by using a single behavior detection method based on a deep network, and helps to improve the credibility of security monitoring in actual behavior monitoring applications; (3) As long as data is re-collected and the model is trained according to the application scenario, and new behavioral prior logical reasoning rules are established, the method can be very conveniently modified to be promoted and applied to the detection of other human behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of the smoking and phone call detection method based on deep learning and behavioral priors of the present invention

[0026] Figure 2 This is the logical reasoning diagram of the smoking and phone call detection method based on deep learning and behavioral priors of the present invention. DETAILED DESCRIPTION

[0027] The following describes the implementation of the present invention in detail with reference to the accompanying drawings and examples, but the examples should not be construed as limiting the present invention.

[0028] See also Figure 1 The present invention provides a smoking and phone call detection method based on deep learning and behavior prior, which includes two processes, offline and online. The offline process trains a multi-task target detection deep convolutional neural network by using a self-built smoking and phone call behavior image dataset. The online process uses the trained deep network model to perform forward reasoning on the input image or video frame after face detection, first preliminarily predicts the label, confidence and location information of the smoking or phone call behavior, and also predicts the label, confidence and location information of specific targets related to these behaviors, such as human hands, cigarettes or mobile phones, etc., and then establishes logical reasoning rules between these information based on the prior knowledge when the behavior occurs, and further determines whether the smoking or phone call behavior occurs.

[0029] Specifically, the present invention provides a smoking and phone call detection method based on deep learning and behavioral priors, and its offline process includes the following steps:

[0030] Step 1: Collect training videos or images, and use face detection methods to filter out video frames or images containing face information as valid training samples; Step 2: Label the filtered valid training samples, including labels and corresponding bounding box information for smoking, making phone calls or normal behaviors, as well as labels and corresponding bounding box information for targets related to smoking and making phone calls, such as hands, cigarettes or mobile phones; Step 3: Use data enhancement methods on the labeled samples to obtain more samples, which together form a training sample set; Step 4: Use all training samples and annotation information to train a multi-task target detection deep convolutional neural network based on the principles of deep learning.

[0031] In the above technical solution, the data collection method in step one is to record human behavior at different locations and lighting conditions indoors and outdoors, record videos of different people smoking or making phone calls, and also record some videos of not smoking or making phone calls as normal behavior samples; in addition, images downloaded from the Internet or images directly photographed of different behaviors can also be used as training data; in order to establish the correlation between behavior and people and consider the redundancy between consecutive video frames, the data screening method is to collect one frame of the video file every few frames and use the face detection algorithm to process it, and directly use the face detection algorithm to process the image file, and only retain those images in which faces can be detected as valid training samples.

[0032] In the above technical solution, the method for labeling the effective training samples in step 2 is: on the one hand, the behavior information is labeled, and a larger image area containing the face is framed as a behavior boundary box. When smoking and making phone calls occur, the corresponding labels are set to smoking and calling respectively, otherwise it is regarded as a normal behavior and the label is set to normal; on the other hand, the target information related to smoking and making phone calls is also labeled, that is, when targets such as human hands, cigarettes, and mobile phones appear in the image, their boundary boxes are marked, and the labels are set to hand, cigarette, and phone accordingly.

[0033] In the above technical solution, the data enhancement methods used in step three include image scaling, horizontal mirror flipping, random adjustment of brightness and hue, etc., keeping the label information of each behavior or target unchanged while updating the bounding box coordinate information according to the corresponding geometric transformation method.

[0034] In the above technical solution, the multi-task target detection network used in the step 4 can be modified based on the existing network structure in the field, such as Fast / Faster R-CNN, SSD or YOLO series, and share the backbone network structure to achieve simultaneous training of the behavior detection classifier and the corresponding target detection classifier. The behavior detection classifier is used to predict the label, confidence and location information of smoking, making a phone call or normal behavior, while the target detection classifier is used to predict the label, confidence and location information of a hand, a cigarette or a mobile phone. Here, the behavior detection problem is also regarded as a target detection problem, and the loss function form of the two tasks is the same during training.

[0035] The present invention provides a smoking and phone call detection method based on deep learning and behavioral priors, wherein the online process includes the following steps:

[0036] Step 1: For the input surveillance video or single image, use the face detection method to filter out the video frames or images containing face information as valid test samples; Step 2: Send the valid test samples to the multi-task target detection network trained in the offline process for forward reasoning, and predict the behavior, i.e. smoking, making a phone call or normal behavior, as well as the label, confidence and location information of the target related to the behavior, i.e. human hands, cigarettes or mobile phones; Step 3: Based on the prior knowledge when the behavior occurs, establish the logical reasoning rules between these predicted information to further determine whether the smoking or phone calling behavior occurs.

[0037] In the above technical solution, the face detection method used in the offline process is used in step 1, and the video frame or image containing the face information is sent to the deep network model as a valid test sample for forward reasoning, and the position information of the face therein is recorded for the logical reasoning in step 3;

[0038] In the above technical solution, in step 2, when using the trained deep network model for forward reasoning, the behavior label L, L∈{smoking, calling, normal}, the confidence level p are predicted for the valid test samples at the same time. 0 , location information (x, y, h, w), i.e., the horizontal coordinate, vertical coordinate, width and height of the center point of the behavior detection box, and the target label L′, L′∈{hand, cigarette, phone}, related to the behavior, and the confidence p 0 ′, position information (x′, y′, w′, h′), namely the horizontal coordinate, vertical coordinate, width and height of the center point of the target detection box.

[0039] In the above technical solution, the prior knowledge related to smoking and making phone calls used in step three includes: (1) the predicted behavior box should include the face area. For the situation where multiple people may appear in the image at the same time, the face included in the behavior box indicates that the behavior corresponds to the person; (2) when smoking or making phone calls occur in real life, the positional relationship between the face, hand, and object, i.e., cigarette or mobile phone, also has certain constraints. When the confidence corresponding to the behavior label predicted by the trained network model is low or the actual behavior is missed or misdetected, this constraint relationship can be used to establish a logical reasoning rule based on behavior priors to further perform behavior judgment.

[0040] Let Dist(face, object), Dist(hand, object) and Dist(face, hand) represent the distance between the face and the object, i.e., the cigarette or mobile phone, the distance between the hand and the object, i.e., the cigarette or mobile phone, and the distance between the face and the hand, respectively. The distance can be obtained by calculating the distance between the center points of the detection frame, and the possibility of smoking or making a phone call in the image is associated with these distance information. Since the absolute distance between pixels will change with the image scale, the side length of the detected square face frame Len(face) is used as the reference distance here, and the following rules are established:

[0041] (1) When Dist(face, object) ≤ a·Len(face), the confidence of smoking or making phone calls increases by p 1 ;

[0042] (2) When Dist(hand, object)≤b·Len(face), the confidence of smoking or making phone calls increases by p 2 ;

[0043] (3) When Dist(face, hand) ≤ c·Len(face), the confidence of smoking or making phone calls increases by p 3 ;

[0044] When determining the parameters a, b, and c, we can first perform statistical analysis on the labeled information of the training samples and then fine-tune based on human experience to determine the parameter p. 1 , p 2 , p 3 When, according to people's experience, according to the contribution degree p to the occurrence of smoking or phone calls 1 ≥p 2 >>p 3 ≥0, and the above three conditions are met at the same time p 1 +p 2 +p 3 =1, for example, p1 =0.5, p 2 =0.4, p 3 =0.1.

[0045] like Figure 2 As shown in the figure, when it is necessary to determine whether a specific behavior, such as smoking or making a phone call, occurs in an image, the label L is used to represent the behavior, and T is used to represent the confidence threshold of the occurrence of the behavior. According to the behavior predicted by the target detection network and the label, confidence, and location information of the related target, the processing is carried out according to the situation:

[0046] (1) When the detection result predicts a specific behavior label L and the confidence level p 0 Higher i.e. p 0 >T, directly determine that behavior L occurs;

[0047] (2) When the detection result predicts the behavior label L and the confidence level p 0 Lower that is p 0 ≤T, it is necessary to re-judge whether the behavior L occurs based on the distance relationship between the behavior-related target and the judgment rule (i.e. Figure 2 Rule 2) is: Calculate the distance information based on the relevant position information, determine whether the above three distance conditions are met, and obtain the confidence increase of the behavior L as p 1 , p 2 , p 3 , then the confidence of behavior L is corrected to p 0 +p 1 +p 2 +p 3 , if the corrected confidence is higher than the threshold T, then it is determined that behavior L occurs, otherwise behavior L does not occur;

[0048] (3) When the detection result does not predict the behavior label L, p 0 = 0, it is also necessary to re-judge whether the behavior L occurs based on the distance relationship between the behavior-related target and the judgment rule (i.e. Figure 2 Rule 1) is: Calculate the distance information based on the relevant position information, determine whether the above three conditions are met, and obtain the confidence increase of behavior L as p 1 , p 2 , p 3 , then the confidence of behavior L is calculated as p 1 +p 2 +p 3 , if the confidence is higher than the threshold T, it is determined that behavior L occurs, otherwise behavior L does not occur.

[0049] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

[0050] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.

Claims

1. A smoking and phone call detection method based on deep learning and behavioral priors, It is characterized in that The method is divided into two processes, offline and online. The offline process trains a multi-task target detection deep convolutional neural network by using a self-built smoking and phone call behavior image dataset. The online process uses the trained deep network model to perform forward reasoning on the input image or video frame after face detection. It first preliminarily predicts the label, confidence and location information of the smoking or phone call behavior, and also predicts the label, confidence and location information of specific targets related to these behaviors, namely, hands, cigarettes or mobile phones. Then, based on the prior knowledge when the behavior occurs, logical reasoning rules between the above information are established to further determine whether the smoking or phone call behavior occurs; the prior knowledge when the behavior occurs includes: (1) the predicted behavior box should contain the face area, and (2) when the smoking or phone call behavior occurs, there is a distance constraint relationship between the positions of the face, hands, cigarettes or mobile phones.

2. A method for detecting smoking and making phone calls based on deep learning and behavioral priors according to claim 1, It is characterized in that The offline process includes the following four steps: step R1, collecting training videos or images, and using a face detection method to screen out video frames or images containing face information as valid training samples; step R2, labeling the screened valid training samples, including labels and corresponding bounding box information of smoking, making phone calls or normal behaviors, and labels and corresponding bounding box information of targets related to smoking and making phone calls, namely, hands, cigarettes or mobile phones; step R3, using data enhancement methods for the labeled samples to obtain more samples, which together form a training sample set; step R4, using all training samples and labeling information, based on the principle of deep learning, to train a multi-task target detection deep convolutional neural network.

3. The method for detecting smoking and making phone calls based on deep learning and behavioral priors according to claim 1, It is characterized in that In the offline process described in claim 2, the data collection method of step R1 is to record the behavior of people at different locations and lighting conditions indoors and outdoors, record the behavior videos of different people smoking or making phone calls, and also record some videos of not smoking and not making phone calls as normal behavior samples; in addition, images downloaded from the Internet or directly photographed images of different behaviors can also be used as training data; in order to establish the correlation between behavior and people and consider the redundancy between consecutive video frames, the data screening method is to collect one frame of the video file every few frames and use the face detection algorithm to process it, and directly use the face detection algorithm to process the image file, and only retain those images in which faces can be detected as effective training samples; In the offline process, the method of labeling the effective training samples in step R2 is: on the one hand, the behavior information is labeled, and a larger image area containing the face is framed as the behavior boundary box. When smoking and calling behaviors occur, the corresponding labels are set to smoking and calling respectively, otherwise it is regarded as a normal behavior and the label is set to normal; on the other hand, the target information related to smoking and calling is also labeled, that is, when a human hand, cigarette, or mobile phone target appears in the image, its boundary box is marked, and the labels are set to hand, cigarette, and phone accordingly; In the offline process, the data enhancement method used in step R3 includes image scaling, horizontal mirror flipping, random adjustment of brightness and hue, keeping the label information of each behavior or target unchanged while updating the bounding box coordinate information according to the corresponding geometric transformation method; In the offline process, the multi-task target detection network used in step R4 can be modified based on the existing network structure in the field, such as Fast / Faster R-CNN, SSD or YOLO series, and share the backbone network structure to achieve simultaneous training of the behavior detection classifier and the corresponding target detection classifier. The behavior detection classifier is used to predict the label, confidence and location information of smoking, making a phone call or normal behavior, while the target detection classifier is used to predict the label, confidence and location information of a hand, a cigarette or a mobile phone. Here, the behavior detection problem is also regarded as a target detection problem, and the loss function form of the two tasks is the same during training.

4. The method for detecting smoking and making phone calls based on deep learning and behavioral priors according to claim 1, It is characterized in that The online process includes the following three steps: Step S1: For the input surveillance video or single image, use the face detection method to filter out the video frames or images containing face information as valid test samples; Step S2: Send the valid test samples to the multi-task target detection network trained in the offline process for forward reasoning, and predict the behavior, i.e. smoking, making a phone call or normal behavior, as well as the label, confidence and location information of the target related to the behavior, i.e. human hand, cigarette or mobile phone; Step S3: Based on the prior knowledge when the behavior occurs, establish the logical reasoning rules between these predicted information to further determine whether the smoking or phone call behavior occurs.

5. The method for detecting smoking and making phone calls based on deep learning and behavioral priors according to claim 1, It is characterized in that In the online process described in claim 4, step S1 uses the same face detection method as the offline process, sends the video frame or image containing face information as a valid test sample to the deep network model for forward reasoning, and records the position information of the face therein for logical reasoning in step S3; In the online process, step S2 predicts the behavior label L, L∈{smoking, calling, normal}, the confidence level p, and the like for the valid test samples when using the trained deep network model for forward reasoning. 0 , location information (x, y, h, w), namely the horizontal coordinate, vertical coordinate, width and height of the center point of the behavior detection box, and the target label L related to the behavior ′ ,L ′ ∈{hand,cigarette,phone}, confidence p 0 ′ 、Location information(x ′ ,y ′ ,w ′ ,h ′ ) is the horizontal coordinate, vertical coordinate, width and height of the center point of the target detection frame; In the online process, step S3 uses the prior knowledge related to smoking and phone calls to establish a logical reasoning rule based on the prior knowledge of the behavior, and further conducts behavior judgment. The specific method is: Let Dist(face,object), Dist(hand,object) and Dist(face,hand) represent the distance between the face and the object, i.e., the cigarette or mobile phone, the distance between the hand and the object, i.e., the cigarette or mobile phone, and the distance between the face and the hand, respectively. The distance can be obtained by calculating the distance between the center points of the detection frame, and the possibility of smoking or making a phone call in the image is associated with these distance information. Since the absolute distance between pixels will change with the image scale, the side length of the detected square face frame Len(face) is used as the reference distance here to establish the following logical inference rules: (1) When Dist(face,object)≤a·Len(face), the confidence of smoking or making phone calls increases by p 1 ; (2) When Dist(hand,object)≤b·Len(face), the confidence of smoking or making phone calls increases by p 2 ; (3) When Dist(face, hand) ≤ c·Len(face), the confidence of smoking or making phone calls increases by p 3 ; When determining the parameters a, b, and c, we can first perform statistical analysis on the labeled information of the training samples and then fine-tune based on human experience to determine the parameter p. 1 ,p 2 ,p 3 When, according to people's experience, according to the contribution degree p to the occurrence of smoking or phone calls 1 ≥p 2 >>p 3 ≥0, and the above three conditions are met at the same time p 1 +p 2 +p 3 =1; When judging whether a certain behavior, such as smoking or making a phone call, occurs in an image, the label L is used to represent the behavior, and the confidence threshold of the occurrence of the behavior is represented by T. The behavior predicted by the target detection network and the labels, confidence, and location information of the related targets are processed according to different situations: (1) When the detection result predicts a specific behavior label L and the confidence level p 0 Higher i.e. p 0 >T, directly determine that behavior L occurs; (2) When the detection result predicts the behavior label L and the confidence level p 0 Lower that is p 0 ≤T, it is necessary to re-judge whether the behavior L occurs based on the distance relationship with the behavior-related target. The judgment rule is: calculate the distance information based on the relevant position information, determine whether the above three distance conditions are met, and obtain the confidence increase of the occurrence of behavior L as p 1 ,p 2 ,p 3 , then the confidence of behavior L is corrected to p 0 +p 1 +p 2 +p 3 , if the corrected confidence is higher than the threshold T, then it is determined that behavior L occurs, otherwise behavior L does not occur; (3) When the detection result does not predict the behavior label L, at this time p 0 = 0, then it is also necessary to re-determine whether the behavior L occurs based on the distance relationship between the behavior-related targets. The determination rule is: calculate the distance information according to the relevant position information, determine whether the above 3 conditions are satisfied, and obtain the confidence increase of the occurrence of behavior L as p 1 , p 2 , p 3 , then calculate the confidence of behavior L as p 1 + p 2 + p 3 . If this confidence is higher than the threshold T, it is determined that the behavior L occurs; otherwise, the behavior L does not occur.

Citation Information

Patent Citations

  • A sleep behavior detection method based on deep learning

    CN109472226A

  • Double-branch abnormity detection method based on crowd behavior priori knowledge

    CN110378233A