Multi-task cascade low-false-alarm smoking monitoring method

Through the multi-task cascaded smoking monitoring method, combined with head detection and smoking attribute recognition, the problems of high false alarm rate and poor scenario adaptability in the prior art are solved, and the smoking monitoring with low false alarm rate in diverse scenarios is realized, which improves detection accuracy and user experience.

CN119992446APending Publication Date: 2025-05-13NEWLAND DIGITAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411984245.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems with high false alarm rates and poor scenario adaptability in smoking monitoring, especially in diversified scenarios, resulting in low monitoring efficiency and poor user experience.

Method used

The low false alarm smoking monitoring method of multitasking cascade is adopted. Through monitoring image acquisition, head detection recognition, smoking attribute recognition and smoking detection, combined with the yolo data format and MobileNetv4 architecture, head detection and smoking attribute recognition are carried out to reduce the false alarm rate and improve detection accuracy.

Benefits of technology

Smoking monitoring with low false alarm rate in diverse scenarios is achieved, which improves monitoring accuracy and adaptability, reduces false alarm rate, and improves user experience and monitoring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992446A_ABST
    Figure CN119992446A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task cascaded low-false-alarm smoking monitoring method. The method specifically comprises the following steps: acquiring a monitoring image; performing head detection and recognition to obtain head coordinates in the image; utilizing the obtained head coordinates to expand and cut out a head sub-graph, and carrying out smoking attribute identification to obtain smoking attribute judgment of each head; performing external expansion matting on the human head area with the attribute judged to be smoking, performing smoking detection, and if the confidence coefficient of the cigarette is detected to be greater than a threshold value, judging that the frame has a smoking behavior; and comprehensively judging whether a person smokes or not according to continuous multi-frame results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applied to the field of video monitoring and image processing, and specifically is a multi-task cascade low false alarm smoking monitoring method. Background Art

[0002] As global attention to public health issues increases, the monitoring and management of smoking behavior has become an important issue that governments and public health agencies need to address urgently. According to statistics from the World Health Organization (WHO), smoking causes more than 8 million deaths worldwide each year, including about 1.2 million non-smokers who die from secondhand smoke exposure. Secondhand smoke not only poses a serious threat to children and pregnant women, but also increases the risk of chronic diseases such as cardiovascular disease and cancer in adults. In addition, with the continuous increase in safety needs, identifying and preventing potentially dangerous behaviors in key areas has become an important issue. Among them, smoking is strictly prohibited in certain specific places (such as gas stations, warehouses, hospitals, etc.). Therefore, reducing the prevalence of smoking and protecting public health and safety has become a top priority.

[0003] Smoking bans in public places have been widely implemented in many countries. Although the existing smoking bans have reduced the smoking rate to a certain extent, there are still many challenges in the implementation process, which are mainly reflected in the following aspects: First, the limitations of manual monitoring: Traditional methods that rely on manual supervision are easily affected by human factors, resulting in low monitoring efficiency and poor consistency. Especially during peak hours, monitoring personnel may not be able to provide comprehensive coverage. Second, the shortcomings of sensor technology: existing smoke sensors can usually only detect the concentration of smoke in the air and cannot distinguish between smoking and other smoke sources (such as barbecues, fires, etc.). This is particularly evident in complex public places, resulting in high false alarm rates and false negatives. Third, user experience issues: high false alarm rates will lead to dissatisfaction among those being monitored, affecting the smooth implementation of policies. In addition, frequent false alarms may also lead to a decrease in trust in the monitoring system.

[0004] In order to meet the above challenges, research in recent years has gradually turned to the use of advanced computer vision and machine learning technologies. Although existing studies have made some progress in smoking detection, when faced with diverse scenarios, the direct use of target detection still has problems such as high false detection rate and poor scene adaptability due to the small proportion of cigarettes in the monitoring screen. For example, some actions similar to smoking behaviors (such as drinking water, hands close to the mouth) will cause misjudgment, and some white strip-like objects similar to cigarettes (such as toothpicks, straws, sticks) will cause misjudgment, especially the smoking attribute model will be more sensitive to white strip-like objects. Therefore, a new multi-task cascaded low false alarm smoking monitoring algorithm is proposed to achieve high accuracy and low false alarm smoking behavior detection, and promote public health management. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a multi-task cascade low false alarm smoking monitoring method in view of the deficiencies in the prior art.

[0006] In order to solve the above technical problems, a multi-task cascade low false alarm smoking monitoring method of the present invention specifically includes the following steps:

[0007] Monitoring image acquisition;

[0008] Perform head detection and recognition to obtain the head coordinates in the image;

[0009] Using the acquired head coordinates to expand and cut out the head sub-image, and perform smoking attribute recognition to obtain the smoking attribute determination of each head;

[0010] The head area judged as smoking is expanded and cut out, and smoking detection is performed. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame;

[0011] Determine whether someone is smoking based on the results of multiple consecutive frames.

[0012] As a possible implementation, further, the step of performing head detection and recognition to obtain the head coordinates in the image specifically includes:

[0013] In order to adapt to different scenes, weather, and lighting, we use public head datasets and self-collected data to construct head detection training sets, and generate yolo data format for training;

[0014] Replace the backbone of yolov11 with MobileNetv4 for training;

[0015] After the training, conduct another self-distillation training;

[0016] During the test, the surveillance video image is input into the network. If the confidence is greater than the threshold, the candidate box is considered to be a head detection box.

[0017] As a possible implementation, further, the steps of using the acquired head coordinates to expand and cut out head sub-images, and performing smoking attribute recognition to obtain the smoking attribute determination step of each head include:

[0018] Collect public data and self-collected data for annotation, use polygons to annotate the cigarette area, and generate corresponding annotation mask images;

[0019] Filter out the head area, expand the head area N times and cut out the image. The corresponding Mask image needs to cut out the same position as the training data. If the corresponding position contains a cigarette, its label is smoking, otherwise it is non-smoking;

[0020] The original MobileNetv4 only supports classification. On this basis, we add branches to MobileNetv4 to predict both the smoking category and the cigarette segmentation map during training. If the input is a C×H×W image, the cigarette segmentation map is of size H / 2×W / 2, and the training loss is: Loss = L cls +γ mask L mask ;

[0021] Among them, γ mask is the weight coefficient, L cls is the smoking classification loss, L cls =-logp k , sample laebl is the kth class, L mask is the cigarette segmentation loss, in Where K is the number of pixels H / 2×W / 2 on the cigarette segmentation map, p is the confidence value of the cigarette segmentation map, α is the balance factor which can be 0.25, and β is the focus factor;

[0022] During the test, the candidate box of the head detection is expanded N times and input for smoking attribute recognition. If its confidence value is greater than the threshold, it is considered that the person is smoking. Smoking attribute recognition can provide clearer target information for subsequent detection.

[0023] As a possible implementation, further, the head region of the person whose attribute is determined to be smoking is expanded and cut out, and smoking detection is performed. If the confidence of the detected cigarette is greater than a threshold, the step of determining that there is smoking behavior in this frame specifically includes:

[0024] Collect public data and self-collected data, use the coordinates of the smoking head obtained by head detection and smoking attribute recognition, expand the detection box to cut out the image and use it together with the original large image as the training set;

[0025] Replace the backbone of yolov11 with MobileNetv4 for training, change the training task from smoking detection alone to smoking detection plus cigarette segmentation, and change its IoU loss to focaler-IoU loss, improving performance by focusing on regression samples of different difficulty levels;

[0026] The focaler-IoU loss is as follows:

[0027] Where d is the lower threshold, μ is the upper threshold, B is the predicted target box, and B gt is the real target frame;

[0028] During the test, the coordinates of the smoking head obtained by head detection and smoking attribute recognition are used to expand the detection frame and perform cutout input for smoking detection. If the confidence value is greater than the threshold, a cigarette detection frame is obtained. If the obtained cigarette detection frame is within the detection frame input into the smoking attribute recognition module of Task 2, it is considered that the person is smoking, otherwise he is not smoking.

[0029] As a possible implementation, further, the step of comprehensively determining whether someone is smoking based on the continuous multi-frame results includes: fusing the multi-frame results as a judgment, and reporting only when the number of frames with the same result is set to be greater than a preset frame number threshold; if there is no continuous multi-frame input, directly using the single-frame result as the final result.

[0030] A multi-task cascade low false alarm smoking monitoring system, specifically comprising:

[0031] Image acquisition module, monitoring image acquisition;

[0032] A head coordinate acquisition module performs head detection and recognition to obtain the head coordinates in the image;

[0033] The smoking attribute recognition module uses the acquired head coordinates to expand and cut out the head sub-image, and performs smoking attribute recognition to obtain the smoking attribute judgment of each head;

[0034] The smoking detection module expands and cuts out the head area of ​​the person whose attribute is determined to be smoking, and performs smoking detection. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame;

[0035] The multi-frame judgment module comprehensively determines whether someone is smoking based on the results of multiple consecutive frames.

[0036] The present invention adopts the above technical solution and has the following beneficial effects:

[0037] 1. Able to handle real-time processing of embedded devices.

[0038] 2. Use multi-task cascade method to further reduce the false alarm rate of smoking monitoring

[0039] 3. This solution can be used to adapt to both indoor and outdoor scenes, which can greatly improve the shortcomings of smoke sensors that are only suitable for indoor use. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0041] Figure 1 This is a simplified structural flow chart of the smoking monitoring system of the present invention;

[0042] Figure 2 A Mask diagram is marked for the embodiment of the present invention;

[0043] Figure 3 This is a simplified diagram of smoking attribute recognition according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0045] Example 1

[0046] The present invention provides a multi-task cascade low false alarm smoking monitoring method, which specifically includes the following steps:

[0047] Monitoring image acquisition;

[0048] Perform head detection and recognition to obtain the head coordinates in the image;

[0049] Using the acquired head coordinates to expand and cut out the head sub-image, and perform smoking attribute recognition to obtain the smoking attribute determination of each head;

[0050] The head area judged as smoking is expanded and cut out, and smoking detection is performed. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame;

[0051] Determine whether someone is smoking based on the results of multiple consecutive frames.

[0052] The steps of performing head detection and recognition to obtain the head coordinates in the image specifically include:

[0053] In order to adapt to different scenes, weather, and lighting, we use public head datasets and self-collected data to construct head detection training sets, and generate yolo data format for training;

[0054] Replace the backbone of yolov11 with MobileNetv4 for training;

[0055] After the training, conduct another self-distillation training;

[0056] During the test, the surveillance video image is input into the network. If the confidence is greater than the threshold, the candidate box is considered to be a head detection box.

[0057] The steps of using the acquired head coordinates to expand and cut out the head sub-image, and performing smoking attribute recognition to obtain the smoking attribute determination steps of each head include:

[0058] Collect public data and self-collected data for annotation, use polygons to annotate the cigarette area, and generate corresponding annotation mask images;

[0059] Filter out the head area, expand the head area N times and cut out the image. The corresponding Mask image needs to cut out the same position as the training data. If the corresponding position contains a cigarette, its label is smoking, otherwise it is non-smoking;

[0060] The original MobileNetv4 only supports classification. On this basis, we add branches to MobileNetv4 to predict both the smoking category and the cigarette segmentation map during training. If the input is a C×H×W image, the cigarette segmentation map is of size H / 2×W / 2, and the training loss is: Loss = L cls +γ mask L mask ;

[0061] Among them, γ mask is the weight coefficient, L cls is the smoking classification loss, L cls =-logp k , sample laebl is the kth class, L mask is the cigarette segmentation loss, in Where K is the number of pixels H / 2×W / 2 on the cigarette segmentation map, p is the confidence value of the cigarette segmentation map, α is the balance factor which can be 0.25, and β is the focus factor;

[0062] During the test, the candidate box of the head detection is expanded N times and input for smoking attribute recognition. If its confidence value is greater than the threshold, it is considered that the person is smoking. Smoking attribute recognition can provide clearer target information for subsequent detection.

[0063] The steps of expanding and cutting out the head region of the person whose attribute is determined to be smoking and performing smoking detection, and if the confidence of the detected cigarette is greater than a threshold, determining that there is smoking behavior in this frame specifically include:

[0064] Collect public data and self-collected data, use the coordinates of the smoking head obtained by head detection and smoking attribute recognition, expand the detection box to cut out the image and use it together with the original large image as the training set;

[0065] Replace the backbone of yolov11 with MobileNetv4 for training, change the training task from smoking detection alone to smoking detection plus cigarette segmentation, and change its IoU loss to focaler-IoU loss, improving performance by focusing on regression samples of different difficulty levels;

[0066] The focaler-IoU loss is as follows:

[0067] Where d is the lower threshold, μ is the upper threshold, B is the predicted target box, and B gt is the real target frame;

[0068] During the test, the coordinates of the smoking head obtained by head detection and smoking attribute recognition are used to expand the detection frame and perform cutout input for smoking detection. If the confidence value is greater than the threshold, a cigarette detection frame is obtained. If the obtained cigarette detection frame is within the detection frame input into the smoking attribute recognition module of Task 2, it is considered that the person is smoking, otherwise he is not smoking.

[0069] Among them, the step of comprehensively judging whether someone is smoking based on the continuous multi-frame results includes: fusing the multi-frame results as a judgment, and reporting only when the number of frames with the same result is greater than a preset frame number threshold; if there is no continuous multi-frame input, directly using the single-frame result as the final result.

[0070] A multi-task cascade low false alarm smoking monitoring system, such as Figure 1 Specifically shown include:

[0071] Image acquisition module, monitoring image acquisition;

[0072] A head coordinate acquisition module performs head detection and recognition to obtain the head coordinates in the image;

[0073] The smoking attribute recognition module uses the acquired head coordinates to expand and cut out the head sub-image, and performs smoking attribute recognition to obtain the smoking attribute judgment of each head;

[0074] The smoking detection module expands and cuts out the head area of ​​the person whose attribute is determined to be smoking, and performs smoking detection. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame;

[0075] The multi-frame judgment module comprehensively determines whether someone is smoking based on the results of multiple consecutive frames.

[0076] Example 2

[0077] The present invention is a multi-task cascade low false alarm smoking monitoring algorithm, which specifically includes the following:

[0078] Monitoring image acquisition

[0079] Obtain the coordinates of all heads in the image through the task head detection module

[0080] Expand the head coordinates obtained from Task 1 to extract all head sub-images, and input them into the smoking attribute recognition module of Task 2 to obtain the smoking attribute of each head.

[0081] The head area of ​​the person whose attribute is determined to be smoking is expanded and cut out, and then input into the smoking detection module of Task 3. If the confidence of the detected cigarette is greater than the threshold, it is considered that someone is smoking in this frame

[0082] Comprehensively determine whether someone is smoking based on the results of multiple consecutive frames

[0083] The following is a detailed description of the specific detection algorithm and its multi-task modules:

[0084] Monitoring image acquisition module:

[0085] Install the camera in a smoking-free area or designated smoking area (either indoors or outdoors), adjust the monitoring equipment to cover the entire area as much as possible, and allow clear viewing of smoking.

[0086] Task: Head detection module

[0087] When a user smokes, the cigarette is usually near the head and the cigarette is usually small. Directly using the smoking detection model is likely to cause false alarms. The task head detection module can be used to filter out the areas that need further judgment.

[0088] The task here is to use a head detection module to train a head detection model using a deep learning method. The current mainstream target detection algorithms mainly include the following categories: two-stage detection algorithms (such as the R-CNN series), single-stage detection algorithms (such as YOLO, SSD) and Transformer-based detection algorithms (such as DETR). Considering the real-time performance, we use the latest detection model of the current YOLO series, yolov11, for improvement and training. The steps are as follows:

[0089] In order to adapt to different scenes, weather, and lighting, we used public head datasets such as CrowdHuman, HollywoodHeads, and CityPersons and self-collected data to construct a head detection training set, and generated the yolo data format for training.

[0090] MobileNetv4 is a general and efficient architecture designed for mobile devices. We replace the backbone of yolov11 with MobileNetv4 for training to achieve a balance between speed and accuracy.

[0091] After the training, perform another self-distillation training to improve performance

[0092] During the test, the surveillance video image can be input into the network. If the confidence is greater than the threshold (which can be set to 0.6), the candidate box is considered to be a head detection box.

[0093] Task 2: Smoking attribute recognition module

[0094] After the head area is screened out, the smoking attribute can be identified to determine whether the person is smoking. The smoking attribute recognition module of Task 2 here is also trained using deep learning methods. MobileNetv4 can be used for training here. The steps are as follows:

[0095] Collect public data and self-collected data for annotation, use polygons to annotate the cigarette area, and generate a corresponding annotation mask map, such as Figure 2 As shown, the cigarette part is assigned a value of 255.

[0096] 3.2 Use the head detection module to filter out the head area, expand the head area by N times (1.5 is used here) for cropping, and the corresponding Mask image needs to be cropped out of the same position as training data. If the corresponding position contains a cigarette, its label is smoking, otherwise it is not smoking.

[0097] The original MobileNetv4 only supports classification. On this basis, we add branches to MobileNetv4 to predict both the smoking category and the cigarette segmentation map during training. If the input is a C×H×W image, the cigarette segmentation map is of size H / 2×W / 2, and the training loss is Loss=L cls +γ mask L mask , where γ mask is the weight coefficient, which can be 0.5, L cls is the smoking classification loss, L cls =-logp k , sample laebl is the kth class, L mask is the cigarette segmentation loss, K is the number of pixels H / 2×W / 2 on the cigarette segmentation map, p is the confidence value of the cigarette segmentation map, α is the balance factor, which can be 0.25, and β is the focus factor, which can be 2. The mask cigarette segmentation branch helps to determine the smoking category. In actual testing, the auxiliary mask branch will be trimmed and only the classification part, that is, the dotted box part in the figure below, will be used.

[0098] 3.4 During the test, the candidate box of the head detection is expanded N times (1.5 times is proposed here) and input into the task 2 smoking attribute recognition module. If its confidence value is greater than the threshold, it is considered that the person is smoking. The task 2 smoking attribute recognition module can provide more clear target information for subsequent detection, thereby improving the accuracy of detection.

[0099] 4. Task 3: Smoking Detection Module

[0100] Because the smoking attribute module is used directly, it is easy to misdetect white strip-like objects similar to cigarettes, such as straws. In order to improve its accuracy and reduce false detections, the detection frame of the head that is judged as smoking is expanded (here it is planned to be expanded by 2 times) to capture the contextual information around the target, thereby helping the model to better understand the relationship between the target and the environment, so as to eliminate false positives such as straws and white strips.

[0101] The steps are as follows:

[0102] 4.1 Collect public data and self-collected data, use the smoker head coordinates obtained by the head detection module of task one and the smoking attribute recognition module of task two, expand the detection box (here 3 is proposed to be used) to cut out the image and use it together with the original large image as the training set.

[0103] 4.2 Replace the backbone of yolov11 with MobileNetv4 for training, and change the training task from smoking detection alone to smoking detection plus cigarette segmentation, and change its IoU loss to focaler-IoU loss, by focusing on regression samples of different difficulty levels to improve performance.

[0104] The focaler-IoU loss is as follows:

[0105] Where d is the lower threshold value which can be 0.1, μ is the upper threshold value which can be 0.95, B is the predicted target box, B gt is the real target box

[0106] 4.3 During the test, the coordinates of the smoking head obtained by the head detection module of Task 1 and the smoking attribute recognition module of Task 2 are used to expand the detection frame input into Task 2 (here it is planned to be 2 times) for clipping and input into the smoking detection module of Task 3. If the confidence value is greater than the threshold, the cigarette detection frame is obtained. If the obtained cigarette detection frame is within the detection frame input into the smoking attribute recognition module of Task 2, it is considered that this person smokes, otherwise it is considered that he does not smoke.

[0107] 5.Multi-frame judgment module

[0108] In order to make the result more accurate, multiple frame results can be fused as judgment. Here, continuous multiple frame results (3 frames are used here) are directly used to report when smoking. If there is no continuous multiple frame input, the single frame result is directly used as the final result.

[0109] The above are embodiments of the present invention. For ordinary technicians in this field, according to the teachings of the present invention, all equivalent changes, modifications, substitutions and variations made within the scope of the patent application of the present invention without departing from the principles and spirit of the present invention should fall within the scope of the present invention.

Claims

1. A multi-task cascade low false alarm smoking monitoring method, characterized in that: The specific steps include: Monitoring image acquisition; Perform head detection and recognition to obtain the head coordinates in the image; Using the acquired head coordinates to expand and cut out the head sub-image, and perform smoking attribute recognition to obtain the smoking attribute determination of each head; The head area judged as smoking is expanded and cut out, and smoking detection is performed. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame; Determine whether someone is smoking based on the results of multiple consecutive frames.

2. A multi-task cascade low false alarm smoking monitoring method according to claim 1, characterized in that: The step of performing head detection and recognition to obtain the head coordinates in the image specifically includes: In order to adapt to different scenes, weather, and lighting, we use public head datasets and self-collected data to construct head detection training sets, and generate yolo data format for training; Replace the backbone of yolov11 with MobileNetv4 for training; After the training, conduct another self-distillation training; During the test, the surveillance video image is input into the network. If the confidence is greater than the threshold, the candidate box is considered to be a head detection box.

3. The multi-task cascade low false alarm smoking monitoring method according to claim 1, characterized in that: The step of using the acquired head coordinates to expand and cut out a head sub-image, and performing smoking attribute recognition to obtain a smoking attribute determination step for each head includes: Collect public data and self-collected data for annotation, use polygons to annotate the cigarette area, and generate corresponding annotation mask images; Filter out the head area, expand the head area N times and cut out the image. The corresponding Mask image needs to cut out the same position as the training data. If the corresponding position contains a cigarette, its label is smoking, otherwise it is non-smoking; The original MobileNetv4 only supports classification. On this basis, we add branches to MobileNetv4 to predict both the smoking category and the cigarette segmentation map during training. If the input is a C×H×W image, the cigarette segmentation map is of size H / 2×W / 2, and the training loss is: Loss = L cls +γ mask L mask ; Among them, γ mask is the weight coefficient, L cls is the smoking classification loss, L cls =-logp k , sample laebl is the kth class, L mask is the cigarette segmentation loss, in Where K is the number of pixels H / 2×W / 2 on the cigarette segmentation map, p is the confidence value of the cigarette segmentation map, α is the balance factor which can be 0.25, and β is the focus factor; During the test, the candidate box of the head detection is expanded N times and input for smoking attribute recognition. If its confidence value is greater than the threshold, it is considered that the person is smoking. Smoking attribute recognition can provide clearer target information for subsequent detection.

4. The multi-task cascade low false alarm smoking monitoring method according to claim 1, characterized in that: The step of expanding and cutting out the head region of the person whose attribute is determined to be smoking, and performing smoking detection, and if the confidence of the detected cigarette is greater than a threshold, determining that there is smoking behavior in this frame specifically includes: Collect public data and self-collected data, use the coordinates of the smoking head obtained by head detection and smoking attribute recognition, expand the detection box to cut out the image and use it together with the original large image as the training set; Replace the backbone of yolov11 with MobileNetv4 for training, change the training task from smoking detection alone to smoking detection plus cigarette segmentation, and change its IoU loss to focaler-IoU loss, improving performance by focusing on regression samples of different difficulty levels; The focaler-IoU loss is as follows: Where d is the lower threshold, μ is the upper threshold, B is the predicted target box, and B gt is the real target frame; During the test, the coordinates of the smoking head obtained by head detection and smoking attribute recognition are used to expand the detection frame and perform cutout input for smoking detection. If the confidence value is greater than the threshold, a cigarette detection frame is obtained. If the obtained cigarette detection frame is within the detection frame input into the smoking attribute recognition module of Task 2, it is considered that the person is smoking, otherwise he is not smoking.

5. The multi-task cascade low false alarm smoking monitoring method according to claim 1, characterized in that: The step of comprehensively judging whether someone is smoking based on the continuous multi-frame results includes: fusing the multi-frame results as a judgment, and reporting only when the number of frames with the same result is greater than a preset frame number threshold; if there is no continuous multi-frame input, directly using the single-frame result as the final result.

6. A multi-task cascade low false alarm smoking monitoring system, characterized in that: Specifically include: Image acquisition module, monitoring image acquisition; A head coordinate acquisition module performs head detection and recognition to obtain the head coordinates in the image; The smoking attribute recognition module uses the acquired head coordinates to expand and cut out the head sub-image, and performs smoking attribute recognition to obtain the smoking attribute judgment of each head; The smoking detection module expands and cuts out the head area of ​​the person whose attribute is determined to be smoking, and performs smoking detection. If the confidence of the detected cigarette is greater than the threshold, it is determined that there is smoking behavior in this frame; The multi-frame judgment module comprehensively determines whether someone is smoking based on the results of multiple consecutive frames.

Citation Information

Cited By

  • Coal mining machine roller tracking system and method

    CN121010931A