Smoking detection method and device and computer program product
By using the pre-trained object detection model and feature fusion module, real-time scene images in complex scenarios are analyzed, and the problem of poor accuracy and real-time accuracy of smoking detection results in the prior art is solved, and more efficient smoking behavior detection is achieved.
Patent Information
- Application Number
- CN202510280537.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing smoking detection technology has poor accuracy and real-time accuracy in complex scenarios.
The pre-trained object detection model is used to analyze the real-time scene images, and the convolution kernel adjustment module based on the adaptive mechanism, the context-aware module and the feature fusion module based on the feature pyramid network are determined whether the image contains smoking behavior.
It improves the accuracy and real-timeness of smoking behavior detection, enhances the robustness of the model in complex scenarios, and reduces resource consumption.
Smart Images

Figure CN120198958A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular, to a smoking detection method, an apparatus, and a computer program product. Background Art
[0002] Currently, for the detection of smoking behavior, it is usually to install a smoke sensor under a monitoring camera or use image processing technology to detect the presence of smoke to determine whether there is a smoking behavior. Specifically: Installing a smoke sensor is to detect whether smoke exists. If it exists, an alarm is issued and relevant events are recorded. The accuracy of this detection result depends on the accuracy and sensitivity of the smoke sensor, and the accuracy and sensitivity of the smoke sensor may be interfered by external environmental factors such as the position of the camera and ambient light, resulting in a high false alarm rate and low accuracy; while image processing technology is to use image processing algorithms to analyze the video frames captured by the monitoring camera to detect whether there is smoke in the image. Although this method can achieve real-time monitoring of smoke, the detection accuracy in complex scenarios is relatively low, and the algorithm complexity is relatively high, with insufficient real-time performance, unable to meet the requirements of real-time monitoring of smoke.
[0003] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide a smoking detection method, an apparatus, and a computer program product to at least solve the technical problems of poor accuracy and real-time performance of smoking behavior detection results in complex scenarios in related smoking detection technologies.
[0005] According to one aspect of the embodiments of the present application, a smoking detection method is provided, including: obtaining a real-time scene image in a target area; using a pre-trained object detection model to analyze the real-time scene image to obtain whether the real-time scene image contains a smoking behavior, where the object detection model is used to determine whether the real-time scene image contains a smoking behavior according to the probability distribution of smoking targets related to smoking behavior included in multiple detection frames in the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and the smoking actions of smokers, and the object detection model at least includes: a convolution kernel adjustment module based on an adaptive mechanism, a context awareness module, and a feature fusion module based on a feature pyramid network.
[0006] Optionally, the training process of the object detection model includes: obtaining multiple sets of training samples, where each set of training samples includes: an image sample, a sample label corresponding to the image sample, and the position information and size of the true detection box of the smoking target related to the smoking behavior contained in the image sample, and the sample label is used to represent whether the image sample contains the smoking behavior; constructing a deep convolutional neural network model based on the YOLOv5 architecture; and iteratively training the deep convolutional neural network model with multiple sets of training samples to obtain the object detection model.
[0007] Optionally, the deep convolutional neural network model further includes: a static convolution module including multiple convolutional kernels with fixed sizes and strides, a dynamic convolution module including multiple convolutional kernels with dynamic sizes and strides, a classification module, and a regression module. Among them, iteratively training the deep convolutional neural network model with multiple sets of training samples to obtain the object detection model includes: for each image sample, calling the static convolution module to perform a convolution operation on the image sample to obtain a corresponding initial feature map; calling the convolution kernel adjustment module to determine the sizes and strides of the convolutional kernels in the dynamic convolution module according to the size of the image sample and the size of the initial feature map; calling the dynamic convolution module to perform a convolution operation on the initial feature map according to the adjusted sizes and strides of the convolutional kernels to obtain multiple feature maps with different depths, and forming a feature pyramid network from the multiple feature maps, where the depth corresponding to each layer of the feature map in the feature pyramid network decreases layer by layer from top to bottom; calling the feature fusion module based on the feature pyramid network to fuse the feature maps in the feature pyramid network to obtain a target fusion feature map; calling the context awareness module to extract the context feature map of the image sample and fuse the context feature map with the target fusion feature map to obtain a target enhanced fusion feature map; calling the classification module and the regression module to analyze the target enhanced fusion feature map respectively to obtain the predicted label corresponding to the image sample and the position information and size of the predicted anchor box of the image sample; constructing a classification loss function based on the sample label and the predicted label corresponding to each image sample, and constructing a regression loss function based on the position information and size of the true detection box in each training sample and the position information and size of the corresponding predicted anchor box, and constructing an object loss function based on the classification loss function and the regression loss function; calculating the gradient of the object loss function with respect to the model parameters, and updating the model parameters based on the gradient until the model converges to obtain the trained object detection model.
[0008] Optionally, calling the convolution kernel adjustment module to determine the sizes and strides of the convolutional kernels in the dynamic convolution module according to the size of the image sample and the size of the initial feature map includes: calling the convolution kernel adjustment module to determine the sizes and strides of the convolutional kernels in the dynamic convolution module according to the size of the image sample and the size of the initial feature map according to the following formula:
[0009]
[0010] Where K d represents the size of the adjusted convolutional kernel, and S d represents the adjusted stride, and K b represents the size of the convolutional kernel before adjustment, and S b represents the stride before adjustment, and S img represents the size of the image sample, and S feat represents the size of the initial feature map.
[0011] Optionally, call the feature fusion module based on the feature pyramid network to fuse each feature map in the feature pyramid network to obtain the target fusion feature map, including: for each feature layer in the feature pyramid network, call the feature fusion module based on the feature pyramid network to upsample the feature map in the feature layer to match its resolution with that of the shallower feature map with a lower depth, and perform weighted fusion on the feature map and the adjacent shallower feature map to obtain the corresponding fusion feature map; determine all the feature values in the fusion feature map within the feature layer, determine the quality score of the fusion feature map based on all the feature values, and determine the weight of the feature layer based on the quality score; calculate the target fusion feature map according to the following formula based on the quality score of the fusion feature map in each feature layer and the weight of each feature layer:
[0012]
[0013] Where F feature represents the target fusion feature map, M represents the total number of layers of the feature pyramid network, and W i represents the weight of the i-th layer fusion feature map in the feature pyramid, and F i represents the i-th layer fusion feature map in the feature pyramid.
[0014] Optionally, perform weighted fusion on the feature map and the adjacent shallower feature map to obtain the corresponding fusion feature map, including: perform weighted fusion on the feature map and the adjacent shallower feature map according to the following formula to obtain the corresponding fusion feature map:
[0015]
[0016] Where F i represents the i-th layer fusion feature map in the feature pyramid network, F s represents the shallower feature map below the i-th layer feature map in the feature pyramid network, F d represents the i-th layer feature map, α, β represent the weighting coefficients, and β = 1 - α, where γ represents the adjustment parameter, and S fs represents the statistical information of the shallower feature map.
[0017] Optionally, determine the quality score of the fused feature map based on all feature values, and determine the weight of the feature layer based on the quality score, including: based on all feature values, use the following formula to determine the quality score of the fused feature map:
[0018]
[0019] In the formula, Q i represents the quality score of the i-th layer fused feature map in the feature pyramid network, F i,j represents the j-th feature value in the i-th layer fused feature map in the feature pyramid network, N represents the total number of feature values in the fused feature map, and Var(·) represents the variance of the pixel values in the i-th layer fused feature map in the feature pyramid network; based on the quality score of the fused feature map, calculate the weight of the feature layer according to the following formula:
[0020]
[0021] In the formula, W i represents the weight of the i-th layer fused feature map in the feature pyramid network, and Q i represents the quality score of the i-th layer fused feature map in the feature pyramid network.
[0022] Optionally, call the context-aware module to extract the context feature map of the image sample, and fuse the context feature map with the target fused feature map to obtain the target enhanced fused feature map, including: call the context-aware module, and extract the feature map corresponding to the context information in the image sample according to the following formula to obtain the context feature map:
[0023] F context = ContextModule(F feature , ContextualInfo)
[0024] In the formula, F context represents the context feature map, ContextModule represents the context-aware module, and ContextualInfo represents the context information in the image sample; perform weighted fusion on the context feature map and the target fused feature map to obtain the target enhanced fused feature map.
[0025] Optionally, before the classification module and the regression module are called to analyze the target fusion feature map respectively, the method further includes: determining a target detection box containing a smoking target in each image sample; clustering a plurality of target detection boxes with the optimization objective of minimizing the total distance between the initial anchor boxes and the respective target detection boxes to obtain a plurality of initial anchor boxes; determining the matching degree between each target detection box and its affiliated initial anchor box, and adjusting the position information and size of each initial anchor box through a regression loss function with the optimization objective of maximizing the matching degree between each target detection box and its affiliated initial anchor box to obtain a plurality of adjusted anchor boxes.
[0026] Optionally, clustering a plurality of target detection boxes to obtain a plurality of initial anchor boxes includes: clustering a plurality of target detection boxes according to the following formula to obtain a plurality of initial anchor boxes:
[0027]
[0028] In the formula, D(C,B) represents the total distance between a plurality of initial anchor boxes C and a plurality of target detection boxes B, N represents the total number of target detection boxes, B p represents p target detection boxes, C q =(w q ,h q ) represents the q-th initial anchor box, and w q represents the width of the q-th initial anchor box, h q represents the height of the q-th initial anchor box.
[0029] According to another aspect of the embodiments of the present application, there is also provided a smoking detection device, including: an acquisition module for acquiring a real-time scene image in a target area; a detection module for analyzing the real-time scene image by using a pre-trained target detection model to obtain whether the real-time scene image contains a smoking behavior, wherein the target detection model is used to determine whether the real-time scene image contains a smoking behavior according to the probability distribution of smoking targets related to the smoking behavior in a plurality of detection boxes in the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and the smoking actions of smokers, and the target detection model at least includes: a convolution kernel adjustment module based on an adaptive mechanism, a context awareness module, and a feature fusion module based on a feature pyramid network.
[0030] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes: a computer program, wherein when the computer program is executed by a processor, the above-mentioned smoking detection method is implemented.
[0031] In the embodiments of the present application, a pre-trained object detection model is used to analyze a real-time scene image to determine whether the real-time scene image contains a smoking behavior. Among them, the object detection model can determine whether the real-time scene image contains a smoking behavior based on the probability distribution of smoking targets related to the smoking behavior in multiple detection frames within the real-time scene image. And since the model integrates a convolutional kernel adjustment module based on an adaptive mechanism, an enhanced context awareness module, and a feature fusion module based on a feature pyramid network. Among them, the convolutional kernel adjustment module based on the adaptive mechanism can dynamically adjust the convolutional kernel size and stride to improve the model's detection ability for small targets; while the context awareness module enhances the feature representation ability and improves the recognition accuracy and robustness of the model for smoking behavior in complex scenes; the feature fusion module based on the feature pyramid network optimizes the combination of multi-scale features by fusing feature maps at different levels, so as to maintain excellent detection quality when dealing with large and small targets. Thus, the technical effect of highly accurate detection of smoking behavior in real-time scene images is achieved, the monitoring accuracy is improved, real-time monitoring is realized, and resource consumption is reduced, effectively solving the problems existing in the prior art solutions in terms of accuracy, real-time performance, and robustness. Furthermore, the technical problem of poor accuracy and real-time performance of the detection results of smoking behavior in complex scenes in related smoking detection technologies is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0033] Figure 1 is a schematic flowchart of an optional smoking detection method according to an embodiment of the present application;
[0034] Figure 2 is a schematic structural diagram of an optional smoking detection device according to an embodiment of the present application;
[0035] Figure 3 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0038] To better understand the embodiments of the present application, the following is a translation and explanation of some nouns or terms that appear in the description process of the embodiments of the present application:
[0039] Convolution: In deep learning, convolution is a mathematical operation, usually used to extract features of input data (such as images). By sliding a convolution kernel (i.e., filter) on the input data and performing a dot product operation at each position, a feature map can be obtained. The convolution operation is the convolution neural network.
[0040] Depthwise Separable Convolution: It decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution. Among them, depthwise convolution performs independent convolution operations on each channel of the input, and pointwise convolution uses a 1*1 convolution kernel to perform weighted combination on the features in the depth direction.
[0041] Anchor Box: It is a predefined rectangular box used in the object detection model. The model predicts the bounding boxes of the target objects in the image by adjusting the positions and sizes of these anchor boxes. In YOLOv5, the anchor boxes are used to locate objects and estimate their sizes.
[0042] Feature Pyramid: The feature pyramid is a structure used to process features of different scales. By extracting features at different levels, it improves the detection ability of the model for objects of different sizes. In YOLOv5, the feature pyramid is used to process the information of large and small objects simultaneously.
[0043] Detection Head: The detection head is the last part of the object detection model, responsible for mapping the extracted features to specific prediction results. It generates the final classification results and bounding box coordinates. In YOLOv5, the detection head is used to predict the class and location of objects.
[0044] Generative Adversarial Networks (GANs): A deep learning model mainly consisting of two parts: a generative model (Generator) and a discriminative model (Discriminator). Among them, the generative model is responsible for generating data, while the discriminative model is responsible for evaluating the authenticity of the data. The core idea of GANs is to enable the generative model to produce increasingly realistic data through the adversarial process of the two models.
[0045] YOLOv5: A single-stage object detection algorithm that adds some new improvement ideas on the basis of YOLOv4, greatly improving its speed and accuracy. Generally, the YOLOv5 object detection algorithm can be divided into four general modules: the input end (for inputting the image to be detected), the backbone network (usually some networks with beneficial performance in classifiers), the Neck network (located in the middle of the backbone network and the head network, which can be used to further improve the diversity and robustness of features), and the Head output end (for completing the output of the object detection results, including a classification branch and a regression branch).
[0046] Embodiment 1
[0047] According to the embodiments of the present application, a smoking detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0048] Figure 1 It is a schematic flowchart of a smoking detection method provided according to the embodiments of the present application. As Figure 1 shown, the method includes the following steps:
[0049] Step S102, obtaining a real-time scene image within the target area.
[0050] In the technical solution provided in the above step S102, the above target area can be various no-smoking areas such as public places, office buildings, hospitals, schools, etc. Therefore, the above real-time scene image can be a static picture captured by an intelligent camera or other image acquisition devices or an image frame within a real-time video stream in the no-smoking area.
[0051] In addition, to improve the accuracy of subsequent feature extraction and object detection, the acquired real-time scene image can also be preprocessed, such as decoding (used to convert image data into a format recognizable by a computer), denoising (used to improve image quality), resizing (used to adjust the image size to meet the model input requirements), color space conversion (ensuring consistent color representation), etc., so that the data format input into the model subsequently is consistent and of good quality.
[0052] Step S104: Analyze the real-time scene image using a pre-trained object detection model to obtain whether the real-time scene image contains a smoking behavior.
[0053] In the technical solution provided in the above step S104, the above object detection model determines whether the real-time scene image contains a smoking behavior based on the probability distribution of smoking targets related to the smoking behavior in multiple detection frames within the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and the smoking actions of smokers. That is to say, the object detection model is trained with a large number of data sets marked with smoking behaviors and can identify and locate the smoking targets related to the smoking behavior in the image to be analyzed, including smoke, cigarette butts, and the smoking actions of smokers. In addition, the core mechanisms of the object monitoring model include: a convolutional kernel adjustment module based on an adaptive mechanism, a context awareness module, and a feature fusion module based on a feature pyramid network, where: the convolutional kernel adjustment module based on an adaptive mechanism can dynamically adjust the convolutional kernel size and stride to improve the model's detection ability for small targets; the context awareness module can combine object features with surrounding environmental information to enhance the feature representation ability; the feature fusion module based on a feature pyramid network improves the detection performance by fusing feature maps at different levels.
[0054] The following describes each step of the smoking detection method in combination with a specific implementation process.
[0055] As an optional implementation method, the training process of the object detection model can include:
[0056] Step S1: Obtain multiple groups of training samples. Among them, each group of training samples includes: an image sample, a sample label corresponding to the image sample, and the position information and size of the real detection frame of the smoking target related to the smoking behavior included in the image sample, and the sample label is used to characterize whether the image sample contains a smoking behavior.
[0057] Specifically, for the sample acquisition scheme provided in the above step S1, it can be implemented through the following steps S11 - S14, including:
[0058] Step S11: Obtain multiple image samples in different regions collected by an image acquisition device, and determine the sample labels for each image sample. Among them, these image samples are divided into those containing smoking behavior and those not containing smoking behavior, and these regions can include public places, indoor office areas, schools, etc. under different lighting conditions and at different angles. That is, by using a surveillance camera or a mobile device to record videos or take static images containing and not containing smoking behavior in the above scenarios, ensure that the quality of the collected videos or images is high and the lighting conditions vary, so as to cover different environmental conditions and enable the model to generalize to various actual environments.
[0059] Step S12: Determine the sample labels for each image sample, and for the image samples containing smoking behavior, mark the true detection boxes of the smoking targets related to the smoking behavior. This step can usually be performed by professional annotators or based on the results of a preliminary detection algorithm. That is, select the positions and sizes of the targets (including cigarettes, smoke, smoking actions, etc.) on the image and mark their categories (e.g., "smoking", "not smoking").
[0060] Step S13: Perform preprocessing operations on each image sample. Among them, the preprocessing operations include but are not limited to: image scaling, cropping, color adjustment, etc. Thus, it adapts to the input requirements of the model and increases the diversity of the data to help the model better learn the characteristics of smoking behavior.
[0061] Step S14: Integrate the collected image samples, the corresponding sample labels, and the positions and sizes of the true detection boxes within the image samples into multiple groups of training samples.
[0062] Through the above steps S11 - S14, a high-quality training sample set containing rich scenarios and smoking behavior types can be constructed.
[0063] In addition, during the above process of obtaining training samples, additional training samples can be generated through generative adversarial networks (GANs). Among them, GANs consist of two parts: a generator and a discriminator. The generator is responsible for generating synthetic images that are indistinguishable from real images, while the discriminator tries to distinguish between the generated images and real images. In the context of smoking detection, GANs can be used to generate images containing specific smoking behaviors or complex backgrounds to expand the training data set and improve the generalization ability of the model. At the same time, the images generated by GANs can simulate various lighting conditions, occlusion situations, and different smoking postures, thus ensuring that the model can accurately detect smoking behavior in multiple scenarios.
[0064] Step S2: Construct a deep convolutional neural network model based on the YOLOv5 architecture.
[0065] Specifically, the deep convolutional neural network model specifically includes the following modules: a static convolution module (which includes multiple convolution kernels with fixed sizes and strides), a context awareness module (which generates a context feature map by analyzing the target features and surrounding environment information), a convolution kernel adjustment module based on an adaptive mechanism (which adaptively adjusts the parameters of the convolution kernels by analyzing the resolution of the input image and the target size), a dynamic convolution module (which includes multiple convolution kernels with dynamic sizes and strides), a feature fusion module based on a feature pyramid network (which fuses feature maps at different levels through a preset weighted fusion strategy), a classification module (for classifying the detected targets), and a regression module (for adjusting the position and size of the anchor boxes).
[0066] Step S3: Use multiple groups of training samples to iteratively train the deep convolutional neural network model to obtain a target detection model.
[0067] Optionally, in the iterative training scheme provided for the above step S3, it can be implemented through the following steps:
[0068] Step S31: Traverse each image sample and loop through the following steps S311 - S31:
[0069] Step S311: Invoke the static convolution module to perform a convolution operation on the image sample to obtain a corresponding initial feature map.
[0070] Specifically, for a model based on the YOLOv5 architecture, its feature map is generated by the backbone network of the model. Therefore, the above static convolution module can be understood as the (deep) feature maps output by multiple convolution layers and pooling layers within the backbone network, and this (deep) feature map needs to contain potential information such as the size and position of the target. Among them, the sizes and strides of the convolution kernels in these convolution layers are all fixed parameters set at the initial stage of the network. Additionally, the above convolution operation can be a depthwise separable convolution, which greatly reduces the number of model parameters and computational complexity while maintaining the feature extraction ability.
[0071] Step S312: Invoke the convolution kernel adjustment module to determine the sizes and strides of the convolution kernels within the dynamic convolution module based on the size of the image sample and the size of the initial feature map.
[0072] Optionally, for the specific scheme provided in the above step S312, it can be implemented in the following way:
[0073] Invoke the convolution kernel adjustment module to determine the sizes and strides of the convolution kernels within the dynamic convolution module based on the size of the image sample and the size of the initial feature map according to the following formula:
[0074]
[0075] In the formula, K d represents the size of the adjusted convolution kernel, and S d represents the adjusted stride, and K b represents the size of the convolution kernel before adjustment (i.e., the standard size selected during network design), and S b represents the stride before adjustment, and S img represents the size of the image sample (which can be the width or height of the image sample), and S feat represents the size of the initial feature map (i.e., the size of the feature map after convolution and pooling operations).
[0076] The core idea of the above adjustment strategy is: when the resolution of the input image is relatively high or the target size is relatively small, the size of the convolution kernel and the stride can be appropriately reduced to capture more detailed information; while when the resolution of the input image is relatively low or the target size is relatively large, the size of the convolution kernel and the stride can be appropriately increased to speed up the processing speed and reduce the computational amount. This dynamic adjustment mechanism helps to improve the detection accuracy and efficiency of the model for small targets (such as smoke and cigarettes), especially in complex backgrounds and dynamic environments.
[0077] Step S313, call the dynamic convolution module, and perform convolution operations on the initial feature map according to the size and stride of the adjusted convolution kernel to obtain multiple feature maps with different depths, and form a feature pyramid network from the multiple feature maps. Among them, the depth corresponding to each feature map in the feature pyramid network decreases layer by layer from top to bottom.
[0078] That is to say, each convolution layer in the dynamic convolution module continues to perform convolution operations on the initial feature map according to the size and stride of the adjusted convolution kernel, so that each convolution layer will generate new and optimized feature maps. Feature maps with different depths are used to construct a feature pyramid (Feature Pyramid Network, FPN) from low to high according to their levels in the backbone network. Similarly, the above convolution operation can also be a depthwise separable convolution, which greatly reduces the number of model parameters and computational complexity while maintaining the feature extraction ability.
[0079] Step S314, call the feature fusion module based on the feature pyramid network to fuse each feature map in the feature pyramid network to obtain the target fusion feature map.
[0080] Since the feature pyramid contains feature maps at multiple levels, and each level corresponds to a different scale. Among them, the feature maps at higher levels (i.e., deeper levels) usually contain more semantic information but may lack details; the feature maps at lower levels (i.e., shallower levels) provide more detailed information but less semantic information. Therefore, the embodiments of the present application propose to apply a dynamic weighting strategy to feature maps of different scales. This strategy dynamically adjusts the weights according to the importance of the feature maps in order to more effectively integrate the feature information of each layer.
[0081] Optionally, the specific solution provided in step S314 above can be implemented in the following manner:
[0082] The first step: For each feature layer in the feature pyramid network, loop through the following steps until all feature layers in the feature pyramid are traversed, including:
[0083] First, call the feature fusion module based on the feature pyramid network to upsample the feature maps in the feature layer to match the resolution of the shallower feature maps with lower depth, and perform weighted fusion on the feature maps and adjacent shallower feature maps to obtain the corresponding fused feature maps.
[0084] Specifically, each layer of the deep convolutional neural network generates a feature map of a specific scale. The sizes of these feature maps gradually decrease as the network depth increases, while the semantic information of the features gradually increases. Therefore, in order to fuse the feature maps of different depths in the feature pyramid, starting from the deepest feature map in the feature pyramid (i.e., the last layer at the top of the feature pyramid), perform upsampling operations (such as nearest neighbor interpolation, bilinear interpolation, etc.) step by step from top to bottom to gradually restore the resolution of this feature map to be the same as the size of the shallower feature map (i.e., the feature map of the second-to-last layer of the feature pyramid) until all levels of feature maps are integrated at the top of the pyramid to generate the final feature representation for object detection. Among them, after upsampling, the feature pyramid can perform horizontal connection between the shallower feature maps and the upsampled deeper feature maps, so that the feature maps of each level contain a combination of features of different depths, thereby obtaining more comprehensive object information at each scale.
[0085] Generally, horizontal connection can be performed by element-wise addition or element-wise multiplication. For example, the following formula can be used to perform weighted fusion on the feature map and the adjacent shallower feature map to obtain the corresponding fused feature map:
[0086]
[0087] In the formula, F i represents the fused feature map of the i-th layer in the feature pyramid network, F s represents the shallower feature map below the feature map of the i-th layer in the feature pyramid network, Fd denotes the feature map of the i-th layer, and α and β denote the weighting coefficients. The distribution is used to control the contribution values of the shallow feature map and the deep feature map.
[0088] Among them, in the embodiments of the present application, according to different scenario requirements or statistical information of the feature map, the ratio of the two is automatically adjusted through an adaptive adjustment strategy to ensure the optimal effect of feature fusion. The expression of the above adaptive adjustment strategy can be written as:
[0089]
[0090] β = 1 - α
[0091] where γ represents a regulation parameter used to control the balance of feature fusion, and S fs represents the statistical information of the shallow feature map (such as the mean and variance of each pixel point in the shallow feature map).
[0092] Through the above fusion strategy, the fused feature map F i has the high-resolution spatial information of the shallow feature and also has the strong semantic expression ability of the deep feature, enhancing the detection ability of the model for multi-scale targets, especially the precise positioning of small targets, thereby significantly improving the robustness and accuracy of the detection model.
[0093] Next, all the feature values in the fused feature map within the feature layer are determined, the quality score of the fused feature map is determined based on all the feature values, and the weight of the feature layer is determined based on the quality score.
[0094] Optionally, the quality score of the fused feature map within each feature layer is determined by the clarity and information amount of the feature map. Therefore, based on all the feature values, the quality score of each layer of the fused feature map is determined using the following formula:
[0095]
[0096] In the formula, Q i represents the quality score of the i-th layer of the fused feature map within the feature pyramid network, F i,j represents the j-th feature value in the i-th layer of the fused feature map within the feature pyramid network, N represents the total number of feature values in the fused feature map, and Var(·) represents the variance of the pixel values in the i-th layer of the fused feature map within the feature pyramid network (used to measure the information amount of the feature map).
[0097] Finally, based on the quality score of each layer of the fused feature map, the weight of each feature layer is calculated according to the following formula:
[0098]
[0099] In the formula, W iRepresents the weight of the fused feature map at the i-th layer within the feature pyramid network, Q i Represents the quality score of the fused feature map at the i-th layer within the feature pyramid network. M represents the total number of layers of the feature pyramid network, Represents the sum of the quality scores of the fused feature maps of all layers within the feature pyramid.
[0100] Step 2: According to the quality scores of the fused feature maps of each layer and the weights of each feature layer, calculate the target fused feature map according to the following formula:
[0101]
[0102] In the formula, F feature Represents the target fused feature map, W i Represents the weight of the fused feature map at the i-th layer within the feature pyramid, F i Represents the fused feature map at the i-th layer within the feature pyramid, and M represents the total number of layers of the feature pyramid network.
[0103] Through the above feature pyramid weighting strategy, by dynamically adjusting the weights of the fused feature maps of each layer, the feature information of each layer is reasonably reflected in the final fusion process, thereby improving the detection accuracy of the model for targets of different scales.
[0104] Step S315: Invoke the context awareness module to extract the context feature map of the image sample, and fuse the context feature map with the initial feature map to obtain an enhanced feature map. Among them, the context feature map includes the background information and environmental feature information of the image sample.
[0105] Specifically, the embodiment of the present application extracts the context information of the target and its surrounding environment through the context awareness module to improve the accuracy and robustness of target detection. Among them, this process is implemented in the detection head stage, enhancing the feature representation ability and enabling the model to better handle the target detection task under complex backgrounds.
[0106] Optionally, for the technical solution provided in the above step S315, it can be implemented in the following manner:
[0107] First step: Invoke the context awareness module to extract the feature map corresponding to the context information within the image sample according to the following formula to obtain the context feature map:
[0108] F context = ContextModule(F feature , ContextualInfo)
[0109] In the formula, F contextThe ContextFeatureMap represents the context feature map, the ContextModule represents the context awareness module, and the ContextualInfo represents the context information within the image sample (including but not limited to the relative position of the human hand and the cigarette, the contrast between the cigarette smoke and the surrounding environment, the scene features of the ashtray or lighter, etc.).
[0110] It should be noted that a self-attention mechanism can be added during the above context feature extraction process. This is because the self-attention mechanism allows the model to focus on different positions in the feature map. By calculating the attention weights for each position in the feature map, the model can focus more on the most relevant regions, thereby improving the recognition accuracy of the target. Therefore, in smoking detection, the self-attention mechanism can help the model better understand the context of the smoking behavior, such as the hand movements of the person, the position of the cigarette, and the shape of the smoke, thus enhancing the robustness and accuracy of the detection.
[0111] Step 2: Perform weighted fusion on the context feature map and the target fusion feature map to obtain the target enhanced fusion feature map. Among them, the fusion process of the above context feature map and the target fusion feature map can be specifically represented by the following formula:
[0112]
[0113] In the formula, F fused represents the target enhanced fusion feature map, both represent the weighting coefficients, respectively adjusting the contribution ratios of the fusion feature map F feature and the context feature map F context .
[0114] Step S316, call the classification module and the regression module to analyze the target fusion feature map respectively, and obtain the predicted label corresponding to the image sample and the position information and size of the predicted anchor box of the image sample.
[0115] Specifically, the detection head in the target detection model includes two tasks: classification and regression. Therefore, the classification module is responsible for predicting the category of the target in the image sample, while the regression module is used to predict the position information (center point coordinates) and size (width and height) of the boundary detection box of the target.
[0116] When using the classification module for analysis, the classification module applies a fully-connected layer or a convolutional layer at each predefined anchor position to predict the presence of a target at that position and the category of the target. The classification module can output a vector whose length is equal to the number of categories plus 1 (the additional 1 represents the background or no target). Each element in the vector corresponds to a probability distribution for a category. For example, for smoking behavior detection, if two categories (smoking, non-smoking) are set, the classification module will output a vector containing three elements, where the first two elements represent the probabilities of smoking and non-smoking respectively, and the last element represents the probability of the background.
[0117] When using the regression module for analysis, the regression module usually outputs four values, representing the four coordinate points of the bounding box of the relevant smoking target (the x and y coordinates of the center point, and the width and height of the bounding box) or the relative offset and scale information of the bounding box, which are used to adjust the anchor box to more precisely match the target boundary.
[0118] Therefore, by combining the outputs of the classification module and the regression module, prediction labels and prediction anchor boxes are generated for each anchor position. Specifically, if the highest probability category output by the classification module at a certain position is the target category (such as "smoking"), it is considered that there is a target related to the smoking behavior at that position, and the position and size information output by the regression module are used to adjust the anchor box to generate the final predicted bounding box.
[0119] Step S32: Construct a classification loss function based on the sample label and the prediction label corresponding to each image sample, and construct a regression loss function based on the position information and size of the true detection box and the position information and size of the corresponding predicted anchor box within each training sample, and construct an objective loss function based on the classification loss function and the regression loss function.
[0120] Specifically, the classification loss function (Classification Loss) is used to measure the difference between the prediction label output by the model and the true label of the sample. Commonly used classification loss functions include the cross-entropy loss (Cross-Entropy Loss) function, etc. The regression loss function is used to measure the position and size differences between the anchor box predicted by the model and the true detection box. Commonly used regression loss functions include the Smooth L1 Loss function, the Intersection over Union Loss (IoU Loss) function, etc., and the expression of the IoU loss function can be written as:
[0121]
[0122] where, A gt represents the area of the true detection box within the image sample, A predRepresents the area of the predicted anchor box within the image sample.
[0123] Therefore, the expression of the objective loss function can be written as:
[0124] L = L cls + L reg
[0125] where L cls represents the classification loss function, and L reg represents the regression loss function.
[0126] Step S33: By calculating the gradient of the objective loss function with respect to the model parameters and updating the model parameters based on the gradient until the model converges, a trained object detection model is obtained.
[0127] Specifically, the gradient of the objective loss function with respect to the model parameters (such as convolution kernels, weight matrices, etc.) is calculated through backpropagation, and a preset optimization algorithm (such as Stochastic Gradient Descent SGD, Adam, RMSprop, etc.) is used to update the model parameters based on the calculated gradient. The above forward propagation, loss calculation, backpropagation, and parameter update processes are repeated until the model converges, and a trained object detection model is obtained. In addition, during the training process, hyperparameters (such as learning rate, weight decay, etc.) may also need to be adjusted to find the combination of hyperparameters that makes the model performance optimal. This is usually done through methods such as grid search, random search, or Bayesian optimization.
[0128] It should be noted that in the object detection task, the anchor box is used for predicting the position of the candidate object. However, the models of the YOLOv5 architecture usually use fixed anchor boxes for detection, and this design lacks flexible adaptation to the changes in the object scale and shape, resulting in a decrease in the detection localization accuracy when the object scale or shape is different.
[0129] Therefore, the embodiment of this application is through an adaptive anchor box optimization method, aiming to dynamically adjust the size and shape of the anchor box according to the size distribution of the object, so as to improve the detection accuracy and robustness.
[0130] Optionally, the embodiment of this application can adaptively adjust the position information and size of the anchor box in the following way, including:
[0131] The first step: Determine the object detection box containing the smoking object in each image sample.
[0132] The second step: Taking the minimization of the total distance between the initial anchor box and each object detection box as the optimization goal, clustering multiple object detection boxes to obtain multiple initial anchor boxes. Among them, these initial anchor boxes can cover smoking objects of different sizes in the training sample set.
[0133] Specifically, the embodiments of the present application can cluster multiple target detection boxes according to the following formula to obtain multiple initial anchor boxes:
[0134]
[0135] In the formula, D(C,B) represents the total distance between multiple initial anchor boxes C and multiple target detection boxes B, N represents the total number of target detection boxes, B p represents p target detection boxes, and C q =(w q ,h q ) represents the q-th initial anchor box. The clustering result is a set of anchor boxes C = {(w q ,h q ), q = 1, 2,..., k}, where w q represents the width of the q-th initial anchor box, h q represents the height of the q-th initial anchor box, and k represents the number of clustering clusters.
[0136] Step 3: Determine the matching degree between each target detection box and its corresponding initial anchor box, and take maximizing the matching degree between each target detection box and its corresponding initial anchor box as the optimization goal. Adjust the position information and size of each initial anchor box through a regression loss function to obtain multiple adjusted anchor boxes.
[0137] Specifically, in the target detection process of the embodiments of the present application, in order to maximize the IoU value, the regression loss function can be continuously optimized during the iteration process to adjust the position information and size of the anchor box.
[0138] For example, for each anchor box, the size (including height and width) of the anchor box is dynamically adjusted through feedback learning during the training process. The specific formula is as follows:
[0139] w new =w old ·(1 + η·δ w )
[0140] h new =h old ·(1 + η·δ h )
[0141] In the formula, w new , h new respectively represent the adjusted width and height of the target anchor box, w old , h new respectively represent the width and height of the target anchor box before adjustment, η represents the learning rate, and δ w , δh The updated gradients representing the width and height respectively are obtained by calculating based on the loss function of the model.
[0142] Through the above adaptive anchor box adjustment mechanism, the model can adaptively adjust the anchor box according to the size and shape of different targets, enabling it to better adapt to the scale changes of different scenarios and targets, and greatly improving the robustness and detection accuracy of the model.
[0143] Therefore, through the training process of the object detection model provided by the above steps S1 - S3, an object detection model that can perform well on a given task can be obtained. Therefore, the trained object detection model is applied to real-time or non-real-time image or video streams to achieve the object detection task. Among them, in practical applications, the object detection model may run on embedded devices, servers or cloud platforms to meet different computing and latency requirements.
[0144] In addition, the embodiment of the present application can also trigger corresponding warning prompt information according to the detection result of whether the real-time scene image of the detected target area contains a smoking behavior, so as to prompt the management personnel to manage the smoking behavior in the no-smoking area in time.
[0145] Therefore, through the above description of the smoking detection method provided by the embodiment of the present application, it is not difficult to see that compared with the existing object detection solutions, the solution of the present application has the following technical advantages:
[0146] (1) By introducing a convolutional kernel adjustment module based on an adaptive mechanism, the convolutional kernel size and stride can be adjusted in real time according to the resolution of the input image and the size of the smoking target, thereby improving the detection ability of the model for small targets (such as smoke and cigarettes). Especially in real-time video streams or high-resolution images, the flexibility and accuracy of feature extraction are ensured, thus significantly improving the detection accuracy.
[0147] (2) By introducing a feature fusion module based on a feature pyramid network, the combination of deep and shallow features is optimized, and the fusion of shallow and deep features can utilize both detailed and semantic information at the same time, significantly improving the accuracy of object detection. Especially in the case of small targets and complex backgrounds, this fusion enhances the robustness of the model and improves the detection effect.
[0148] (3) By introducing a weighting strategy in the feature pyramid, the weights of each layer of features are dynamically adjusted according to the quality and importance of the feature layer, thereby improving the sensitivity of the model to different feature layers and optimizing the fusion process of features at different scales.
[0149] (4) By adaptively optimizing the anchor boxes, the anchor boxes can flexibly adapt to the scales and shapes of different targets, improving the detection accuracy and effect of the model in various scenarios, especially in the case of diverse target sizes and shapes.
[0150] (5) This solution generates initial anchor boxes through a clustering method and dynamically adjusts the size and proportion of the anchor boxes according to the actual detection error. This adaptive adjustment mechanism enables the model to more flexibly adapt to the scales and shapes of different targets, reducing the false detection rate and missed detection rate, and improving the generality and accuracy of target detection.
[0151] (6) By integrating a context awareness module, the model can fuse target features and surrounding environment information, improving the detection ability in complex scenarios and occlusion situations. This module enhances the understanding of the target context, making the detection more accurate and robust, and maintaining a high detection accuracy even in the case of complex backgrounds or partially occluded targets.
[0152] Embodiment 2
[0153] According to the embodiments of the present application, there is also provided a smoking detection device for implementing the smoking detection method in Embodiment 1, as Figure 2 shown. The smoking detection device at least includes: an acquisition module 22 and a detection module 24, where:
[0154] The acquisition module 22 is used to acquire real-time scene images within a target area;
[0155] The detection module 24 is used to analyze the real-time scene images by using a pre-trained target detection model to obtain whether the real-time scene images contain smoking behavior. The target detection model is used to determine whether the real-time scene images contain smoking behavior based on the probability distribution of smoking targets related to smoking behavior within multiple detection boxes in the real-time scene images, and the smoking targets at least include: smoke, cigarette butts, and the smoking actions of smokers. The target detection model at least includes: a convolutional kernel adjustment module based on an adaptive mechanism, a context awareness module, and a feature fusion module based on a feature pyramid network.
[0156] It should be noted that each module in the smoking detection device in the embodiments of the present application corresponds to each implementation step of the smoking detection method in Embodiment 1. Since the details not shown in this embodiment can be referred to in Embodiment 1, they will not be elaborated here.
[0157] Embodiment 3
[0158] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, the smoking detection method in Embodiment 1 is implemented.
[0159] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. When the device where the non-volatile storage medium is located runs this computer program, the smoking detection method in Embodiment 1 is executed.
[0160] According to an embodiment of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, the smoking detection method in Embodiment 1 is executed.
[0161] According to an embodiment of the present application, there is also provided an electronic device, which includes: a memory and a processor. Among them, a computer program is stored in the memory, and the processor is configured to execute the smoking detection method in Embodiment 1 through the computer program.
[0162] Specifically, when the computer program runs, the following steps are implemented: obtaining a real-time scene image in a target area; using a pre-trained target detection model to analyze the real-time scene image to obtain whether the real-time scene image contains a smoking behavior. Among them, the target detection model is used to determine whether the real-time scene image contains a smoking behavior according to the probability distribution of smoking targets related to the smoking behavior in multiple detection frames within the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and the smoking actions of smokers. The target detection model at least includes: a convolution kernel adjustment module based on an adaptive mechanism, a context awareness module, and a feature fusion module based on a feature pyramid network.
[0163] As an alternative implementation manner, the above-mentioned electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 3 A hardware structure block diagram of an electronic device for implementing the smoking detection method is shown. As Figure 3 shown, the electronic device 30 may include one or more (shown as 302a, 302b,..., 302n in the figure) processors 302 (the processor 302 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand, Figure 3The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 30 may further include more or fewer components than those shown in Figure 3 or have a different configuration from that shown in Figure 3 .
[0164] It should be noted that the above one or more processors 302 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the electronic device 30. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0165] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the smoking detection method in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the vulnerability detection method of the above application program. The memory 304 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 304 may further include a memory remotely set relative to the processor 302, and these remote memories can be connected to the electronic device 30 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0166] The transmission device 306 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the electronic device 30. In one instance, the transmission device 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 306 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0167] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the electronic device 30.
[0168] The above serial numbers of the embodiments are only for description and do not represent the advantages or disadvantages of the embodiments.
[0169] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0170] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0171] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0172] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0174] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A smoking detection method, characterized in that: include: Acquire real-time scene images within the target area; The real-time scene image is analyzed by using a pre-trained target detection model to determine whether the real-time scene image contains smoking behavior, wherein the target detection model is used to determine whether the real-time scene image contains smoking behavior based on the probability distribution of smoking targets related to smoking behavior contained in multiple detection frames in the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and smoking actions of smokers. The target detection model at least includes: a convolution kernel adjustment module based on an adaptive mechanism, a context perception module, and a feature fusion module based on a feature pyramid network.
2. The method according to claim 1, characterized in that The training process of the target detection model includes: Acquire multiple groups of training samples, wherein each group of training samples includes: an image sample, a sample label corresponding to the image sample, and position information and size of a real detection frame containing a smoking target related to smoking behavior in the image sample, and the sample label is used to indicate whether the image sample contains smoking behavior; Build a deep convolutional neural network model based on the YOLOv5 architecture; The deep convolutional neural network model is iteratively trained using the multiple groups of training samples to obtain the target detection model.
3. The method according to claim 2, characterized in that The deep convolutional neural network model also includes: a static convolution module including multiple convolution kernels of fixed sizes and steps, a dynamic convolution module including multiple convolution kernels of dynamic sizes and steps, a classification module, and a regression module, wherein the deep convolutional neural network model is iteratively trained using the multiple groups of training samples to obtain the target detection model, including: For each of the image samples, the static convolution module is called to perform a convolution operation on the image sample to obtain a corresponding initial feature map; the convolution kernel adjustment module is called to determine the size and step size of each convolution kernel in the dynamic convolution module according to the size of the image sample and the size of the initial feature map; the dynamic convolution module is called to perform a convolution operation on the initial feature map according to the adjusted size and step size of the convolution kernel to obtain a plurality of feature maps of different depths, and a feature pyramid network is formed by the plurality of feature maps, wherein the depth corresponding to each layer of the feature map in the feature pyramid network decreases layer by layer from top to bottom; the feature fusion module based on the feature pyramid network is called to fuse the feature maps in the feature pyramid network to obtain a target fused feature map; the context perception module is called to extract the context feature map of the image sample, and the context feature map is fused with the target fused feature map to obtain a target enhanced fused feature map; the classification module and the regression module are called to analyze the target enhanced fused feature map respectively to obtain the predicted label corresponding to the image sample and the position information and size of the predicted anchor box of the image sample; Constructing a classification loss function according to the sample label corresponding to each of the image samples and the predicted label, constructing a regression loss function according to the position information and size of the real detection box in each of the training samples and the position information and size of the corresponding predicted anchor box, and constructing a target loss function based on the classification loss function and the regression loss function; The trained target detection model is obtained by calculating the gradient of the target loss function with respect to the model parameters and updating the model parameters based on the gradient until the model converges.
4. The method according to claim 3, characterized in that Calling the convolution kernel adjustment module to determine the size and step size of each convolution kernel in the dynamic convolution module according to the size of the image sample and the size of the initial feature map, including: The convolution kernel adjustment module is called to determine the size and step size of each convolution kernel in the dynamic convolution module according to the size of the image sample and the size of the initial feature map according to the following formula: Where K d Represents the size of the adjusted convolution kernel, S d represents the adjusted step size, K b Represents the size of the convolution kernel before adjustment, S b represents the step length before adjustment, S img represents the size of the image sample, S feat represents the size of the initial feature map.
5. The method according to claim 3, characterized in that: Calling the feature fusion module based on the feature pyramid network to fuse the feature maps in the feature pyramid network to obtain a target fused feature map, including: For each feature layer in the feature pyramid network, the feature fusion module based on the feature pyramid network is called to upsample the feature map in the feature layer so that its resolution matches that of a shallow feature map with a lower depth, and the feature map and the adjacent shallow feature map are weightedly fused to obtain a corresponding fused feature map; all feature values in the fused feature map in the feature layer are determined, a quality score of the fused feature map is determined based on all the feature values, and a weight of the feature layer is determined based on the quality score; According to the quality score of the fused feature map in each feature layer and the weight of each feature layer, the target fused feature map is calculated according to the following formula: Where F feature represents the target fusion feature map, M represents the total number of layers of the feature pyramid network, and W i represents the weight of the fusion feature map at the i-th layer in the feature pyramid, F i Represents the fused feature map of the i-th layer in the feature pyramid.
6. The method according to claim 5, characterized in that The feature map and the adjacent shallow feature map are weightedly fused to obtain a corresponding fused feature map, including: The feature map and the adjacent shallow feature map are weighted fused according to the following formula to obtain the corresponding fused feature map: Where F i represents the fusion feature map of the i-th layer in the feature pyramid network, F s represents the shallow feature map below the i-th layer feature map in the feature pyramid network, F d represents the i-th layer feature map, α, β represent weighting coefficients, and β=1-α, where γ represents the adjustment parameter, S fs Represents the statistical information of the shallow feature map.
7. The method according to claim 5, characterized in that Determining a quality score of the fused feature map according to all the feature values, and determining a weight of the feature layer according to the quality score, including: Based on all the feature values, the quality score of the fused feature map is determined using the following formula: Where Q i represents the quality score of the fusion feature map at the i-th layer in the feature pyramid network, F i,j represents the jth eigenvalue in the i-th layer fused feature map in the feature pyramid network, N represents the total number of eigenvalues in the fused feature map, Var(·) represents the variance of the pixel value in the i-th layer fused feature map in the feature pyramid network; according to the quality score of the fused feature map, the weight of the feature layer is calculated according to the following formula: Where W i represents the weight of the fusion feature map of the i-th layer in the feature pyramid network, Q i Represents the quality score of the i-th layer fused feature map in the feature pyramid network.
8. The method according to claim 3, characterized in that Calling the context perception module to extract the context feature map of the image sample, and fusing the context feature map with the target fusion feature map to obtain a target enhanced fusion feature map, including: The context perception module is called to extract the feature map corresponding to the context information in the image sample according to the following formula to obtain the context feature map: F context =ContextModule(F feature ,ContextualInfo) Where F context represents the context feature map, ContextModule represents the context perception module, ContextualInfo represents contextual information within the image sample; The context feature map is weightedly fused with the target fusion feature map to obtain the target enhanced fusion feature map.
9. The method according to claim 3, characterized in that: Before calling the classification module and the regression module to analyze the target fusion feature map respectively, the method further includes: Determine an object detection frame containing the smoking object in each of the image samples; Taking minimizing the total distance between the initial anchor point frame and each of the target detection frames as the optimization goal, multiple target detection frames are clustered to obtain multiple initial anchor point frames: Determine the degree of match between each of the target detection frames and the initial anchor frame to which it belongs, and take maximizing the degree of match between each of the target detection frames and the initial anchor frame to which it belongs as the optimization goal, adjust the position information and size of each of the initial anchor frames through the regression loss function, and obtain multiple adjusted anchor frames.
10. The method according to claim 9, characterized in that Clustering the multiple target detection frames to obtain multiple initial anchor frames includes: Cluster the multiple target detection frames according to the following formula to obtain multiple initial anchor frames: Where D(C, B) represents the total distance between the multiple initial anchor boxes C and the multiple target detection boxes B, N represents the total number of target detection boxes, and B p represents p target detection boxes, C q =(w q ,h q ) represents the qth initial anchor box, and w q represents the width of the qth initial anchor box, h q Represents the height of the qth initial anchor box.
11. A smoking detection device, characterized in that: include: An acquisition module is used to acquire real-time scene images in a target area; A detection module is used to analyze the real-time scene image using a pre-trained target detection model to determine whether the real-time scene image contains smoking behavior, wherein the target detection model is used to determine whether the real-time scene image contains smoking behavior based on the probability distribution of smoking targets related to smoking behavior contained in multiple detection frames in the real-time scene image, and the smoking targets at least include: smoke, cigarette butts, and smoking actions of smokers. The target detection model at least includes: a convolution kernel adjustment module based on an adaptive mechanism, a context perception module, and a feature fusion module based on a feature pyramid network.
12. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the smoking detection method according to any one of claims 1 to 10 is implemented.