Smoking Event Object Location Detection Method, Apparatus, Device, and Storage Medium
By converting the real labeled data into Gaussian heat map in the initial position detection model and adjusting the weight loss function, the problem of inaccurate cigarette position detection is solved, the detection accuracy is improved, and it is suitable for real-time detection of edge devices.
Patent Information
- Application Number
- CN202211339075.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-28
AI Technical Summary
In the prior art, the detection accuracy of cigarette locations is low, especially when the vehicle is driving, the identification of the elongated small object of cigarette is inaccurate.
By obtaining sample image data and labeling the real data, it is converted into a Gaussian thermal map, and the initial position detection model is used to train and adjust the weight loss function, combining the cigarette butt orientation and shape characteristics of the cigarette, the accuracy of the detection model is improved.
Improves the accuracy of detection of cigarette locations, reduces false detection, and is suitable for real-time detection needs of edge devices.
Smart Images

Figure CN115619854B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technologies, and in particular, to a method, apparatus, device, and storage medium for detecting the position of a smoking event object. Background Art
[0002] It is generally prohibited for a driver to smoke during vehicle driving. In some other specific scenarios, smoking events of personnel also need to be prohibited. Therefore, a solution for detecting smoking events is required.
[0003] The detection of smoking events often relies on the detection of the position of a smoking event object, such as the cigarette itself. Since the target size of a cigarette is relatively small compared to a person's hand, head, etc., and its shape is long and thin, the recognition of a cigarette using the methods of related technologies is inaccurate, and the accuracy of detecting the position of a cigarette is relatively low. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, apparatus, device, and storage medium for detecting the position of a smoking event object, so as to solve the technical problem of low accuracy in detecting the position of a cigarette in related technologies.
[0005] In view of the above problems, the present invention provides a method for detecting the position of a smoking event object, the method comprising: obtaining sample image data and the true annotation data of the sample image data, the true annotation data being obtained by pre-annotating the smoking event object in the sample image, the smoking event object including at least one of a head, a mouth, a hand, and a cigarette, the true annotation data including true annotation box center position data and true annotation box size data; if the smoking event object includes any one of a head, a mouth, and a hand, converting the true annotation box center position data into a true circular Gaussian heat map, if the smoking event object includes a cigarette, converting the true annotation box center position data into a true elliptical Gaussian heat map, and taking the true circular Gaussian heat map and the true elliptical Gaussian heat map as true Gaussian heat maps; inputting the sample image data into an initial position detection model to obtain the predicted annotation data of the smoking event object, the predicted annotation data including predicted annotation box size data and a predicted Gaussian heat map, if the smoking event object includes any one of a head, a mouth, and a hand, the predicted Gaussian heat map including a predicted circular Gaussian heat map, if the smoking event object includes a cigarette, the predicted Gaussian heat map including a predicted elliptical Gaussian heat map, the predicted Gaussian heat map being used to represent the predicted annotation box center position data; training the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data; obtaining a to-be-detected image, inputting the to-be-detected image into the trained initial position detection model to obtain the detected annotation box center position data and the detected annotation box size data of the smoking event object in the to-be-detected image, so as to detect the position of the smoking event object in the to-be-detected image. In an embodiment of the present invention, before training the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data, the method for detecting the position of the smoking event object further comprises: generating a dimensional weight loss function according to preset head weights, cigarette weights, hand weights, and mouth weights, the cigarette weight being greater than the head weight, the hand weight, and the mouth weight, the head weight being less than the hand weight and the mouth weight, so as to train the initial position detection model by means of the dimensional weight loss function.
[0006] In an embodiment of the present invention, generating a dimensional weight loss function according to preset head weights, cigarette weights, hand weights, and mouth weights comprises:
[0007]
[0008] Among them, Loss is the loss function with dimensional weights. The value of cls is 0, 1, 2, or 3, which respectively represent that the object of the smoking event is a cigarette, the head, the hand, or the mouth. Gama is the coordination factor of Loss, and p^ is the predicted probability. If the value of cls is 0, then weight is the preset cigarette weight; if the value of cls is 1, weight is the preset head weight; if the value of cls is 2, weight is the preset hand weight; if the value of cls is 3, weight is the preset mouth weight. Ln represents the natural logarithm function.
[0009] In an embodiment of the present invention, before converting the central position data of the true annotation box into a true elliptical Gaussian heat map, the method for detecting the position of the object of the smoking event further includes: if the object of the smoking event includes a cigarette, obtaining the orientation of the cigarette butt, and determining the cigarette butt angle of the cigarette butt based on the true annotation box size data and the orientation of the cigarette butt, so as to adjust the initial elliptical Gaussian heat map determined based on the central position data of the true annotation box through the cigarette butt angle.
[0010] In an embodiment of the present invention, converting the central position data of the true annotation box into a true elliptical Gaussian heat map includes: determining the semi-major axis and semi-minor axis of the Gaussian ellipse and the Gaussian distribution variance according to a preset scaling ratio and the true annotation box size data, and generating the initial elliptical Gaussian heat map; rotating the initial elliptical Gaussian heat map according to the cigarette butt angle to obtain the true elliptical Gaussian heat map.
[0011] In an embodiment of the present invention, determining the semi-major axis and semi-minor axis of the Gaussian ellipse and the Gaussian distribution variance according to a preset scaling ratio and the true annotation box size data, and generating the initial elliptical Gaussian heat map includes: determining a reference width according to the true annotation box width of the cigarette and the true annotation box height of the cigarette, and the true annotation box width of the cigarette and the true annotation box height of the cigarette are determined according to the true annotation box size data; determining a reference height according to the reference width and a preset cigarette butt length-width ratio; respectively determining the semi-minor axis and semi-major axis of the Gaussian ellipse based on the preset scaling ratio, the reference width, and the reference height, and determining the short axis and long axis of the Gaussian ellipse; respectively determining the short-axis sub-Gaussian distribution variance and the long-axis sub-Gaussian distribution variance based on the short axis and long axis of the Gaussian ellipse as the Gaussian distribution variance; determining the initial elliptical Gaussian heat map according to the Gaussian distribution variance and the two-dimensional point coordinates of the semi-minor axis and semi-major axis of the Gaussian ellipse in the true annotation box of the cigarette.
[0012] In an embodiment of the present invention, inputting the image to be detected into the trained initial position detection model to obtain the center position data of the detection annotation box of the smoking event object in the image to be detected includes: inputting the image to be detected into the trained initial position detection model to obtain the detection annotation data of the smoking event object to be detected in the image to be detected, where the detection annotation data includes a predicted Gaussian heat map; determining suspected position data of the center positions of a plurality of suspected annotation boxes and the probability values of the suspected position data based on the predicted Gaussian heat map; filtering the suspected position data based on the probability values and a preset threshold; and determining the center position data of the detection annotation box of the smoking event object in the image to be detected based on the probability values of the filtered suspected position data.
[0013] In an embodiment of the present invention, after detecting the position of the smoking event object in the image to be detected, the method for detecting the position of the smoking event object includes: determining the position information of the detection target box according to the center position data of the detection annotation box and the size data of the detection annotation box; determining the smoking state based on a preset determination rule and the position information of the detection target box of the smoking event object in the image to be detected; where the preset determination rule includes at least one of the following. If the smoking event object in the image to be detected does not include a cigarette, the smoking state is determined to be not smoking. If the smoking event object in the image to be detected includes a cigarette and at least one smoking-related object, and it is determined that there is an intersection between the cigarette and the smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of at least one smoking-related object, the smoking state is determined to be smoking, and the smoking-related object includes at least one of the head, hand, and mouth. If the smoking event object in the image to be detected includes a cigarette and at least one smoking-related object, and it is determined that there is no intersection between the cigarette and the smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of at least one smoking-related object, the smoking state is determined to be not smoking.
[0014] An embodiment of the present invention further provides a smoking event object position detection device, which includes: a sample acquisition module, configured to acquire sample image data and true annotation data of the sample image data, where the true annotation data is obtained by pre-annotating the smoking event object in the sample image, and the smoking event object includes at least one of a head, a mouth, a hand, and a cigarette, and the true annotation data includes true annotation box center position data and true annotation box size data; a true image conversion module, configured to, if the smoking event object includes any one of a head, a mouth, and a hand, convert the true annotation box center position data into a true circular Gaussian heat map, and if the smoking event object includes a cigarette, convert the true annotation box center position data into a true elliptical Gaussian heat map, and use the true circular Gaussian heat map and the true elliptical Gaussian heat map as true Gaussian heat maps; a model prediction module, configured to input the sample image data into an initial position detection model to obtain predicted annotation data of the smoking event object, where the predicted annotation data includes predicted annotation box size data and a predicted Gaussian heat map, and if the smoking event object includes any one of a head, a mouth, and a hand, the predicted Gaussian heat map includes a predicted circular Gaussian heat map, and if the smoking event object includes a cigarette, the predicted Gaussian heat map includes a predicted elliptical Gaussian heat map, and the predicted Gaussian heat map is used to represent the predicted annotation box center position data; a model training module, configured to train the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data; a position detection module, configured to acquire a to-be-detected image, input the to-be-detected image into the trained initial position detection model, and obtain the detected annotation box center position data and detected annotation box size data of the smoking event object in the to-be-detected image, so as to detect the position of the smoking event object in the to-be-detected image.
[0015] An embodiment of the present invention further provides an electronic device, including a processor, a memory, and a communication bus; the communication bus is used to connect the processor and the memory; the processor is configured to execute a computer program stored in the memory to implement the method according to any one of the above embodiments.
[0016] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to cause a computer to execute the method according to any one of the above embodiments.
[0017] As described above, a smoking event object position detection method, device, equipment, and storage medium provided by the present invention have the following beneficial effects:
[0018] This method obtains sample image data and true annotation data obtained by annotating smoking event objects therein, converts the center position data of the true annotation box into a true Gaussian heat map, inputs the sample image data into an initial position detection model to obtain the predicted annotation box size data and the predicted Gaussian heat map of the smoking event object, trains the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data, obtains an image to be detected, and inputs it into the trained initial position detection model to obtain the center position data and the detection annotation box size data of the smoking event object in the image to be detected, so as to detect the position of the smoking event object in the image to be detected. By converting the center position of the annotation box of the smoking event object with a slender shape such as a cigarette into an elliptical Gaussian heat map for prediction, the trained initial position detection model can be made more suitable for the recognition of small targets such as cigarette butts, and the detection accuracy of the position of the cigarette is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a system architecture diagram shown in an exemplary embodiment of the present application.
[0020] Figure 2 is a flowchart of a method for detecting the position of a smoking event object shown in an exemplary embodiment of the present application.
[0021] Figure 3 is a schematic structural diagram of an initial position detection model shown in an exemplary embodiment of the present application.
[0022] Figure 4 is a block diagram of a device for detecting the position of a smoking event object shown in an exemplary embodiment of the present application.
[0023] Figure 5 is a schematic structural diagram of an electronic device provided in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0025] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0026] Please refer to Figure 1 , Figure 1 which is a system architecture diagram shown in an exemplary embodiment of the present application. As Figure 1 shown, the system includes an edge chip 101 and a central processor (CPU shown in the figure) 102. The edge chip can be disposed in security devices such as cameras. The central processor can be disposed in the same security device as the edge chip, or transmit relevant data to other terminal devices through a network, and the CPU installed in other terminal devices is used for further processing. An initial position detection model is installed in the edge chip 101. By obtaining sample images and real annotation data obtained by annotating smoking event objects therein, the central position data of the real annotation box is converted into a real Gaussian heat map. The sample image data is input into the initial position detection model to obtain the predicted annotation box size data and the predicted Gaussian heat map of the smoking event object. The initial position detection model is trained based on the predicted Gaussian heat map, the real Gaussian heat map, the real annotation box size data, and the predicted annotation box size data. After the initial position detection model converges, the training is stopped. Based on the obtained image to be detected, through the trained initial position detection model, the detected annotation box size data and the predicted Gaussian heat map of the smoking event object in the image to be detected are obtained. Due to the accuracy problem of the edge chip itself and the sensitivity of the chip quantization model to a large number of 0 values, the model accuracy loss is serious. At this time, the predicted Gaussian heat map can be transmitted to the CPU 102, and sigmoid and subsequent decoding are separately implemented in the CPU using c++ code. And the sigmoid layer, threshold filtering, and maxpool layer in the decoding process are integrated into a loop for processing, reducing the number of loops and better maintaining the locality of accessing memory, which can reduce the time consumption. Among them, during the threshold filtering (filtering the probability value through a preset threshold), different preset thresholds can be set for different smoking event objects. Through the processing of the CPU, the central position data and the detected annotation box size data of the smoking event object in the image to be detected are obtained to detect the position of the smoking event object in the image to be detected. In this way, the detection accuracy of the smoking event object can be further improved, and then the detection accuracy of the smoking event can be improved.
[0027] Generally, it is prohibited for drivers to smoke during vehicle driving. In some other specific scenarios, smoking events of personnel also need to be prohibited. Therefore, a solution for detecting smoking events needs to be provided.
[0028] The detection of smoking events often relies on the detection of the position of smoking event objects, such as the cigarette itself. Since the target size of the cigarette is relatively small compared to a person's hand, head, etc., and its shape is long and thin, the method of using related technologies to identify the cigarette is inaccurate, and the accuracy of detecting the position of the cigarette is low.
[0029] To solve the above technical problems, the embodiments of the present application provide a method for detecting the position of a smoking event object, a device for detecting the position of a smoking event object, an electronic device, and a computer-readable storage medium. Please refer to Figure 2 , Figure 2 is a flowchart of the method for detecting the position of a smoking event object shown in an exemplary embodiment of the present application. As Figure 2 shown, in an exemplary embodiment, this method can be applied to Figure 1 the implementation environment shown in Figure 1 and is specifically implemented by the edge chip in
[0030] Step S201: Obtain sample image data and the true annotation data of the sample image data.
[0031] Among them, the true annotation data is obtained by pre-annotating the smoking event objects in the sample image. The smoking event objects include, but are not limited to, at least one of the head, mouth, hand, cigarette, etc. The true annotation data includes the center position data of the true annotation box and the size data of the true annotation box, etc. The smoking event objects may also include, such as halos, bright spots, flames, etc. At this time, the true annotation data can be annotated by manual annotation or other annotation methods known to those skilled in the art. The accuracy of this true annotation data is high and can be used as the true value of the subsequent initial position detection model to calibrate the initial position detection model.
[0032] In one embodiment, before obtaining the sample image data, this method further includes:
[0033] Obtain multiple original images;
[0034] Detect and screen each original image through a preset human head detection model to obtain a sample image including a human head image. The sample image data includes the image data of one or more sample images.
[0035] In this embodiment, after detecting and screening each original image through a preset human head detection model, this method further includes:
[0036] Determine the density of human heads in the original image;
[0037] Expand the filtered intermediate image according to a preset expansion ratio based on the density of human heads, and use the expanded image as a sample image.
[0038] For example, the sample image data includes the image data of multiple sample images. The sample image and the subsequent to-be-detected image can, on the premise of complying with relevant laws and regulations and with the consent of relevant departments and personnel, obtain videos, images, etc. of people from street views and the Internet as the original images, and filter out the images containing human head images through a preset human head detection model.
[0039] Also, for example, to make the retention of the smoking event object included in the sample image more complete, the image containing the human head image output by the preset human head detection model can be expanded according to a preset expansion ratio, and the expanded image is used as a sub-sample image. Among them, the preset expansion ratio can be a value set by those skilled in the art, or can be determined by the density of human heads in the original image. Generally, it is expanded by 1.5 - 4 times depending on the density of human heads. For example, preset expansion ratios corresponding to different human head density gradients are set.
[0040] In one embodiment, during the process of annotating the smoking event object in the sample image, before converting the center position data of the true annotation box into a true elliptical Gaussian heat map, the smoking event object position detection method further includes:
[0041] If the smoking event object includes a cigarette, obtain the orientation of the cigarette butt, and determine the cigarette butt angle of the cigarette butt based on the true annotation box size data and the orientation of the cigarette butt, so as to adjust the initial elliptical Gaussian heat map determined based on the center position data of the true annotation box. Among them, the true annotation box size data includes the length and height (i.e., width) of the true annotation box.
[0042] Taking the orientation of the cigarette butt as left as an example, if the determination method of the cigarette butt angle tilted to the left includes:
[0043]
[0044] Among them, theta is the cigarette butt angle, w is the width of the true annotation box of the cigarette, and h is the height of the true annotation box of the cigarette.
[0045] Taking the orientation of the cigarette butt as left as an example, if the determination method of the cigarette butt angle tilted to the right includes:
[0046]
[0047] Among them, theta is the cigarette butt angle, w is the width of the true annotation box of the cigarette, and h is the height of the true annotation box of the cigarette.
[0048] Step S202: If the smoking event object includes any one of the head, mouth, and hand, convert the central position data of the true annotation box into a true circular Gaussian heat map. If the smoking event object includes a cigarette, convert the central position data of the true annotation box into a true elliptical Gaussian heat map, and use the true circular Gaussian heat map and the true elliptical Gaussian heat map as the true Gaussian heat map.
[0049] To make the smoking event object position detection method provided by the embodiments of the present application better applicable to scenarios such as edge chips, and to avoid the problem of serious loss of model accuracy caused by the quantization model of edge chips being sensitive to a large number of 0 values, in this embodiment, the output of the trained initial position detection model for the central position data of the annotation box can be in the form of a Gaussian heat map, rather than directly outputting the central position data of the annotation box. To better train the initial position detection model, the central position data of the true annotation box can be converted into a true Gaussian heat map for subsequent model training comparison.
[0050] The conversion method of the true circular Gaussian heat map can be implemented by the methods of related technologies, which will not be limited here. A conversion method of a true elliptical Gaussian heat map is as follows:
[0051] Since the loss of the hm heat map in the original FairMot and centernet is isotropic, and the contour lines of the heat map are circular. This is okay for large objects or when the area of the object without the map occupying the target box is very large, but for slender objects like cigarette butts, it is best to be consistent with the detected object, which can accelerate the convergence speed and training effect of model training. A determination method of a true elliptical Gaussian heat map is as follows:
[0052]
[0053] Among them, H is the true procrastination Gaussian heat map, x and y are the two-dimensional point coordinates of the true Gaussian ellipse with the semi-major axis R h , semi-minor axis R w in the true annotation box of the cigarette, and sigma w and sigma h are the variances of the Gaussian distribution, and exp represents the natural exponential function.
[0054] Among them, the determination method of the Gaussian distribution variance can be implemented with reference to the following formulas (2)-(9).
[0055]
[0056] D w = 2*R w + 1 Formula (3),
[0057]
[0058]
[0059] D h = 2 * R h +1 formula (6),
[0060]
[0061]
[0062]
[0063] where alpha is a preset scaling ratio, generally taken as 3 - 10, w and h are the width and height of the cigarette category target (which can be the true annotation box of the cigarette, etc.), beta is the preset aspect ratio of the cigarette butt, generally taken as 20 - 40. sigma w and sigma h are the variances of the Gaussian distribution. Then rotate the obtained initial Gaussian heatmap by theta degrees (the cigarette butt angle) to be consistent with the cigarette butt direction.
[0064] It should be noted that when the above - mentioned set of formulas (Formula 1 - Formula 9) for determining the true elliptical Gaussian heatmap is applied to determine the true elliptical Gaussian heatmap, the w and h data use the size data of the true annotation box of the cigarette (the true annotation box size data, including length and width),
[0065] Step S203, input the sample image data into the initial position detection model to obtain the predicted annotation data of the smoking event object.
[0066] It should be noted that in one embodiment, the execution sequence between Step S202 and Step S203 is not limited.
[0067] Among them, the predicted annotation data includes predicted annotation box size data and a predicted Gaussian heatmap. If the smoking event object includes any one of the head, mouth, and hand, the predicted Gaussian heatmap includes a predicted circular Gaussian heatmap. If the smoking event object includes a cigarette, the predicted Gaussian heatmap includes a predicted elliptical Gaussian heatmap. The predicted Gaussian heatmap is used to represent the center position data of the predicted annotation box. The prediction method of the predicted annotation box size data can be implemented in a manner known to those skilled in the art and is not limited herein.
[0068] In one embodiment, referring to Formulas (1) - (9), if the smoking event object includes a cigarette, converting the center position data of the true annotation box into a true elliptical Gaussian heatmap includes:
[0069] Determine the semi-major axis R of the Gaussian ellipse according to the preset scaling ratio alpha and the true annotation box size data h and the semi-minor axis R of the Gaussian ellipse w , and determine the variances of the Gaussian distribution (sigma w and sigma h ), generate the initial elliptical Gaussian heat map, and the determination method of the initial elliptical Gaussian heat map can refer to the above formulas (1)-(9), or be implemented in other ways known to those skilled in the art;
[0070] Rotate the initial elliptical Gaussian heat map according to the cigarette angle theta to obtain the true elliptical Gaussian heat map.
[0071] In one embodiment, continue to refer to formulas (1)-(9) to determine the semi-major axis of the Gaussian ellipse, the semi-minor axis of the Gaussian ellipse according to the preset scaling ratio and the true annotation box size data, and determine the variance of the Gaussian distribution. Generating the initial elliptical Gaussian heat map includes:
[0072] Determine the reference width L according to the true annotation box width w of the cigarette and the true annotation box height h of the cigarette w , and the true annotation box width and the true annotation box height of the cigarette are determined according to the true annotation box size data;
[0073] According to the reference width L w and the preset cigarette length-width ratio beta to determine the reference height L h ;
[0074] Based on the preset scaling ratio alpha, the reference width L w and the reference height L h respectively determine the semi-minor axis R of the Gaussian ellipse w and the semi-major axis R of the Gaussian ellipse h , and determine the short axis D of the Gaussian ellipse W and the long axis D of the Gaussian ellipse h ;
[0075] Based on the short axis D of the Gaussian ellipse W and the long axis D of the Gaussian ellipse h respectively determine the variance sigma of the short-axis sub-Gaussian distribution w and the variance sigma of the long-axis sub-Gaussian distribution h , as the variance of the Gaussian distribution;
[0076] Determine the initial elliptical Gaussian heat map according to the variance of the Gaussian distribution, the semi-minor axis of the Gaussian ellipse, the semi-major axis of the Gaussian ellipse, and the two-dimensional point coordinates in the true annotation box of the cigarette.
[0077] The center point of the cigarette butt target box is transformed into a rotating elliptical Gaussian heat map, making the predicted position more accurate. Transforming the center point of the cigarette butt target box into a rotating elliptical Gaussian heat map while keeping those of other categories unchanged allows the model to better learn the features of the cigarette butt, accelerate convergence, and achieve better model performance.
[0078] Step S204: Train the initial position detection model based on the predicted Gaussian heat map, the real Gaussian heat map, the real annotation box size data, and the predicted annotation box size data.
[0079] During the training of the initial position detection model, corresponding predicted Gaussian heat maps and real Gaussian heat maps are used for comparison according to different smoking event objects. For example, the predicted elliptical Gaussian heat map and the real elliptical Gaussian heat map of the cigarette are compared.
[0080] In one embodiment, before training the initial position detection model based on the predicted Gaussian heat map, the real Gaussian heat map, the real annotation box size data, and the predicted annotation box size data, the smoking event object position detection method further includes:
[0081] Generate a dimensional weight loss function according to the preset head weight, cigarette weight, hand weight, and mouth weight. The cigarette weight is greater than the head weight, hand weight, and mouth weight, and the head weight is less than the hand weight and mouth weight, so as to train the initial position detection model through the dimensional weight loss function.
[0082] In one embodiment, generating a dimensional weight loss function according to the preset head weight, cigarette weight, hand weight, and mouth weight includes:
[0083]
[0084] Where Loss is the dimensional weight loss function, cls takes values of 0, 1, 2, 3, representing the smoking event object as a cigarette, head, hand, and mouth respectively. Gama is the coordination factor of Loss, p^ is the predicted probability. If cls takes the value of 0, weight is the preset cigarette weight; if cls takes the value of 1, weight is the preset head weight; if cls takes the value of 2, weight is the preset hand weight; if cls takes the value of 3, weight is the preset mouth weight. Ln represents the natural logarithm function.
[0085] Since cigarette butts are very small targets and are easily interfered by white or reflective objects in the background (such as railings, suspension lines of masks), and the actual pixel range of cigarette butts in the picture only occupies a part of the cigarette butt frame, it is necessary to make special modifications to the network architecture of the detection model and the objective function of training. For example, by setting a weighted 2, Focal loss, setting the weight of the cigarette butt to be very large, the weight of the human head to be very small, and the weights of the hand, human mouth, and human head to be relatively small. This can make the training of the cigarette butt category converge faster and suppress overfitting of other categories.
[0086] Step S205: Obtain the image to be detected, input the image to be detected into the trained initial position detection model, and obtain the center position data of the detection annotation frame and the size data of the detection annotation frame of the smoking event object in the image to be detected, so as to detect the position of the smoking event object in the image to be detected.
[0087] Through the trained initial position detection model, more accurate center position data of the detection annotation frame and size data of the detection annotation frame of the smoking event object can be obtained. Furthermore, based on the center position data of the detection annotation frame and the size data of the detection annotation frame of each smoking event object, the position occlusion relationship of each smoking event object can be determined, and the smoking event of whether smoking is determined.
[0088] In one embodiment, as the post-processing logic 1 of the initial position detection model, the method further includes inputting the image to be detected into the trained initial position detection model, and the center position data of the detection annotation frame of the smoking event object obtained in the image to be detected includes:
[0089] Input the image to be detected into the trained initial position detection model to obtain the to-be-detected annotation data of the to-be-detected smoking event object in the image to be detected, and the to-be-detected annotation data includes a predicted Gaussian heat map;
[0090] Based on the predicted Gaussian heat map, determine the suspected to-be-detected position data of the center positions of multiple suspected annotation frames and the probability values of each suspected to-be-detected position data;
[0091] Filter the suspected to-be-detected position data based on the probability value and a preset threshold;
[0092] Based on the probability values of the filtered suspected to-be-detected position data, determine the center position data of the detection annotation frame of the smoking event object in the image to be detected.
[0093] The process of converting the predicted heat map into the data of the center position of the detection annotation box can be implemented on a processor such as a CPU to avoid the problem of serious loss of model accuracy and poor data accuracy caused by the accuracy problem of the edge chip itself and the sensitivity of the chip quantization model to a large number of 0 values when processed by the edge chip. This process can be implemented through a post-processing logic. During inference, no sigmoid layer is set after hm (the feature map of the center position of the annotation box) in the initial position detection model, while a sigmoid layer is retained after wh (the feature map of the annotation box size). Because of the accuracy problem of the edge chip itself and the sensitivity of the chip quantization model to a large number of 0 values, resulting in serious loss of model accuracy, the sigmoid after hm in the network can be not set, and the sigmoid and subsequent decoding are separately implemented in C++ code on the CPU. And the sigmoid layer, threshold filtering, and maxpool layer in the decoding process are integrated into a loop for processing, reducing the number of loops and better maintaining the locality of memory access, which can reduce the time consumption. Moreover, the NMS algorithm for the conventional predicted target boxes can be removed because maxpool and the threshold have actually filtered out most of the boxes that are very close. Where pool is the size of the kernel of maxpool, mat is the hm feature map, and threshold is the threshold. The threshold corresponding to each smoking event object can be set to different values. Here is only an example. The pseudo code is as follows:
[0094]
[0095]
[0096]
[0097] In one embodiment, as the post-processing logic 2 of the initial position detection model, after detecting the position of the smoking event object in the image to be detected, the method for detecting the position of the smoking event object includes:
[0098] Determine the position information of the detection target box according to the data of the center position of the detection annotation box and the data of the size of the detection annotation box;
[0099] Determine the smoking state based on the preset determination rule and the position information of the detection target box of the smoking event object in the image to be detected;
[0100] Wherein, the preset determination rule includes at least one of the following,
[0101] If the smoking event object in the image to be detected does not include a cigarette, determine the smoking state as determined not to be smoking;
[0102] If the smoking event objects in the image to be detected include a cigarette and at least one smoking-related object, and it is determined that there is an intersection between the cigarette and the smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of at least one smoking-related object, then the smoking status is determined to be smoking, and the smoking-related object includes at least one of the head, hand, and mouth;
[0103] If the smoking event objects in the image to be detected include a cigarette and at least one smoking-related object, and it is determined that there is no intersection between the cigarette and the smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of at least one smoking-related object, then the smoking status is determined to be non-smoking.
[0104] In the above manner, it is possible to determine whether it is smoking based on the positional relationship between the detected head target box, hand target box, mouth target box and the cigarette butt target, and eliminate unreasonable false detection situations.
[0105] For example, the following logical judgment is performed on the prediction result of an input image by the trained initial position detection model:
[0106] 1) If there is an intersection between the cigarette butt box and the human face detection box, and the area of the intersection is greater than A times the area of the cigarette butt box, it is determined to be smoking. Generally, A is taken between 0.2 and 0.5.
[0107] 2) If there is an intersection between the cigarette butt box and the hand detection box, and the area of the intersection is greater than B times the area of the cigarette butt box, it is determined to be smoking. Generally, B is taken between 0.2 and 0.5.
[0108] 3) If there is an intersection between the cigarette butt box and the mouth detection box, and the area of the intersection is greater than C times the area of the cigarette butt box, it is determined to be smoking. Generally, C is taken between 0.1 and 0.5.
[0109] 4) If there is no cigarette butt box, it is determined that there is no smoking. If none of the conditions 1), 2), and 3) are satisfied, it is determined that there is no smoking.
[0110] Because in a street view scene or a crowded scene, a certain human face may be incomplete, the human face may not be detected, only the mouth or hand of a person is detected, and this person may be smoking.
[0111] Please refer to Figure 3 , Figure 3 which is a partial structure schematic diagram of an initial position detection model. As Figure 3 shown, there are two network branch diagrams of the detection head during prediction. The left one is the heat map hm branch for predicting the center point position of the target box, and the right one is the width wh branch for predicting the width and height of the target box. This structure schematic diagram is only an example, and those skilled in the art can also adopt other known structures to implement the method in this embodiment.
[0112] The downsampling factor of the last feature layer of the conventional detection algorithm is more than 4 times, even 32 times, compared to the original image, but this will result in inaccurate positioning of particularly small targets. In the embodiments of the present application, the feature layer of the initial position detection model is downsampled by 4 times, and this layer fuses the feature information from downsampling by 4, 8, 16, and 32 times. Taking the DLA34 network in the FairMot paper as an example, the Tree of the DLA34 network is pruned. The original feature extraction depth levels([1,1,1,2,2,1]) is pruned to [1,1,1,1,1,1], and the original number of channels channels([16,32,64,128,256,512]) is pruned to [16,24,32,48,64,128]. After pruning, the parameters change from 61.1MB to 5.16MB, which is 1 / 11.8 times of the original. Although the measured accuracy and recall are 2-3 points lower than those before pruning, this can not only greatly reduce the parameters of the model, save memory, but also reduce the computational amount of the model, improve the inference speed, facilitate transplantation, and be better applicable to edge devices, such as camera chips and mobile phone chips.
[0113] In the embodiments of the present application, the centernet of the initial position detection model has two output heads, namely the center point feature map (heat map) hm of the target box and the size wh of the target box. Since the downsampling factor in the embodiments of the present application is small, and adding the offset reg of the center point of the target box makes the model more difficult to train on small targets such as cigarette butts, and the training effect is worse than that after removing reg, so the offset reg head of the center point of the target box is not used in the embodiments of the present application.
[0114] During inference, the input video or picture at the front end is input into the above model (initial position detection model) to obtain the output result 1 of the model; the hm branch in the result 1 passes through the post-processing logic 1 and then through decoding to restore to the coordinate point position of the original image, obtaining the output result 2; the wh branch in the result 1 passes through decoding to restore the original size of the image. The result before decoding is the normalized width and height, obtaining the true width and height of the predicted target box, denoted as the output result 3; the output result 2 and the result 3 are merged to obtain the predicted target box; finally, through the post-processing logic 2, it outputs whether smoking, and outputs the results such as cigarette butt, mouth, human head, hand box and confidence.
[0115] The method provided in the above embodiments can be regarded as including two parts: training and prediction. There are 4 categories of target boxes for training: cigarette, human head, hand, and mouth. The training method includes the following parts: composition of training data, model framework, and objective function. The overall implementation steps are as follows: 1) Obtain videos and images of people from street views and the Internet, use a human head detection model to screen out images containing human heads, and perform annotation. 2) After appropriately enhancing the annotated data, send it into the model for training until the loss reaches the convergence condition and then stop training. Among them, the center point coordinates of the detection box need to be converted into a heat map, and the midpoint of the center line of the cigarette butt detection box needs to be converted into an elliptical Gaussian heat map after rotation. 3) Accelerate model inference. Remove the sigmoid after the hm heat map branch of the trained model, and then quantize the model. 4) Fuse the sigmoid, maxpool, and threshold filtering in the subsequent steps of the hm heat map branch, which is denoted as post-processing logic 1. 5) Perform a logical judgment on the overlapping ratio of the boxes of the 4 categories obtained by model inference, and refer to the above post-processing logic 2.
[0116] The method for detecting the position of a smoking event object provided in this embodiment obtains sample image data and true annotation data obtained by annotating the smoking event object therein, converts the center position data of the true annotation box into a true Gaussian heat map, inputs the sample image data into the initial position detection model to obtain the predicted annotation box size data and predicted Gaussian heat map of the smoking event object, trains the initial position detection model based on the predicted Gaussian heat map, true Gaussian heat map, true annotation box size data, and predicted annotation box size data, obtains the image to be detected, and inputs it into the trained initial position detection model to obtain the center position data and detection annotation box size data of the detection annotation box of the smoking event object in the image to be detected, so as to detect the position of the smoking event object in the image to be detected. By converting the center position of the annotation box of the smoking event object in the shape of a slender cigarette into an elliptical Gaussian heat map for prediction, the trained initial position detection model can be made more suitable for the recognition of small targets such as cigarette butts, and the detection accuracy of the position of the cigarette is improved. It can also solve the problems of time-consuming multi-target inference, time-consuming nms (non-maximum suppression) for filtering redundant target boxes, difficult training of small targets such as cigarette butts, and poor effects.
[0117] Among them, the network of the initial position detection model mainly uses centernet. By pruning the network, after pruning, the number of parameters is small and the speed is fast. By modifying the objective function and the Gaussian heat map of the center point of the cigarette butt target box, it can be better applicable to the recognition of small targets such as cigarette butts. By fusing the post-processing of the network, the inference speed can be accelerated.
[0118] Please refer to Figure 4 ,Figure 4 It is a block diagram of a smoking event object position detection device shown in an exemplary embodiment of the present application. As Figure 4 shown, this embodiment provides a smoking event object position detection device 400, including:
[0119] A sample acquisition module 401, configured to acquire sample image data and true annotation data of the sample image data. The true annotation data is obtained by pre-annotating the smoking event object in the sample image. The smoking event object includes at least one of a head, a mouth, a hand, and a cigarette. The true annotation data includes true annotation box center position data and true annotation box size data;
[0120] A true map conversion module 402, configured to, if the smoking event object includes any one of a head, a mouth, and a hand, convert the true annotation box center position data into a true circular Gaussian heat map, and if the smoking event object includes a cigarette, convert the true annotation box center position data into a true elliptical Gaussian heat map, and use the true circular Gaussian heat map and the true elliptical Gaussian heat map as the true Gaussian heat map;
[0121] A model prediction module 403, configured to input the sample image data into an initial position detection model to obtain prediction annotation data of the smoking event object. The prediction annotation data includes prediction annotation box size data and a prediction Gaussian heat map. If the smoking event object includes any one of a head, a mouth, and a hand, the prediction Gaussian heat map includes a prediction circular Gaussian heat map. If the smoking event object includes a cigarette, the prediction Gaussian heat map includes a prediction elliptical Gaussian heat map. The prediction Gaussian heat map is used to represent the prediction annotation box center position data;
[0122] A model training module 404, configured to train the initial position detection model based on the prediction Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the prediction annotation box size data;
[0123] A position detection module 405, configured to acquire an image to be detected, input the image to be detected into the trained initial position detection model, and obtain the detection annotation box center position data and the detection annotation box size data of the smoking event object in the image to be detected, so as to detect the position of the smoking event object in the image to be detected.
[0124] In this embodiment, the device essentially sets multiple modules to execute the method in any of the above embodiments. For the specific functions and technical effects, refer to the above embodiments, and details are not described herein again.
[0125] See Figure 5 , this embodiment of the present invention also provides an electronic device 500, including a processor 501, a memory 502, and a communication bus 503;
[0126] The communication bus 503 is used to connect the processor 501 and the memory connection 502;
[0127] The processor 501 is configured to execute a computer program stored in the memory 502 to implement the method according to one or more of the above embodiments.
[0128] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored,
[0129] The computer program is used to cause a computer to execute the method according to any one of the above Embodiment 1.
[0130] An embodiment of the present application further provides a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can be caused to execute instructions (instructions) for the steps included in Embodiment 1 of the embodiments of the present application.
[0131] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code included on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0132] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device.
[0133] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0135] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology may modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the art within the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for detecting the position of a smoking event object, characterized in that The method includes: Obtaining sample image data and the true annotation data of the sample image data, where the true annotation data is obtained by pre-annotating the smoking event objects in the sample image, and the smoking event objects include at least one of a head, a mouth, a hand, and a cigarette. The true annotation data includes the central position data of the true annotation box and the size data of the true annotation box; If the smoking event object includes any one of a head, a mouth, and a hand, converting the central position data of the true annotation box into a true circular Gaussian heat map. If the smoking event object includes a cigarette, converting the central position data of the true annotation box into a true elliptical Gaussian heat map, and taking the true circular Gaussian heat map and the true elliptical Gaussian heat map as the true Gaussian heat map; Inputting the sample image data into an initial position detection model to obtain the predicted annotation data of the smoking event object, where the predicted annotation data includes the predicted annotation box size data and the predicted Gaussian heat map. If the smoking event object includes any one of a head, a mouth, and a hand, the predicted Gaussian heat map includes a predicted circular Gaussian heat map. If the smoking event object includes a cigarette, the predicted Gaussian heat map includes a predicted elliptical Gaussian heat map. The predicted Gaussian heat map is used to represent the central position data of the predicted annotation box; Training the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data; Obtaining a to-be-detected image, inputting the to-be-detected image into the trained initial position detection model to obtain the central position data of the detection annotation box and the size data of the detection annotation box of the smoking event object in the to-be-detected image, so as to detect the position of the smoking event object in the to-be-detected image; Before training the initial position detection model based on the predicted Gaussian heat map, the true Gaussian heat map, the true annotation box size data, and the predicted annotation box size data, the smoking event object position detection method further includes: Generating a dimension-weighted loss function according to preset head weights, cigarette weights, hand weights, and mouth weights, where the cigarette weight is greater than the head weight, the hand weight, and the mouth weight, and the head weight is less than the hand weight and the mouth weight, so as to train the initial position detection model through the dimension-weighted loss function; Generating a dimension-weighted loss function according to preset head weights, cigarette weights, hand weights, and mouth weights, including: wherein Among them, Loss is the loss function with dimensional weights. cls takes values of 0, 1, 2, and 3, representing that the smoking event object is a cigarette, head, hand, and mouth respectively. gama is the coordination factor of Loss, p^ is the predicted probability. If cls takes the value of 0, weight is the preset cigarette weight; if cls takes the value of 1, weight is the preset head weight; if cls takes the value of 2, weight is the preset hand weight; if cls takes the value of 3, weight is the preset mouth weight. ln represents the natural logarithm function.
2. The method for detecting the position of a smoking event object according to claim 1, wherein Before converting the central position data of the true annotation box into a true elliptical Gaussian heat map, the smoking event object position detection method further includes: If the smoking event object includes a cigarette, obtain the orientation of the cigarette butt, and determine the butt angle of the cigarette butt based on the true annotation box size data and the orientation of the cigarette butt, so as to adjust the initial elliptical Gaussian heat map determined based on the central position data of the true annotation box through the butt angle.
3. The method for detecting the position of a smoking event object according to claim 2, wherein, Converting the central position data of the true annotation box into a true elliptical Gaussian heat map includes: Determine the semi-major axis of the Gaussian ellipse, the semi-minor axis of the Gaussian ellipse, and determine the Gaussian distribution variance according to the preset scaling ratio and the true annotation box size data, and generate the initial elliptical Gaussian heat map; Rotate the initial elliptical Gaussian heat map according to the butt angle to obtain the true elliptical Gaussian heat map.
4. The method for detecting the position of a smoking event object according to claim 3, wherein Determining the semi-major axis of the Gaussian ellipse, the semi-minor axis of the Gaussian ellipse, and determining the Gaussian distribution variance according to the preset scaling ratio and the true annotation box size data, and generating the initial elliptical Gaussian heat map includes: Determine the reference width according to the true annotation box width of the cigarette and the true annotation box height of the cigarette. The true annotation box width of the cigarette and the true annotation box height of the cigarette are determined according to the true annotation box size data; Determine the reference height according to the reference width and the preset cigarette length-width ratio; Based on the preset scaling ratio, the reference width, and the reference height, determine the semi-minor axis of the Gaussian ellipse and the semi-major axis of the Gaussian ellipse respectively, and determine the short axis and the long axis of the Gaussian ellipse; Based on the short axis and the long axis of the Gaussian ellipse, determine the short-axis sub-Gaussian distribution variance and the long-axis sub-Gaussian distribution variance respectively as the Gaussian distribution variance; Determine the initial elliptical Gaussian heat map according to the Gaussian distribution variance and the two-dimensional point coordinates of the Gaussian ellipse semi-minor axis and the Gaussian ellipse semi-major axis in the true annotation box of the cigarette.
5. The method for detecting the position of a smoking event object according to any one of claims 1-4, characterized in that, Inputting the image to be detected into the trained initial position detection model to obtain the central position data of the detection annotation box of the smoking event object in the image to be detected includes: Input the image to be detected into the trained initial position detection model to obtain the to-be-detected annotation data of the to-be-detected smoking event object in the image to be detected. The to-be-detected annotation data includes a predicted Gaussian heat map; Based on the predicted Gaussian heat map, determine the suspected to-be-detected position data of the central positions of multiple suspected annotation boxes and the probability values of each of the suspected to-be-detected position data; Filter the suspected to-be-detected position data based on the probability value and the preset threshold; Based on the probability value of the filtered suspected position data to be detected, determine the center position data of the detection annotation box of the smoking event object in the image to be detected.
6. The method for detecting the position of a smoking event object according to any one of claims 1 to 4, characterized in that, After detecting the position of the smoking event object in the image to be detected, the smoking event object position detection method includes: Determine the position information of the detection target box according to the center position data of the detection annotation box and the size data of the detection annotation box; Determine the smoking state based on the preset determination rule and the position information of the detection target box of the smoking event object in the image to be detected; Among them, the preset determination rule includes at least one of the following: If the smoking event object in the image to be detected does not include a cigarette, determine the smoking state as not smoking; If the smoking event object in the image to be detected includes a cigarette and at least one smoking-related object, and it is determined that there is an intersection between the cigarette and the at least one smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of the at least one smoking-related object, determine the smoking state as smoking, and the smoking-related object includes at least one of the head, hand, and mouth; If the smoking event object in the image to be detected includes a cigarette and at least one smoking-related object, and it is determined that there is no intersection between the cigarette and the at least one smoking-related object based on the position information of the detection target box of the cigarette and the position information of the detection target box of the at least one smoking-related object, determine the smoking state as not smoking.
7. A smoking event object position detection device, characterized in that, The smoking event object position detection device includes: A sample acquisition module for acquiring sample image data and the true annotation data of the sample image data. The true annotation data is obtained by pre-annotating the smoking event object in the sample image. The smoking event object includes at least one of the head, mouth, hand, and cigarette. The true annotation data includes the center position data of the true annotation box and the size data of the true annotation box; A true image conversion module for converting the center position data of the true annotation box into a true circular Gaussian heat map if the smoking event object includes any one of the head, mouth, and hand, and converting the center position data of the true annotation box into a true elliptical Gaussian heat map if the smoking event object includes a cigarette, and taking the true circular Gaussian heat map and the true elliptical Gaussian heat map as the true Gaussian heat map; A model prediction module for inputting the sample image data into an initial position detection model to obtain the predicted annotation data of the smoking event object. The predicted annotation data includes the predicted annotation box size data and the predicted Gaussian heat map. If the smoking event object includes any one of the head, mouth, and hand, the predicted Gaussian heat map includes a predicted circular Gaussian heat map. If the smoking event object includes a cigarette, the predicted Gaussian heat map includes a predicted elliptical Gaussian heat map. The predicted Gaussian heat map is used to represent the center position data of the predicted annotation box. A model training module, configured to train the initial position detection model based on the predicted Gaussian heatmap, the ground-truth Gaussian heatmap, the ground-truth bounding box size data, and the predicted bounding box size data; and generate a dimension-weighted loss function according to preset head weights, cigarette weights, hand weights, and mouth weights, where the cigarette weight is greater than the head weight, the hand weight, and the mouth weight, and the head weight is less than the hand weight and the mouth weight, and train the initial position detection model through the dimension-weighted loss function; the process of generating the dimension-weighted loss function according to the preset head weights, cigarette weights, hand weights, and mouth weights includes: Among them Among them, Loss is a dimensionality-weighted loss function. The value of cls is 0, 1, 2, or 3, representing that the object of the smoking event is a cigarette, head, hand, or mouth respectively. Gama is the coordination factor of Loss, and p^ is the predicted probability. If the value of cls is 0, then weight is the preset cigarette weight; if cls is 1, weight is the preset head weight; if cls is 2, weight is the preset hand weight; if cls is 3, weight is the preset mouth weight. Ln represents the natural logarithm function; A position detection module, configured to obtain a to-be-detected image, input the to-be-detected image into the trained initial position detection model, and obtain the center position data and the size data of the detection bounding box of the smoking event object in the to-be-detected image, so as to detect the position of the smoking event object in the to-be-detected image.
8. An electronic device, characterized in that, Comprising a processor, a memory, and a communication bus; The communication bus is used to connect the processor and the memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, On which a computer program is stored, The computer program is used to cause the computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Smoking behavior detection method based on human body posture estimation and image classification
CN112528960A
Attribute recognition model training method and device, attribute recognition method and device and equipment
CN114445683A