A smoking behavior detection method and device
The method uses computer vision to accurately detect smoking behavior by identifying smoke within human body regions and calculating distances to predefined body parts, enhancing detection precision and preventing fires.
Patent Information
- Application Number
- CN202110144146.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-02-02
AI Technical Summary
The prior art has low universality in smoking behavior detection, especially in the case of smoking on the side of the human body, the false alarm and missed rate is high, the sensor sensitivity is insufficient, and it is not suitable for high dust or high temperature and high humidity environments. Deep learning methods have difficulties in sample collection and generalization capabilities.
Using computer vision technology, combined with deep learning and traditional image algorithms, through human body detection, smoke detection and smoke recognition, a multi-feature fusion algorithm is used to identify the human body area, smoke position and smoke in the image, and a strict distance measurement criteria and multi-task algorithm scheme are designed to achieve high-precision detection of front and side smoking.
It improves the accuracy and universality of smoking behavior detection, avoids misjudgment and misjudgment, can effectively identify smoking behaviors in various environments, provide timely warnings, and prevent fire risks.
Smart Images

Figure CN114842498B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a smoking behavior detection method and device. Background Art
[0002] Smoking behavior generally refers to the state in which a person holds a cigarette in his mouth and continuously performs the action of inhaling, which usually produces continuously spreading smoke. At present, there are two main ways to detect smoking behavior: one is to detect smoke through sensors, and the other is to directly analyze whether the image meets the characteristics of smoking based on computer vision technology, which is mainly divided into machine learning methods and deep learning methods. In the process of realizing the present invention, the inventor found that the prior art has at least the following problems:
[0003] 1. The smoke produced by smoking is generally not very significant, and it disappears quickly due to diffusion. The sensor is generally only sensitive when the smoke concentration is high enough, and is not suitable for environments with a lot of dust or high temperature and high humidity. In fact, there are many reasons for the generation of smoke, and even if smoke is detected, it cannot be concluded that there is smoking behavior.
[0004] 2. Machine learning methods are not very universal for artificial features, resulting in more false positives and missed negatives. Deep learning methods are difficult to collect samples because real-life scenarios usually involve a large number of people, and the people are often constantly changing. In addition, this algorithm is usually not suitable for pictures of people smoking from the side; smoking from the side cannot show the complete mouth area, so it is impossible to effectively distinguish whether there is smoking behavior through the mouth area. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a smoking behavior detection method and device, which can at least solve the problem that the prior art has low universality and is not suitable for side smoking detection of the human body.
[0006] To achieve the above object, according to one aspect of an embodiment of the present invention, a smoking behavior detection method is provided, comprising:
[0007] Determine the human body area in the image and detect whether there is smoke in each human body area;
[0008] If a cigarette is detected, the distance between the position of the cigarette and a preset part of the human body is identified and calculated. If the distance is less than or equal to a preset threshold, it is determined that there is a smoking behavior; wherein the preset part of the human body is the mouth or the tip of the nose;
[0009] If no smoke is detected or the distance is greater than the preset threshold, it is identified whether there is smoke in the image. If so, it is determined that there may be smoking behavior, otherwise it is determined that there is no smoking behavior.
[0010] Optionally, determining the human body region in the image further includes: for a single human body region, determining the width and height of the single human body region, and taking the product of the width and the expansion coefficient as the expansion value in the horizontal direction; based on the expansion value, expanding the single human body region left and right in the horizontal direction to obtain an expanded single human body region.
[0011] Optionally, after obtaining the expanded human body region, it further includes: based on the resolution of the image in the horizontal and vertical directions, correcting the boundary coordinates of the expanded single human body region to obtain an expanded and corrected single human body region.
[0012] Optionally, the width and height are the width and height of the smallest regular rectangle that can cover the single human body region.
[0013] Optionally, detecting whether there is a smoke body in each human body region includes: for a single human body region, using a first convolution module to extract the first image feature of the single human body region, and then using a first detection module to detect whether the first image feature contains a smoke body of a first size. If it contains, determine the smoke body position information; using a first transition module to perform a transition process on the first image feature, using a second convolution module to extract a second image feature from the transitioned first image feature, and then using a second detection module to detect whether the second image feature contains a smoke body of a second size. If it contains, determine the smoke body position information; using a second transition module to perform a transition process on the second image feature, using a third convolution module to extract a third image feature from the transitioned second image feature, and then using a third detection module to detect whether the third image feature contains a smoke body of a third size. If it contains, determine the smoke body position information; summarizing the above detection information to obtain the detection results and position information of smoke bodies of different sizes in the single human body region.
[0014] Optionally, it further includes: obtaining all the marked anchor box sizes, and obtaining multiple anchor box sizes with the largest clustering results through a clustering method; allocating the multiple anchor box sizes to the first detection module, the second detection module, and the third detection module, and identifying the smoke body region by fine-tuning the anchor box sizes.
[0015] Optionally, it further includes: using a human body key point algorithm to detect all human body key points in the image, and clustering each key point into the corresponding individual region to determine the preset part of each human body region; where the preset part is the tip of the nose.
[0016] Optionally, the identifying and calculating the distance between the position of the smoke body and the preset part of the human body, and if the distance is less than or equal to a preset threshold, it is determined that there is a smoking behavior, including: determining the coordinate value of the upper left corner of a single smoke body area, and calculating the distance between the upper left corner and the preset part of the human body in the vertical direction; wherein, the preset part is the tip of the nose; calculating the product of the width of a single human body area and a preset distance measurement coefficient, and if the distance is less than or equal to the product, it is determined that there is a smoking behavior in the single smoke body area.
[0017] Optionally, the identifying whether there is smoke in the image includes: based on a smoke recognition algorithm of multi-feature fusion, identifying the probability that the image has smoke; wherein, the smoke recognition algorithm of multi-feature fusion includes a deep learning algorithm and an image algorithm.
[0018] Optionally, the image algorithm includes a histogram of oriented gradients algorithm and local binary pattern; the identifying the probability that the image has smoke based on the smoke recognition algorithm of multi-feature fusion includes: using the deep learning algorithm to extract the feature map of the image, and converting it into a one-dimensional first vector with a first length through a flattening layer; wherein, the first length is the product of the sizes of the feature map in three dimensions; converting the first vector into a second vector with a second length through a first fully connected layer, and then converting it into a third vector with a third length through a second fully connected layer; using the local binary pattern to extract the vector of the image, and converting it into a fourth vector with a fourth length through a third fully connected layer; using the histogram of oriented gradients algorithm to extract the vector of the image, and converting it into a fifth vector with a fifth length through a fourth fully connected layer; fusing the third vector, the fourth vector and the fifth vector to obtain a total vector; wherein, the length of the total vector is the sum of the third length, the fourth length and the fifth length; converting the total vector into a seventh vector with a seventh length through a fifth fully connected layer, converting the seventh vector into an eighth vector with a length of 1 through a sixth fully connected layer, and taking the modulus of the eighth vector as the probability that there is smoke in the image.
[0019] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a smoking behavior detection device, including: a detection module, configured to determine the human body area in the image and detect whether there is a smoke body in each human body area; a distance judgment module, configured to, if a smoke body is detected, identify and calculate the distance between the position of the smoke body and the preset part of the human body, and if the distance is less than or equal to a preset threshold, determine that there is a smoking behavior; wherein, the preset part of the human body is the position of the mouth or the tip of the nose; a smoke recognition module, configured to, if no smoke body is detected or the distance is greater than the preset threshold, identify whether there is smoke in the image, and if there is, determine that there may be a smoking behavior, otherwise determine that there is no smoking behavior.
[0020] Optionally, the detection module is further configured to: for a single human body region, determine the width and height of the single human body region, and use the product of the width and the expansion coefficient as the horizontal expansion value; based on the expansion value, expand the single human body region horizontally to the left and right respectively to obtain the expanded single human body region.
[0021] Optionally, the detection module is configured to: based on the resolution of the image in the horizontal and vertical directions, correct the boundary coordinates of the expanded single human body region to obtain the expanded and corrected single human body region.
[0022] Optionally, the width and height are the width and height of the smallest regular rectangle that can cover the single human body region.
[0023] Optionally, the detection module is configured to: for a single human body region, use a first convolutional module to extract the first image feature of the single human body region, and then use a first detection module to detect whether the first image feature contains a smoke body of a first size. If it does, determine the smoke body position information; use a first transition module to perform a transition process on the first image feature, use a second convolutional module to extract a second image feature from the transitioned first image feature, and then use a second detection module to detect whether the second image feature contains a smoke body of a second size. If it does, determine the smoke body position information; use a second transition module to perform a transition process on the second image feature, use a third convolutional module to extract a third image feature from the transitioned second image feature, and then use a third detection module to detect whether the third image feature contains a smoke body of a third size. If it does, determine the smoke body position information; summarize the above detection information to obtain the detection results and position information of smoke bodies of different sizes in the single human body region.
[0024] Optionally, the detection module is further configured to: obtain all the marked anchor box sizes, and obtain multiple anchor box sizes with the largest clustering results through a clustering method; allocate the multiple anchor box sizes to the first detection module, the second detection module, and the third detection module, and identify the smoke body region by fine-tuning the anchor box sizes.
[0025] Optionally, the detection module is further configured to: adopt a human key point algorithm to detect all human key points in the image, and cluster each key point into the corresponding individual region to determine the preset part of each human body region; where the preset part is the tip of the nose.
[0026] Optionally, the spacing judgment module is configured to: determine the coordinate value of the upper left corner of a single cigarette body area, and calculate the spacing between the upper left corner and a preset part of the human body in the vertical direction; wherein, the preset part is the tip of the nose; calculate the product of the width of a single human body area and a preset distance measurement coefficient, and if the spacing is less than or equal to the product, it is determined that there is a smoking behavior in the single cigarette body area.
[0027] Optionally, the smoke recognition module is configured to: based on a smoke recognition algorithm integrating multiple features, recognize the probability that the image has smoke; wherein, the smoke recognition algorithm integrating multiple features includes a deep learning algorithm and an image algorithm.
[0028] Optionally, the image algorithm includes a histogram of oriented gradients algorithm and local binary patterns; the smoke recognition module is configured to: use the deep learning algorithm to extract the feature map of the image, and convert it into a one-dimensional first vector with a first length through a flattening layer; wherein, the first length is the product of the sizes of the feature map in three dimensions; convert the first vector into a second vector with a second length through a first fully connected layer, and then convert it into a third vector with a third length through a second fully connected layer; use the local binary patterns to extract the vector of the image, and convert it into a fourth vector with a fourth length through a third fully connected layer; use the histogram of oriented gradients algorithm to extract the vector of the image, and convert it into a fifth vector with a fifth length through a fourth fully connected layer; fuse the third vector, the fourth vector and the fifth vector to obtain a total vector; wherein, the length of the total vector is the sum of the third length, the fourth length and the fifth length; convert the total vector into a seventh vector with a seventh length through a fifth fully connected layer, convert the seventh vector into an eighth vector with a length of 1 through a sixth fully connected layer, and use the modulus of the eighth vector as the probability that there is smoke in the image.
[0029] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided an electronic device for detecting smoking behavior.
[0030] The electronic device according to the embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the smoking behavior detection method described in any one of the above.
[0031] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the smoking behavior detection method described in any one of the above is implemented.
[0032] According to the solution provided by the present invention, one embodiment of the above invention has the following advantages or beneficial effects: Based on computer vision technology, a high-precision smoking behavior detection algorithm solution applicable to both frontal and lateral smoking is established. This solution comprehensively considers the strict definition of smoking behavior and the accompanying smoke phenomenon, utilizes various algorithm ideas of deep learning, establishes a rigorous algorithm logic, accurately judges smoking behavior, avoids various misjudgments and missed judgments, and achieves the purpose of effectively preventing fires.
[0033] The further effects of the above non-conventional optional methods will be described below in combination with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0035] Figure 1 is a main flowchart of a smoking behavior detection method according to an embodiment of the present invention;
[0036] Figure 2 is a flowchart of a human body region detection method according to an embodiment of the present invention;
[0037] Figure 3 is a flowchart of a smoke body detection method according to an embodiment of the present invention;
[0038] Figure 4 is a network structure diagram of a smoke body detection algorithm;
[0039] Figure 5 is a flowchart of a smoke recognition method according to an embodiment of the present invention;
[0040] Figure 6 is a network structure diagram of a multi-feature fusion algorithm;
[0041] Figure 7(a) is a flowchart of an algorithm for smoking behavior detection;
[0042] Figure 7(b) is a schematic diagram of the algorithm logic of distance measurement;
[0043] Figures 8(a) and (b) are two smoking behavior recognition results;
[0044] Figure 9 is a main module diagram of a smoking behavior detection device according to an embodiment of the present invention;
[0045] Figure 10 is an exemplary system architecture diagram to which the embodiment of the present invention can be applied;
[0046] Figure 11It is a schematic diagram of the structure of a computer system of a mobile device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0048] Smoking may lead to fires, especially in warehouses, stations, distribution centers and other places where there are many flammable materials. Fires often cause significant economic losses and even endanger personal safety. In order to prevent and eliminate fires, smoking is prohibited in many places. However, in reality, some people often ignore relevant laws and regulations and smoke in inappropriate environments. Therefore, it is of great significance to implement automatic detection of smoking in some important places to detect and prevent smoking as early as possible, which is of great significance to the safe production of enterprises. Here is a detailed description of existing computer vision technology and its shortcomings:
[0049] There are two main methods for detecting smoking behavior: one is to detect smoke through sensors. Smoke alarms are a widely used automatic smoke detection technology. They detect smoke that may exist in the surrounding environment through built-in sensors and send out alarm signals in time. However, they are only highly sensitive when the smoke concentration is high enough, so they are not very useful for early warning of smoking behavior. The other method is to directly analyze whether the image meets the characteristics of smoking based on computer vision technology. By using the surveillance images monitored by the camera as the input of the relevant algorithm, possible smoking behavior can be detected without interference.
[0050] The application of computer vision technology in smoking behavior detection is mainly divided into machine learning methods and deep learning methods.
[0051] 1) Machine learning methods are generally based on artificially designed image feature extraction algorithms, combined with certain detection logic to determine whether the image meets the characteristics of smoking. However, these artificial features are not very universal, resulting in more false positives and false negatives.
[0052] 2) Deep learning methods have achieved rapid development in recent years, and various excellent algorithms are constantly emerging. Currently, the application of deep learning in smoking detection mainly uses classification algorithms, which perform binary classification on the entire image or the area near the human mouth, and output the conclusion of whether there is smoking in the image. The main problem of the classification algorithm is that if the model is to have strong generalization ability, the sample coverage must be wide enough; in addition, the samples generally need to include both the smoking and non-smoking states of relevant personnel in the frontal situation at the same time.
[0053] Since real-world scenarios usually involve a large number of people and the people are constantly changing, it is very difficult to collect samples. In addition, this algorithm is usually not applicable to the situation of pictures of people smoking on the side; smoking on the side cannot show the complete mouth area, so it is impossible to effectively distinguish whether there is a smoking behavior through the mouth area. In short, the recognition rate of the binary classification method is often not very high, and the smoking behavior detection algorithm based on deep learning still needs to be further improved.
[0054] See Figure 1 , which shows the main flowchart of a smoking behavior detection method provided by an embodiment of the present invention, including the following steps:
[0055] S101: Determine the human body area in the image, and detect whether there is a smoke body in each human body area;
[0056] S102: If a smoke body is detected, identify and calculate the distance between the position of the smoke body and a preset part of the human body. If the distance is less than or equal to a preset threshold, it is determined that there is a smoking behavior; where the preset part of the human body is the mouth or the tip of the nose;
[0057] S103: If no smoke body is detected or the distance is greater than the preset threshold, identify whether there is smoke in the image. If there is, it is determined that there may be a smoking behavior, otherwise it is determined that there is no smoking behavior.
[0058] In the above implementation, for step S101, after the image is collected, the human body area in the image is first detected by a human body detection algorithm (see the description shown later Figure 2 ), and for each human body area, it is detected whether there is a smoke body, and the possible position of the smoke body is output (see the description shown later Figure 3 ). If the human body area contains a smoke body, the mouth position of the human body is further detected, that is, the area directly contacted by the smoke body. However, since the mouth area usually has a large coverage range, the point of the tip of the nose adjacent to it is selected as an alternative.
[0059] The detection of the nose tip mainly uses the human key point algorithm, which will simultaneously detect some key points of the human body, including the nose tip, left eye, right eye, left ear, right ear, left shoulder, right shoulder, etc. This algorithm consists of two processes: key point detection (in addition to including the main human joint points, it also includes some key positions such as the nose tip) and key point clustering, that is, first detect all human key points in the image, and then cluster these key points into the corresponding individual regions respectively. It should be noted that the human key point algorithm does not use ordinary clustering algorithms, which contains some special strategies and will assign each key point to the corresponding personal region.
[0060] For step S102, the judgment of the smoking behavior is based on the distance between the nose tip and the cigarette body. If this distance is less than the preset threshold, it is determined that there is a smoking behavior. For a human body region, assuming that the coordinate of the nose tip position output by the human key point detection algorithm is (qx, qy), and the upper left corner coordinate of the j-th cigarette body region output by the cigarette body detection algorithm is (s[j][0], s[j][1]), and the lower right corner coordinate is (s[j][2], s[j][3]). For each human body region in the image, with the width and height of the k-th human body region being w k and h k (after external expansion and correction as w' k and h' k ), the general judgment criterion for smoking behavior is as follows:
[0061] |s[j][1] - q y | < βw' k
[0062] where β is the distance metric coefficient, and its value is selected according to the actual situation, generally between 0 and 1, such as 0.3. However, if it is far from the mouth / nose tip position, such as holding a cigarette in the hand, it can be appropriately increased. Since the width of the human body region is generally smaller than the height, the horizontal distance between the cigarette body and the nose tip is usually relatively small. Therefore, the judgment criterion mainly calculates whether the vertical distance between the two is less than or equal to the preset threshold. On the other hand, to ensure that the judgment criterion is applicable to various image resolutions of different human body regions, the relative distance between the two is of general significance. Usually, the picture generally contains the complete human body in the horizontal direction, while in the vertical direction, it may only contain part of the body, such as the upper body. Therefore, the judgment criterion selects the width w' k of the human body region as the reference standard, and the distance judgment operation needs to be performed for each human body region.
[0063] For step S103, distance measurement is a relatively strict determination of smoking behavior. In reality, it is also necessary to give early warnings for situations where smoking behavior may exist, such as when both people and smoke are present in the image. Therefore, if the human body in the image can be detected in the distance measurement scheme, but the smoke body is not detected or the distance measurement criterion is not met, then smoke recognition is further performed on the image.
[0064] This solution combines the advantages of deep learning algorithms and traditional image algorithms, and uses a smoke recognition algorithm based on multi-feature fusion to identify whether there is smoke in the image. Image classification algorithms based on deep learning usually have strong capabilities for extracting deep-level features of images, but these features are not interpretable and are even close to being a "black box"; at the same time, due to the discreteness and variability of the smoke form, deep learning models often have insufficient generalization ability. On the other hand, the smoke area usually presents a special texture structure, and traditional image algorithms specifically extract local features in certain aspects of the image and have certain advantages in capturing smoke textures. Therefore, the multi-feature fusion model will comprehensively utilize the feature information of different aspects of the image extracted by deep learning algorithms and traditional image algorithms, and can effectively improve the accuracy of various smoke recognition problems.
[0065] Based on computer vision technology, the above embodiments construct a high-precision smoking behavior detection algorithm solution that is applicable to both front and side smoking. This solution comprehensively considers the strict definition of smoking behavior and the accompanying smoke phenomenon, uses various algorithm ideas of deep learning, and establishes a rigorous algorithm logic to accurately judge smoking behavior, avoid various misjudgments and missed judgments, and achieve the purpose of effectively preventing fires.
[0066] See Figure 2 , which shows a schematic flowchart of a human body region detection method according to an embodiment of the present invention, including the following steps:
[0067] S201: Detect the human body in the image to output each region where the human body exists;
[0068] S202: For a single human body region, determine the width and height of the single human body region, and take the product of the width and the expansion coefficient as the expansion value in the horizontal direction;
[0069] S203: Based on the expansion value, expand the single human body region to the left and right in the horizontal direction respectively to obtain the expanded single human body region.
[0070] In the above embodiment, for step S201, this embodiment mainly describes detecting the human body in the image through a human body detection algorithm and outputting each region where the human body exists, so as to limit the detection of the smoke body within the human body region to exclude the interference of irrelevant background regions in the image on the detection of the smoke body.
[0071] In an actual scenario, there may be many long and strip-shaped objects similar to cigarettes. Human body detection essentially belongs to the category of object detection in computer vision. Based on some public datasets containing human body targets (such as COCO, etc.) and personal collected data, the human body regions in the images can be labeled, so as to train a human body detection model.
[0072] Human body detection algorithms usually only strictly output the human body regions in the pictures. However, in some cases, the cigarette body may extend outside the human body, such as smoking on the side. Therefore, to ensure that the possible cigarette bodies are not lost in the output human body regions, it is necessary to expand a certain degree horizontally to the left and right respectively on the basis of the algorithm output of the human body regions. And when a person is in the smoking state, the cigarette generally cannot exceed the person's head. Therefore, there is no need to expand vertically.
[0073] Suppose the width and height of the region occupied by the i-th person in the image are w i and h i , and the upper left corner coordinates of the region are (p[i][0], p[i][1]), and the lower right corner coordinates are (p[i][2], p[i][3]). For the expanded human body region, the calculation methods of its minimum coordinate p[i][0] and maximum coordinate p[i][2] in the horizontal direction are as follows:
[0074] p[i][0] = p[i][0] - αw i
[0075] p[i][2] = p[i][2] + αw i
[0076] Among them, α is the expansion coefficient, generally preferably about 0.1. All human body regions in the image can share one expansion coefficient. It should be noted that usually the human body regions in the image are not regular rectangles. Therefore, the w i and h i here refer to the width and height of the smallest regular rectangle that can cover the complete human body region.
[0077] However, the human body regions in the image may have relatively large sizes. The expansion of the human body regions may cause their boundary coordinates to exceed the coordinate range of the image, resulting in the failure of the output of the corresponding regions. Therefore, it is necessary to correct the boundary coordinates of the expanded human body regions. The specific scheme is as follows:
[0078] p[i][0] = 0, if p[i][0] < 0
[0079] p[i][2] = W, if p[i][2] > W
[0080] Wherein, W is the horizontal resolution of the image. Through the above coordinate adjustment, it is ensured that the finally expanded human body area still falls within the range of the original image.
[0081] The method provided by the above embodiment is based on the human body area output by the human body detection algorithm, expands it and adjusts the coordinates of the boundary area, and can exclude the interference of similar smoke substances in the background on the subsequent smoke detection.
[0082] See Figure 3 , which shows a schematic flowchart of a smoke detection method according to an embodiment of the present invention, including the following steps:
[0083] S301: For a single human body area, use the first convolution module to extract the first image feature of the single human body area, and then use the first detection module to detect whether the first image feature contains a smoke body of a first size. If it contains, determine the smoke body position information;
[0084] S302: Use the first transition module to perform a transition process on the first image feature, use the second convolution module to extract the second image feature from the transitioned first image feature, and then use the second detection module to detect whether the second image feature contains a smoke body of a second size. If it contains, determine the smoke body position information;
[0085] S303: Use the second transition module to perform a transition process on the second image feature, use the third convolution module to extract the third image feature from the transitioned second image feature, and then use the third detection module to detect whether the third image feature contains a smoke body of a third size. If it contains, determine the smoke body position information;
[0086] S304: Summarize the above detection information to obtain the detection results and position information of smoke bodies of different sizes in the single human body area.
[0087] In the above embodiment, for step S301, for each corrected human body area, further identify the smoke body therein through a smoke detection model. Since the detection target is the smoke body in the human body area, the training data of the smoke detection model is also processed by the human body detection algorithm to output the human body area picture therein, and then the smoke body is marked in the human body area, or directly marked in the original picture. Theoretically, the effect of the former algorithm is better, and it can exclude the interference of some irrelevant areas on the smoke detection. Therefore, the former method is preferred.
[0088] To ensure that the algorithm can identify smoke bodies of different sizes, the present invention designs a multi-size detection algorithm scheme, and the specific network structure is as Figure 4 shown:
[0089] Backbone is the backbone network, mainly used for the extraction of image features;
[0090] Conv 1a, Conv 1b, and Conv 1c are three convolutional modules used to extract image features at different scales. Three (the number is only an example) anchor boxes are used for each size. The sizes of the anchor boxes are obtained through a clustering algorithm. Therefore, this design can basically ensure that the sizes of most smoke bodies are covered.
[0091] Trans 1a and Trans 1b are corresponding transition modules used to further process the image features extracted in the previous step by other means, such as changing the number of channels of the features, upsampling, etc. Its output is the processed image features, which are used as the input of the subsequent convolutional modules, helping to improve the overall prediction accuracy of the algorithm.
[0092] Conv 2a, Conv 2b, and Conv 2c are three detection modules of different scales used to detect smoke bodies of different sizes and their positions.
[0093] The input feature maps of them gradually decrease, but the receptive fields gradually become larger. Therefore, they are respectively suitable for detecting smaller, medium-sized, and larger smoke bodies. By summarizing the results of the three detection modules, the sizes and positions of various smoke bodies in the image (not the three types of features) are obtained, and the final smoke body detection conclusion is output. If there is no smoke body in the summary result, the image does not contain a smoke body.
[0094] In addition, based on all the labeled anchor box sizes (i.e., the width and height of the rectangle) in the training data, the n anchor box sizes with the highest probability are obtained through the K-Means clustering algorithm, such as n equals 9. These n anchor boxes are respectively assigned to the above three different detection layers, preferably evenly distributed, that is, the number of anchor boxes in each detection layer is n / 3.
[0095] The loss function, which belongs to the basic knowledge of computer vision object detection algorithms, is mainly used to quantify the deviation of the prediction result relative to the actual situation. It consists of the following parts:
[0096] Loss = L box + L cls + L obj
[0097] Among them, L box is the loss caused by the deviation of the target box size and position, L cls is the loss caused by the wrong prediction category, and L obj is the loss caused by the target confidence. As mentioned above, the algorithm pre-sets 9 anchor box sizes, but it is basically impossible for the actual smoke body size to be exactly equal to these 9 anchor box sizes. By fine-tuning the pre-set 9 anchor box sizes, the anchor box sizes are made to be as close as possible to the actual smoke body size to achieve the recognition of the smoke body area in the image.
[0098] In the method provided by the above embodiments, single-stage object detection algorithm is used for smoke detection, that is, smoke detection is regarded as a regression analysis problem of smoke position and category information, and the detection result is directly output through a neural network model, with relatively high calculation efficiency and good real-time performance.
[0099] See Figure 5 , which shows a schematic flowchart of a smoke recognition method according to an embodiment of the present invention, including the following steps:
[0100] S501: Use a deep learning algorithm to extract the feature map of the image and convert it into a one-dimensional first vector with a first length through a flattening layer; wherein, the first length is the product of the sizes of the feature map in three dimensions;
[0101] S502: Convert the first vector into a second vector with a second length through a first fully connected layer, and then convert it into a third vector with a third length through a second fully connected layer;
[0102] S503: Use local binary pattern to extract the vector of the image and convert it into a fourth vector with a fourth length through a third fully connected layer;
[0103] S504: Use histogram of oriented gradients algorithm to extract the vector of the image and convert it into a fifth vector with a fifth length through a fourth fully connected layer;
[0104] S505: Fuse the third vector, the fourth vector and the fifth vector to obtain a total vector; wherein, the length of the total vector is the sum of the third length, the fourth length and the fifth length;
[0105] S506: Convert the total vector into a seventh vector with a seventh length through a fifth fully connected layer, convert the seventh vector into an eighth vector with a length of 1 through a sixth fully connected layer, and use the modulus of the eighth vector as the probability of the existence of smoke in the image.
[0106] In the above embodiment, for step S501, the deep learning algorithm and the traditional algorithm in the multi-feature fusion algorithm can be selected according to aspects such as actual scene features, algorithm accuracy requirements, and algorithm real-time performance requirements.
[0107] Deep learning algorithms include ResNet, VGG, GoogLeNet, etc.; related traditional image algorithms include the LBP (Local Binary Pattern) algorithm, the HOG (Histogram of Oriented Gradient) algorithm, etc. Among them, the LBP algorithm has high sensitivity to the local texture features of each region of the image and has advantages such as gray-scale invariance and rotation invariance. The HOG algorithm describes the contour features inside the image by statistically calculating the histogram of oriented gradients in the local area of the image.
[0108] Figure 6 A schematic diagram of a network structure based on a multi-feature fusion algorithm is given. The deep learning algorithm uses VGG, and the traditional algorithms selected are LBP and HOG. Dense is a fully connected layer, and its parameter represents the number of neurons, corresponding to the vector dimension; Flatten is a flattening layer, and its function is to convert a multi-dimensional feature map into a one-dimensional feature vector; Concatenate is a concatenation layer, and its function is to concatenate several feature vectors in sequence.
[0109] 1) The size of the feature map extracted by the VGG algorithm for the image is (a, b, c), and it is flattened into a one-dimensional feature vector through the Flatten layer, with a length of n1:
[0110] Feature(a,b,c)→Vetor(1,1,abc)=Vetor(1,1,n1), n1=abc
[0111] 2) After passing through two fully connected layers Dense(n2) and Dense(n3), the length of the one-dimensional feature vector is converted to n3.
[0112] 3) The image feature vectors extracted by the LBP algorithm and the HOG algorithm are respectively converted into feature vectors with lengths of n4 and n5 through the fully connected layers Dense(n4) and Dense(n5).
[0113] 4) Through the Concatenate layer, the feature vectors extracted by the above three algorithms are fused to obtain the total feature vector, with a length of n6:
[0114] Vetor(1,1,n3)+Vetor(1,1,n4)+Vetor(1,1,n5)=Vetor(1,1,n6), n6=n3 + n4 + n5
[0115] 5) Finally, through the conversion of the fully connected layers Dense(n7) and Dense(1), a feature vector with a length of 1 is obtained, and the modulus of this vector is the probability of the existence of smoke in the image, that is, the recognition result.
[0116] In addition, to prevent overfitting problems in network training, Dropout layers are further added after the two fully connected layers of Dense(n2) and Dense(n7) to randomly inactivate some neurons. The loss function is selected as the binary cross-entropy function:
[0117]
[0118] where y i is the class label of sample i, which is 1 if the sample contains smoke and 0 otherwise. p(y i ) is the probability that sample i contains smoke.
[0119] The method provided by the above embodiment can further reveal possible smoking behaviors by identifying smoke in the image, thus comprehensively reflecting the smoking situation on-site and avoiding false negatives.
[0120] Referring to FIG. 7(a), a schematic algorithm flowchart for detecting smoking behavior according to an embodiment of the present invention is shown, which specifically includes the following aspects:
[0121] Adopt a multi-task strategy, that is, combine two detection methods of distance measurement and smoke recognition to jointly determine whether there is a smoking behavior.
[0122] 1) The distance measurement method is shown in FIG. 7(b), and the logic mainly includes several parts: 1) First, detect the human body in the image through a human body detection algorithm and output the regions where the human body exists;
[0123] 2) For each human body region, detect whether there is a smoke body through a smoke body detection model and output the possible positions of the smoke body. If the human body region contains a smoke body, further detect the mouth region of the human body. 3) Finally, according to the distance between the position of the smoke body and the position of the mouth / nose tip, output the conclusion of whether there is a smoking behavior in the image.
[0124] 2) If the conclusion of the distance measurement is that there is a smoking behavior, directly output the final result. Otherwise, on the premise of detecting relevant elements such as the human body, but no smoke body is detected or the distance measurement criterion is not met, the "smoke recognition" method is used to further detect the smoke in the image. If the result is that there is smoke, it is prompted that there may be a smoking behavior.
[0125] 3) Based on the previous results, output the recognition conclusion of the entire algorithm for the input image, that is, "there is a smoking behavior", "there may be a smoking behavior", or "there is no smoking behavior".
[0126] The recognition result of Figure 8(a) is "smoking behavior exists", where the cigarette body is next to the person's mouth, meeting the "distance metric" criterion. The recognition result of Figure 8(b) is "smoking behavior may exist"; this picture does not meet the "distance metric" criterion, but the conclusion of the "multi-feature fusion smoke recognition" module is that smoke exists. The test results are consistent with the actual situation. Therefore, the overall algorithm solution can well recognize various features of smoking behavior in the image, so as to provide comprehensive information on smoking phenomena in the on-site environment in a timely manner. These two pictures here are both from network retrieval.
[0127] Based on the above algorithm solution, corresponding picture data are collected to train the human body detection model, cigarette body detection model, and human body key point detection model respectively; smoke pictures and background pictures are collected to train the multi-feature fusion model for smoke recognition. The computational complexity of deep learning is usually large. The GPU (Graphics Processing Unit) device of NVIDIA Corporation contains a large number of computing units. GPU computing based on CUDA can effectively accelerate the training speed of the model and the prediction speed based on the model. One model of GPU is P40. It is based on the advanced Pascal system architecture, with 3840 computing cores (CUDA Cores), and the single-precision computing performance reaches 12 TeraFLOPS; at the same time, this device has a large video memory (24GB) and video memory bandwidth (346GB / s), which is conducive to data transmission between the CPU and the GPU.
[0128] The method provided by the embodiments of the present invention aims at the related problems of existing computer vision algorithms in smoking behavior detection. Based on the strict definition of smoking behavior and the possible accompanying smoke phenomenon, a complete set of smoking warning algorithm solutions is designed. The beneficial effects are as follows:
[0129] 1) It can simultaneously recognize the general cigarette body detection of smoking from the front and side. Based on the human body area output by the human body detection algorithm, expand it and adjust the boundary area coordinates, so as to exclude the interference of similar cigarette body substances in the background and ensure that cigarette bodies in various situations can be completely retained.
[0130] 2) Based on the distance metric strategy, a strict and general smoking behavior determination criterion is proposed. Through nose positioning based on the human body key point detection algorithm and cigarette body positioning based on the target detection algorithm, this criterion only calculates the distance between the nose and the cigarette body in the vertical direction, and converts it into a relative distance through the width of the human body area, so as to be applicable to various human body area resolutions.
[0131] 3) A multi-task algorithm solution combining distance metric and smoke recognition is proposed. Based on this strict determination of distance metric, by recognizing the smoke in the image, the possible smoking behavior is further revealed, so as to comprehensively reflect the smoking situation on the site and avoid false negatives.
[0132] 4) Propose a smoke recognition algorithm based on the idea of multi-feature fusion. Combining the respective advantages of deep learning and traditional image algorithms, this algorithm correlates the feature vectors output by both and then jointly optimizes them to improve the smoke recognition ability for various smoking situations.
[0133] See Figure 9 , which shows a schematic diagram of the main modules of a smoking behavior detection device 900 provided by an embodiment of the present invention, including:
[0134] A detection module 901, configured to determine the human body region in the image and detect whether there is a smoke body in each human body region;
[0135] A spacing judgment module 902, configured to, if a smoke body is detected, identify and calculate the spacing between the position of the smoke body and a preset part of the human body. If the spacing is less than or equal to a preset threshold, it is determined that there is a smoking behavior; wherein, the preset part of the human body is the position of the mouth or the tip of the nose;
[0136] A smoke recognition module 903, configured to, if no smoke body is detected or the spacing is greater than the preset threshold, identify whether there is smoke in the image. If there is, it is determined that there may be a smoking behavior, otherwise it is determined that there is no smoking behavior.
[0137] In the implementation device of the present invention, the detection module 901 is further configured to:
[0138] For a single human body region, determine the width and height of the single human body region, and use the product of the width and the expansion coefficient as the expansion value in the horizontal direction;
[0139] Based on the expansion value, expand the single human body region to the left and right in the horizontal direction respectively to obtain an expanded single human body region.
[0140] In the implementation device of the present invention, the detection module 901 is configured to: Based on the resolution of the image in the horizontal and vertical directions, correct the boundary coordinates of the expanded single human body region to obtain an expanded and corrected single human body region.
[0141] In the implementation device of the present invention, the width and height are the width and height of the smallest regular rectangle that can cover the single human body region.
[0142] In the implementation device of the present invention, the detection module 901 is configured to:
[0143] For a single human body region, use a first convolution module to extract the first image feature of the single human body region, and then use a first detection module to detect whether the first image feature contains a smoke body of a first size. If it contains, determine the smoke body position information;
[0144] Use the first transition module to perform transition processing on the first image feature, use the second convolution module to extract the second image feature from the transitioned first image feature, and then use the second detection module to detect whether the second image feature contains a smoke body of a second size. If it contains, determine the smoke body position information;
[0145] Use the second transition module to perform transition processing on the second image feature, use the third convolution module to extract the third image feature from the transitioned second image feature, and then use the third detection module to detect whether the third image feature contains a smoke body of a third size. If it contains, determine the smoke body position information;
[0146] Summarize the above detection information to obtain the detection results and position information of smoke bodies of different sizes in the single human body region.
[0147] In the implementation device of the present invention, the detection module 901 is further configured to: obtain all the marked anchor box sizes, and obtain multiple anchor box sizes with the largest clustering results through a clustering method; allocate the multiple anchor box sizes to the first detection module, the second detection module, and the third detection module, and identify the smoke body region by fine-tuning the anchor box sizes.
[0148] In the implementation device of the present invention, the detection module 901 is further configured to: adopt a human body key point algorithm to detect all human body key points in the image, and cluster each key point into the corresponding individual region to determine the preset part of each human body region; where the preset part is the tip of the nose.
[0149] In the implementation device of the present invention, the spacing judgment module 902 is configured to:
[0150] Determine the coordinate value of the upper left corner of a single smoke body region, and calculate the spacing between the upper left corner and the preset part of the human body in the vertical direction; where the preset part is the tip of the nose;
[0151] Calculate the product of the width of a single human body region and the preset distance measurement coefficient. If the spacing is less than or equal to the product, it is determined that there is a smoking behavior in the single smoke body region.
[0152] In the implementation device of the present invention, the smoke recognition module 903 is configured to: based on a smoke recognition algorithm of multi-feature fusion, recognize the probability that there is smoke in the image; where the smoke recognition algorithm of multi-feature fusion includes a deep learning algorithm and an image algorithm.
[0153] In the implementation device of the present invention, the image algorithm includes a histogram of oriented gradients algorithm and a local binary pattern;
[0154] The smoke recognition module 903 is configured to: extract the feature map of the image using the deep learning algorithm, and convert it into a one-dimensional first vector of a first length through a flattening layer; wherein, the first length is the product of the sizes of the feature map in three dimensions;
[0155] Convert the first vector into a second vector of a second length through a first fully connected layer, and then convert it into a third vector of a third length through a second fully connected layer;
[0156] Extract the vector of the image using the local binary pattern, and convert it into a fourth vector of a fourth length through a third fully connected layer;
[0157] Extract the vector of the image using the histogram of oriented gradients algorithm, and convert it into a fifth vector of a fifth length through a fourth fully connected layer;
[0158] Fuse the third vector, the fourth vector, and the fifth vector to obtain a total vector; wherein, the length of the total vector is the sum of the third length, the fourth length, and the fifth length;
[0159] Convert the total vector into a seventh vector of a seventh length through a fifth fully connected layer, convert the seventh vector into an eighth vector of length 1 through a sixth fully connected layer, and use the modulus of the eighth vector as the probability of the existence of smoke in the image.
[0160] In addition, the specific implementation content of the device in the embodiments of the present invention has been described in detail in the above method, so the repeated content will not be described here.
[0161] Figure 10 An exemplary system architecture 1000 to which the embodiments of the present invention can be applied is shown.
[0162] As Figure 10 shown, the system architecture 1000 may include terminal devices 1001, 1002, 1003, a network 1004, and a server 1005 (merely examples). The network 1004 is used as a medium to provide a communication link between the terminal devices 1001, 1002, 1003 and the server 1005. The network 1004 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0163] Users can use the terminal devices 1001, 1002, 1003 to interact with the server 1005 through the network 1004 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 1001, 1002, 1003.
[0164] The terminal devices 1001, 1002, 1003 may be various electronic devices having display screens and supporting web browsing, and the server 1005 may be a server providing various services.
[0165] It should be noted that the method provided in the embodiment of the present invention is generally executed by the server 1005 , and accordingly, the device is generally set in the server 1005 .
[0166] It should be understood that Figure 10 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0167] Reference below Figure 11 , which shows a schematic diagram of the structure of a computer system 1100 of a terminal device suitable for implementing an embodiment of the present invention. Figure 11 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0168] like Figure 11 As shown, the computer system 1100 includes a central processing unit (CPU) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage part 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the system 1100 are also stored. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0169] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read therefrom is installed into the storage section 1108 as needed.
[0170] In particular, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1109, and / or installed from the removable medium 1111. When the computer program is executed by the central processing unit (CPU) 1101, the above functions defined in the system of the present invention are executed.
[0171] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0173] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a detection module, a spacing determination module, and a smoke recognition module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the detection module can also be described as a "human body and smoke body detection module".
[0174] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, the device includes:
[0175] Determine the human body area in the image and detect whether there is a smoke body in each human body area;
[0176] If a smoke body is detected, identify and calculate the spacing between the position of the smoke body and a preset part of the human body. If the spacing is less than or equal to a preset threshold, it is determined that there is a smoking behavior; where the preset part of the human body is the position of the mouth or the tip of the nose;
[0177] If no smoke body is detected or the spacing is greater than the preset threshold, identify whether there is smoke in the image. If there is, it is determined that there may be a smoking behavior; otherwise, it is determined that there is no smoking behavior.
[0178] According to the technical solution of the embodiments of the present invention, in view of the related problems of existing computer vision algorithms in smoking behavior detection, based on the strict definition of smoking behavior and the possible accompanying smoke phenomenon, a complete smoking warning algorithm scheme is designed.
[0179] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A smoking behavior detection method, characterized in that, Including: Determine the human body regions in the image and detect whether there is a smoke body in each human body region; If a smoke body is detected, identify and calculate the distance between the position of the smoke body and a preset part of the human body. If the distance is less than or equal to a preset threshold, it is determined that there is a smoking behavior, including: determining the coordinate value of the upper left corner of a single smoke body region, and calculating the distance between the upper left corner and the preset part of the human body in the vertical direction; calculating the product of the width of a single human body region and a preset distance measurement coefficient. If the distance is less than or equal to the product, it is determined that there is a smoking behavior in the single smoke body region; wherein, the preset part of the human body is the position of the mouth or the tip of the nose; If no smoke body is detected or the distance is greater than the preset threshold, identify whether there is smoke in the image. If there is, it is determined that there may be a smoking behavior, otherwise it is determined that there is no smoking behavior.
2. The method according to claim 1, wherein, The determining of the human body regions in the image further includes: For a single human body region, determine the width and height of the single human body region, and use the product of the width and an expansion coefficient as the expansion value in the horizontal direction; Based on the expansion value, expand the single human body region left and right in the horizontal direction to obtain an expanded single human body region.
3. The method according to claim 2, wherein After obtaining the expanded human body region, it further includes: Based on the resolution of the image in the horizontal and vertical directions, correct the boundary coordinates of the expanded single human body region to obtain an expanded and corrected single human body region.
4. The method according to claim 2, wherein The width and height are the width and height of the smallest regular rectangle that can cover the single human body region.
5. The method according to claim 1, wherein The detecting whether there is a smoke body in each human body region includes: For a single human body region, use a first convolution module to extract the first image feature of the single human body region, and then use a first detection module to detect whether the first image feature contains a smoke body of a first size. If it does, determine the smoke body position information; Use a first transition module to perform a transition process on the first image feature, use a second convolution module to extract a second image feature from the transitioned first image feature, and then use a second detection module to detect whether the second image feature contains a smoke body of a second size. If it does, determine the smoke body position information; Use a second transition module to perform a transition process on the second image feature, use a third convolution module to extract a third image feature from the transitioned second image feature, and then use a third detection module to detect whether the third image feature contains a smoke body of a third size. If it does, determine the smoke body position information; Summarize the above detection information to obtain the detection results and position information of smoke bodies of different sizes in the single human body region.
6. The method according to claim 5, wherein It further includes: Obtain all the marked anchor box sizes, and obtain multiple anchor box sizes with the largest clustering results through a clustering method; Allocate the multiple anchor box sizes to the first detection module, the second detection module, and the third detection module, and identify the smoke body region by fine-tuning the anchor box sizes.
7. The method according to claim 1, characterized in that, It further includes: Using a human key point algorithm, all human key points in the image are detected, and each key point is clustered into the corresponding individual area to determine the preset part of each human area; wherein, the preset part is the tip of the nose.
8. The method according to claim 1, characterized in that, The preset part is the tip of the nose.
9. The method according to claim 1, characterized in that, Identifying whether there is smoke in the image includes: Based on a smoke recognition algorithm of multi-feature fusion, the probability of the image having smoke is recognized; wherein, the smoke recognition algorithm of multi-feature fusion includes a deep learning algorithm and an image algorithm.
10. The method according to claim 9, wherein The image algorithm includes a histogram of oriented gradients algorithm and local binary patterns; The smoke recognition algorithm based on multi-feature fusion for recognizing the probability that the image has smoke includes: Using the deep learning algorithm to extract the feature map of the image and converting it into a one-dimensional first vector of a first length through a flattening layer; wherein, the first length is the product of the sizes of the feature map in three dimensions; Converting the first vector into a second vector of a second length through a first fully connected layer, and then converting it into a third vector of a third length through a second fully connected layer; Using the local binary patterns to extract the vector of the image and converting it into a fourth vector of a fourth length through a third fully connected layer; Using the histogram of oriented gradients algorithm to extract the vector of the image and converting it into a fifth vector of a fifth length through a fourth fully connected layer; Fusing the third vector, the fourth vector, and the fifth vector to obtain a total vector; wherein, the length of the total vector is the sum of the third length, the fourth length, and the fifth length; Converting the total vector into a seventh vector of a seventh length through a fifth fully connected layer, converting the seventh vector into an eighth vector of length 1 through a sixth fully connected layer, and taking the modulus of the eighth vector as the probability of there being smoke in the image.
11. A smoking behavior detection device, characterized in that, Includes: A detection module for determining the human area in the image and detecting whether there is a smoke body in each human area; A spacing judgment module for, if a smoke body is detected, identifying and calculating the spacing between the position of the smoke body and the preset part of the human body, and if the spacing is less than or equal to a preset threshold, determining that there is a smoking behavior, including: determining the coordinate value of the upper left corner of a single smoke body area and calculating the spacing between the upper left corner and the preset part of the human body in the vertical direction; calculating the product of the width of a single human area and a preset distance measurement coefficient, and if the spacing is less than or equal to the product, determining that there is a smoking behavior in the single smoke body area; wherein, the preset part of the human body is the position of the mouth or the tip of the nose; A smoke recognition module for, if no smoke body is detected or the spacing is greater than the preset threshold, identifying whether there is smoke in the image, and if there is, determining that there may be a smoking behavior, otherwise determining that there is no smoking behavior.
12. An electronic device, characterized in that, Includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-10.
13. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
Traffic violation detection method and system
CN106530730A
Target personnel detection method, device and equipment, and storage medium
CN111611966A